Shuichiro Ogawa
日本語

Notes · updated 2026-07-28

What Is a Reviewer Looking At When They Write “Insufficient Novelty”?

The most dispiriting rejection in a review report is neither an error of method nor a small sample, but “insufficient novelty.” But which part of the paper was that sentence written from? The passage that grounded the judgment is often left unstated in the report.

On the practitioner’s side there is a view that folds this ambiguity away. In the architecture of a paper, it says, there are only five places where novelty can reside. Related Work, Problem Definition, Materials, Methods, Results. Of these, Related Work belongs to the review article, so a research paper can compete only on the remaining four. One alone is weak, but with “decent novelty” in two or more, reviewers find the paper hard to reject and other researchers will cite it. And competing on Results is to be avoided, because the money and labor it takes are enormous.

As a rule of thumb, the view works well. The trouble is that the boundary is invisible: how much of it can be backed by the literature, and where it turns into a survival tactic. Each of the four loci has its own independent scholarly lineage. Problem definition has the research on problematization, materials has strategic research materials, methods has tools-to-theories. Yet a formulation that lines the four up as “components of a paper” and maps sections onto types of contribution is nowhere to be found in the literature.

Who Has Counted the Types of Contribution, and How?

The work of counting kinds of contribution has itself been repeated on the scholarly side. Wobbrock and Kientz classified research contributions in HCI into seven types: empirical, artifact, methodological, theoretical, dataset, survey, and opinion1. Oulasvirta and Hornbæk, building on Laudan’s account of problem solving, divided the problems HCI addresses into three types—empirical, conceptual, and constructive—and showed that the criteria for a good solution differ by type2. The CHI submission guide lists eight types: artifact, understanding users, systems and tools, methodology, theory, innovation and vision, argument, and validation and replication3. Whetten decomposed theoretical contribution into four elements—What, How, Why, and Who-Where-When—and held that a contribution can stand on any one of them4. Corley and Gioia typologized contribution along two axes: originality (incremental or revelatory) and utility (scientific or practical)5.

Every one of these typologies cuts by what a paper offers. None cuts by which section of the paper it is written in. Analysis that takes the section as its unit belongs to a different lineage. Kanoksilapatham decomposed all IMRAD sections of 60 biochemistry articles into rhetorical moves and proposed fifteen of them6. But those moves are units of the act of writing, such as summarizing prior work or stating a procedure, not kinds of contribution. The IMRAD structure itself is a relatively recent format, spreading in the 1950s and standardized in the 1980s7.

In this collection, no scholarly formulation that normatively maps each IMRaD section onto a type of contribution was found. The four-locus classification is not a summary of a scholarly formulation; it stands on its own as practitioner knowledge. There is a reason the scholarly typologies ignore sections and look only at the content of the contribution. The same kind of contribution is written in different sections in different fields.

The first thing the four-locus hypothesis cuts away is Related Work. Organizing prior work is the review article’s job, so a research paper’s novelty does not reside there. As the place where novelty is born, that is right. But as the place where novelty reaches the reader, the story reverses.

Swales’s CARS model describes the introduction of a research article in three moves8. Establishing a territory, establishing a niche, occupying the niche. Novelty rises as language in the second move: the position where, having laid out prior work, the writer points to what is missing there, what is contradictory, what should be continued. Swales further divided this move into four steps: counter-claiming, indicating a gap, question-raising, and continuing a tradition. Locke and Golden-Biddle qualitatively analyzed articles in organization studies and showed that introductions rhetorically construct opportunities for contribution through two processes, the construction of an intertextual field and problematization9. Luzón Marco reported that in computer science papers, authors weave their own work into existing knowledge through evaluative vocabulary and lexical cohesion, constructing novelty in a problem-and-solution pattern10. Kwan showed that the literature review chapter of a doctoral thesis, while overlapping with the CARS moves of an introduction, carries its own function of claiming novelty11. Hyland demonstrated that self-citation and self-reference are used to bolster the credibility of knowledge claims and to promote personal contribution12. Patriotta, writing as an editor, argues that novelty is socially constructed as entry into a conversation through citation13.

Related Work is not a locus that generates novelty, but it is a locus where novelty is claimed. Even with novelty somewhere in the four loci, a paper that fails to claim it in Related Work is read as having none. Returning to the review report at the opening, at least part of the material for that judgment is taken from this section. The language of claiming has itself inflated. The frequency of positive words in PubMed abstracts rose from 2.0% in 1974–80 to 17.5% in 2014, and novel, robust, innovative, and unprecedented showed some of the largest increases14.

Problem Definition as a Locus

Of the four loci, the one with the thickest scholarly backing is the definition of the problem. Not only because there is a theory that prescribes. It is also because there is evidence that almost no one actually uses that prescription.

Sandberg and Alvesson analyzed 52 articles in organization studies and reported that nearly every construction of a research question was gap-spotting15. Gap-spotting is the practice of finding and filling omissions, inadequacies, inaccuracies, and unapplied areas in the literature, and it divides into these four types. Problematization, by contrast, identifies and challenges the very assumptions on which existing theory leans16. The assumptions to be challenged are organized into five layers: assumptions internal to a theory, the grounds a field shares, paradigms, ideologies, and the assumptions of the field. Alvesson and Sandberg also analyzed the institutional machinery that boxes researchers into their specialty, and set out three strategies of escape: changing the box, jumping out of the box, and transcending the box17.

The operation of challenging assumptions changes in its result with how far the challenge goes. Davis formulated an “interesting” theory as one that denies part of the assumptions its audience holds18. Deny the assumptions entirely and the theory is taken as absurd; deny nothing and it is taken as obvious. Bartunek et al. confirmed this condition in a survey of AMJ editorial board members, reporting that at the center of interestingness lies the partial denial of assumptions19.

The problem-definition locus has reinforcement from sociology and the philosophy of science as well. Merton placed specified ignorance (the express recognition of what is not yet known but needs to be known in order to lay the foundation for still more knowledge) as the unit of a question20. Not a vague “we don’t know,” but putting non-knowledge into a form where one can name what is unknown, is itself an intellectual achievement. Laudan defined scientific progress as problem solving, distinguished empirical from conceptual problems, and argued that the solving of conceptual problems drives progress21. Nickles analyzed a problem as the set of constraints a solution must satisfy22. On this view, making a question new means recombining constraints. Dorst formulated the core of design reasoning as abduction-2 (not fitting a situation to a known frame but creating the frame itself)23, describing from the design side the same operation as the recombination of constraints. A demonstration that applies this generator to the board of AI research is in ai-research-gaps-abduction and problem-definition-metagame.

While prescriptions abound, practice is thin. Wald et al. content-analyzed 124 articles in five higher education journals and reported that although about 80% carried a gap statement, many were implicit, and about 27% lacked any justification for why the gap mattered24. Pointing out a gap and arguing that filling it is worth something are different jobs. What the four-locus hypothesis calls “novelty in Problem Definition” is best read as pointing to the latter.

Materials as a Locus

The materials locus has been the least discussed of the four. It has only gone undiscussed; as a mechanism it was identified long ago.

Merton placed strategic research materials in the same essay as specified ignorance20. These are research objects that exhibit the phenomenon one wants to explain in a peculiarly advantageous form. Here the order is not that the question comes first and the materials are chosen, but that the materials set the resolution of the question. Rheinberger, with the concepts of the experimental system and epistemic things, described the process by which an experimental system is set up before the question is clarified, and the unexpected behavior of things produces the question25. Galison showed that the material culture of instruments and experimental practice branches the very style of questioning26. de Solla Price argued that scientific instrumentation is the medium that advances science and technology in both directions27. Steinle identified exploratory experimentation, aimed neither at theory-driven work nor at hypothesis testing but at generating concepts and questions, as a type28.

Institutional recognition of the materials locus is very recent. NeurIPS created the Datasets and Benchmarks track in 2021. The reason officially given was that fewer than five dataset papers a year were being accepted at the main conference29. The call for papers states explicitly that audits of existing datasets and collection practices themselves are welcome as contributions30. In 2026 the track was renamed Evaluations & Datasets, and negative results, critical analyses, and methodology were officially recognized as principal contributions31. ACL Rolling Review separates review into two axes, Soundness and Excitement, placing no standalone score for novelty32, while setting Datasets and Software as independent scoring items33.

The materials locus looks attractive within the four-locus hypothesis because its cost is moderate while its spillover into the other loci is large. New materials bring with them questions that only those materials can answer. But the very existence of the new track also shows that material contributions were long hard to get valued in the mainstream review frame.

Methods as a Locus

The methods locus has a self-propagating quality the others lack. When a method is new, the questions that method generates turn out new as well.

Gigerenzer traced the history by which the tools of statistical inference (analysis of variance, significance testing) turned into metaphors for theories of mind itself, and formulated the heuristic of tools-to-theories34. What was introduced as a tool comes in time to be accepted as a model of the mind. Greenwald argued for a two-way synergy in which methods produce unprecedented data and that data induces unprecedented theory, and showed the weight of methodological contributions in a breakdown of Nobel Prizes35. The same Greenwald and colleagues analyzed the conditions under which an orientation toward theory testing breeds confirmation bias and obstructs progress, recommending condition-seeking, the search for a theory’s boundary conditions, as an alternative36.

The assumption that the methods locus and the problem locus can be moved independently carries a caveat from cognitive science. Klahr and Dunbar modeled scientific discovery as a dual search of hypothesis space and experiment space37. Search in the two spaces proceeds alternately, and a move in one changes the range that can be moved in the other. The four-locus hypothesis treats the loci as parallel options, but at least Problem Definition and Methods are observed to move in tandem rather than to be chosen. An account of methodological plurality seen from the design research side is in ai-research-gaps-method-pluralism.

The Cost of the Results Locus

The judgment that competing on Results is to be avoided because the cost is enormous can be supported on the evidence. Before that, though, the word Results has to be split, because it points to two different things.

One is a design that stacks resources in order to produce a large result. Large randomized controlled trials, long-term follow-up, and mass data collection fall here. The other is running into an unexpected result. Merton called the pattern in which the observation of unanticipated, anomalous, and strategic data occasions a new theory the serendipity pattern38. Dunbar carried out a year of participant observation in four molecular biology laboratories and reported that the handling of unexpected results and analogy are the central mechanisms that generate new questions39. Alvesson and Kärreman set the discrepancy between empirical data and expectation (breakdown, mystery) as the starting point of theoretical insight40.

Only the former costs enormously. The latter comes out as a by-product of investment in the materials and methods loci, and cannot be budgeted for on its own. What the four-locus hypothesis says to avoid is limited to the former.

Even for that former, cost is not the only reason to avoid it. Foster et al. classified research strategies in a MEDLINE network of chemical relations and showed that papers taking an innovative strategy stay at about 14% of the whole41. The rest are conservative strategies that search the neighborhood of known substances. The distribution is one in which an innovative strategy carries a high risk of being ignored while also carrying higher probabilities of high impact and of prizes. Rzhetsky et al., on the same data, showed that scientists’ search diverges from the strategy that would maximize collective discovery42. The resources put into corroborating results are scarcer still. Hornbæk et al. surveyed 891 papers across four publication venues and reported that 3% attempted to replicate a prior result43. The RepliCHI panel was put forward as a problem statement that CHI’s culture of prizing novelty suppresses replication research44. The explicit addition of validation and replication to CHI’s contribution types amounts to an institutional response to that suppression3.

Is the Prescription of Two or More Loci Supported?

At the center of the four-locus hypothesis is a claim about combination. One locus alone is weak; with decent novelty across two or more, a paper is hard to reject. There is evidence on both sides, supporting and squarely contradicting.

The supporting side is a run of studies that put the combination of knowledge into numbers. Uzzi et al. analyzed 17.9 million papers by the z-scores of cited-journal pairs and showed that the highest-impact papers hold both exceptional conventionality and an intrusion of atypical combinations45. Papers of this type were twice as likely to become highly cited. Shi and Evans predicted combinations of research contents and contexts with hypergraphs and reported that the most surprising combinations raise the rate of high citation greatly over a random baseline, and that such combinations arise more readily from outsiders in distant fields46. Arts and Veugelers showed in patent networks that the probability of a breakthrough rises only when familiar components are combined in a new way, and does not rise with conventional recombination47. Weitzman formulated the combinatorial proliferation of ideas as the source of growth48.

What these measure, though, is the combination of knowledge elements, not the combination of loci within a paper. Pairs of cited journals, or the correspondence of content and context, have not been substituted for the pair of sections Problem Definition and Methods. Reading them over into loci is an analogy, not a direct test.

Even within the supporting side, conditions attach that amend the prescription. Uzzi et al.’s result takes a shape closer to keeping most of the work conventional and inserting a single atypical point than to making both extremely new. Fleming showed that an unprecedented combination lowers the mean value of an invention while greatly raising its variance49. Boudreau et al., in a randomized-assignment evaluation experiment on research proposals, showed that evaluators score both proposals too close to their own specialty and proposals too high in novelty lower50. Criscuolo et al., on actual data from R&D selection panels, report that organizations fund projects of moderate novelty more readily51. Johnson and Proudfoot showed, through experiments and archival analysis, that the more novel an idea, the greater the variance among evaluators, and that the size of that variance itself works as a risk signal that lowers willingness to invest52. Wang et al. tracked about 1.05 million papers across all fields from 2001 and reported that highly novel papers are disadvantaged in short-term citations and in journal indicators, while over the long run they are more likely to enter the top 1% and are recognized with a delay in adjacent fields53. Trapido shows that the recognition premium for novelty is concentrated among authors who have already built a record of novelty54.

The contradicting evidence is a single item, but it takes aim at the four-locus hypothesis head-on. Zhao et al. used an LLM to classify 15,322 papers published in Nature Communications along three dimensions—theory, method, and results—and reported that papers holding only results-based novelty are cited more than papers holding all three dimensions, and are also more likely to enter the top ranks55. The result runs against both the claim that combination beats a single dimension and the claim that the Results locus should be avoided. This is a preprint, however, and its peer-review status could not be confirmed. And in this collection, no peer-reviewed empirical study directly testing that a combination of novelty across multiple dimensions raises evaluation above a single dimension was found. The central prescription of the four-locus hypothesis stands, on the supporting and the disconfirming side alike, without direct peer-reviewed evidence.

The part that says “with two or more loci, reviewers find the paper hard to reject” carries a caveat from another direction as well. The premise that hard-to-rejectness can be engineered tacitly assumes that review judgments are stable. Cortes and Lawrence, reanalyzing the 2014 NeurIPS consistency experiment, reported that when two committees independently reviewed 170 papers, the rate of disagreement on acceptance was 26%, and an accepted paper’s probability of surviving a second review was about 50%56. In the 2021 re-run the disagreement rate was 23%; the level of noise did not change as the scale grew57. The effect of securing two loci rides on top of noise at this level.

What Spills Over When the Scheme Is Brought into Design Research?

Apply the four-locus classification to design research as it stands, and contributions appear that will not fit inside it. What spills over is the artifact.

Frayling divided research in art and design into three types: into, through, and for58. In through, that is, research carried out by making, the made thing itself becomes the bearer of knowledge. Zimmerman et al. positioned research through design as a research method and set the designer’s strength in handling underconstrained problems as the source of contribution59. Gaver argues that the theory this method produces may be provisional and context-dependent, and need not aim at convergent theory60. Höök and Löwgren defined strong concepts as a form of intermediate-level knowledge, more abstract than a particular artifact yet short of general theory61. Gaver and Bowers proposed annotated portfolios, a form of annotation over a body of work that transmits knowledge without passing through theorization62. Cross located design as a third mode of knowing alongside the sciences and the humanities63, and Stolterman argued that because the complexity of design practice differs from scientific complexity, research should function as preparation rather than prescription64.

The artifact is not Materials. Materials are used in order to answer a question, whereas the artifact of research through design sits on the side of the answer. Nor is it Results. It is not a measured consequence but a possibility presented as a thing. The four loci are a scheme premised on measurement, not a scheme premised on making. The methodological terrain of design research is laid out in design-research-methodology-mainstream and design-epistemology-methodology-foundations.

On the design research side, too, what counts as a contribution is unsettled. Cash conducted a citation analysis of 467 papers published in Design Studies (2004 to 2018), showed that theory building predicts impact better than anything else, and diagnosed a stagnation in theoretical development65. Davis et al., in a consensus document by supervisors from nine doctoral programs, split the requirement for contribution by degree: generalizable knowledge for the PhD, a contribution to design judgment in a specific context for the DDes66. To use the four loci, one has to add artifact contribution explicitly as a separate slot.

Scoring the Four-Locus Hypothesis

Having held it against the literature, here is how far each claim of the four-locus hypothesis is supported.

  • The five-locus structural partition: nothing corresponds to it as a scholarly formulation. No literature normatively mapping IMRaD sections onto types of contribution was found. Independent lineages corresponding to each locus do exist, however.
  • Related Work is not a locus for a research paper: supported as a locus of generation. Denied as a locus of claiming (Move 2 of CARS, Locke and Golden-Biddle, Luzón Marco, Kwan, Hyland, Patriotta).
  • Problem Definition works as a locus: strongly supported. Both the prescriptions (problematization, specified ignorance, conceptual problems, abduction-2) and the evidence for the scarcity of practice (the analysis of 52 articles, the gap-statement analysis of 124 articles) are in place.
  • Materials works as a locus: supported as a mechanism (strategic research materials, experimental systems, instrumentation, exploratory experimentation). Recognition as an institution is new, dating from 2021.
  • Methods works as a locus: supported (tools-to-theories, the synergy by which methods induce theory). But the premise that it can be moved independently of Problem Definition carries a caveat (dual space search).
  • Competing on Results costs enormously: supported for a design that wins by scale (innovative strategies at 14%, replication research at 3%). The generation of questions out of unexpected results is a different thing, and the argument from cost does not apply to it.
  • Two or more loci make a paper hard to reject: there is indirect support from measurements of knowledge combination, but no direct test of combinations of loci exists in the peer-reviewed literature. The only direct measurement, a preprint, reports the opposite. On top of that, a run of suppressors operates on the evaluation side: the distance penalty, the preference for the moderate, variance among evaluators, and short-term disadvantage.

Gaps in the Collection

  1. No empirical study directly measuring whether problematization-type questions are advantaged in citations or in acceptance was found.
  2. No peer-reviewed empirical study directly testing that a combination of novelty across multiple dimensions raises evaluation above a single dimension was found.
  3. No study comparing the citation returns of experimentally expensive and inexpensive research was found. Direct evidence on the cost-effectiveness of the Results locus is missing.
  4. No peer-reviewed paper quantitatively verifying that the renewal of measuring instruments opened up new bodies of research questions was found. The literature that comes close is all historical or philosophical analysis.
  5. No intervention study on how to pose questions, specific to design education, was found. The nearest work stays within agriculture and medicine.
  6. No dedicated philosophical treatise explicitly distinguishing novelty of the question from novelty of the answer was found.
  7. No scholarly formulation normatively mapping each IMRaD section onto a kind of contribution was found.

With LLMs entering on the side of question generation, part of this gap looks likely to be filled before long. Si et al. reported, from blind evaluations by more than 100 NLP researchers, that LLM-generated ideas scored significantly higher in novelty than those of human experts, and slightly lower in feasibility67. On the evaluation side too, there is an attempt to extract statements about novelty from review reports and score them automatically68, and a benchmark built from 1,684 paper-and-review pairs69. A review of the methods for measuring novelty has also been compiled70. Which of the four loci stays scarce once generation becomes cheap has not yet been measured.

Unverified Items

  • The specific effect size in Boudreau et al. (2016) (the drop in evaluation rank per standard deviation of novelty). Not confirmed in the peer-reviewed text. Only the direction is adopted in the body.
  • The DOI and page range of Criscuolo et al. (2017). Undetermined because the publisher site returns 403.
  • The peer-review status of Zhao et al. (2026). It is stated to be a submission to AII-EEKE 2026, but this could not be confirmed.
  • The wording of Merton’s (1987) definitions of specified ignorance and strategic research materials. It relies on secondary quotation, and the original pages have not been checked.
  • The original pages for the serendipity pattern in Merton (1948/1968). They differ by edition.
  • The body of response papers from 2025 to 2026 on the declining trend of the CD index. Primary confirmation of authors, years, and pages is incomplete.

References

Constructing Research Questions and Theoretical Contribution

Novelty Seen from Scientometrics

Philosophy of Science, Sociology of Science, Cognitive Science

Contribution Taxonomies and Review Systems

The Rhetoric of Novelty and the Craft of Questions

Contribution in Design Research

LLMs and the Generation of Questions

Footnotes

  1. Wobbrock, J. O. & Kientz, J. A. (2016). Research contributions in human-computer interaction. Interactions 23(3):38–44. https://doi.org/10.1145/2907069

  2. Oulasvirta, A. & Hornbæk, K. (2016). HCI research as problem-solving. Proc. CHI 2016, 4956–4967. https://doi.org/10.1145/2858036.2858283

  3. CHI 2025 Program Chairs (2024). Contributions to CHI. https://chi2025.acm.org/contributions-to-chi/ 2

  4. Whetten, D. A. (1989). What constitutes a theoretical contribution? Academy of Management Review 14(4):490–495. https://doi.org/10.5465/258554

  5. Corley, K. G. & Gioia, D. A. (2011). Building theory about theory building. Academy of Management Review 36(1):12–32. https://doi.org/10.5465/amr.2009.0486

  6. Kanoksilapatham, B. (2005). Rhetorical structure of biochemistry research articles. English for Specific Purposes 24(3):269–292. https://doi.org/10.1016/j.esp.2004.08.003

  7. Sollaci, L. B. & Pereira, M. G. (2004). The introduction, methods, results, and discussion (IMRAD) structure: a fifty-year survey. Journal of the Medical Library Association 92(3):364–371. https://pmc.ncbi.nlm.nih.gov/articles/PMC442179/

  8. Swales, J. M. (1990). Genre Analysis: English in Academic and Research Settings. Cambridge University Press. ISBN 9780521338134. https://archive.org/details/genreanalysiseng0000swal

  9. Locke, K. & Golden-Biddle, K. (1997). Constructing opportunities for contribution. Academy of Management Journal 40(5):1023–1062. https://doi.org/10.2307/256926

  10. Luzón Marco, M. J. (2000). The construction of novelty in computer science papers. Revista Alicantina de Estudios Ingleses 13:123–140. https://doi.org/10.14198/raei.2000.13.10

  11. Kwan, B. S. C. (2006). The schematic structure of literature reviews in doctoral theses of applied linguistics. English for Specific Purposes 25(1):30–55. https://doi.org/10.1016/j.esp.2005.06.001

  12. Hyland, K. (2003). Self-citation and self-reference: credibility and promotion in academic publication. JASIST 54(3):251–259. https://doi.org/10.1002/asi.10204

  13. Patriotta, G. (2017). Crafting papers for publication: novelty and convention in academic writing. Journal of Management Studies 54(5):747–759. https://doi.org/10.1111/joms.12280

  14. Vinkers, C. H., Tijdink, J. K., Otte, W. M. (2015). Use of positive and negative words in scientific PubMed abstracts between 1974 and 2014. BMJ 351:h6467. https://doi.org/10.1136/bmj.h6467

  15. Sandberg, J. & Alvesson, M. (2011). Ways of constructing research questions: gap-spotting or problematization? Organization 18(1):23–44. https://doi.org/10.1177/1350508410372151

  16. Alvesson, M. & Sandberg, J. (2011). Generating research questions through problematization. Academy of Management Review 36(2):247–271. https://doi.org/10.5465/amr.2009.0188

  17. Alvesson, M. & Sandberg, J. (2014). Habitat and habitus: boxed-in versus box-breaking research. Organization Studies 35(7):967–987. https://doi.org/10.1177/0170840614530916

  18. Davis, M. S. (1971). That’s Interesting! Philosophy of the Social Sciences 1(2):309–344. https://doi.org/10.1177/004839317100100211

  19. Bartunek, J. M., Rynes, S. L., Ireland, R. D. (2006). What makes management research interesting, and why does it matter? Academy of Management Journal 49(1):9–15. https://doi.org/10.5465/AMJ.2006.20785494

  20. Merton, R. K. (1987). Three fragments from a sociologist’s notebooks: establishing the phenomenon, specified ignorance, and strategic research materials. Annual Review of Sociology 13:1–28. https://doi.org/10.1146/annurev.so.13.080187.000245 (The wording of the definitions is quoted at second hand; the original pages have not been checked [primary-source verification needed].) 2

  21. Laudan, L. (1977). Progress and Its Problems: Towards a Theory of Scientific Growth. University of California Press. https://www.ucpress.edu/books/progress-and-its-problems/paper

  22. Nickles, T. (1981). What is a problem that we may solve it? Synthese 47(1):85–118. https://doi.org/10.1007/BF01064267

  23. Dorst, K. (2011). The core of ‘design thinking’ and its application. Design Studies 32(6):521–532. https://doi.org/10.1016/j.destud.2011.07.006

  24. Wald, N., Harland, T., Daskon, C. (2024). The gap statement and justification in higher education research. European Journal of Higher Education 14(2):308–323. https://doi.org/10.1080/21568235.2023.2189132

  25. Rheinberger, H.-J. (1997). Toward a History of Epistemic Things: Synthesizing Proteins in the Test Tube. Stanford University Press. https://www.sup.org/books/theory-and-philosophy/toward-history-epistemic-things

  26. Galison, P. (1997). Image and Logic: A Material Culture of Microphysics. University of Chicago Press. https://galison.scholars.harvard.edu/publications/image-and-logic

  27. de Solla Price, D. J. (1984). The science/technology relationship, the craft of experimental science, and policy for the improvement of high technology innovation. Research Policy 13(1):3–20. https://doi.org/10.1016/0048-7333(84)90003-9

  28. Steinle, F. (1997). Entering new fields: exploratory uses of experimentation. Philosophy of Science 64(Proceedings):S65–S74. https://doi.org/10.1086/392587

  29. NeurIPS Organizers (2021). Announcing the NeurIPS 2021 Datasets and Benchmarks Track. https://blog.neurips.cc/2021/04/07/announcing-the-neurips-2021-datasets-and-benchmarks-track/

  30. NeurIPS Organizers (2021). NeurIPS 2021 Datasets and Benchmarks Track: Call for Papers. https://neurips.cc/Conferences/2021/CallForDatasetsBenchmarks

  31. NeurIPS Organizers (2026). NeurIPS 2026 Evaluations & Datasets Track: Call for Papers. https://neurips.cc/Conferences/2026/CallForEvaluationsDatasets

  32. ACL Rolling Review. Reviewer guidelines. https://aclrollingreview.org/reviewerguidelines

  33. ACL Rolling Review. Review form. https://aclrollingreview.org/reviewform

  34. Gigerenzer, G. (1991). From tools to theories: a heuristic of discovery in cognitive psychology. Psychological Review 98(2):254–267. https://doi.org/10.1037/0033-295X.98.2.254

  35. Greenwald, A. G. (2012). There is nothing so theoretical as a good method. Perspectives on Psychological Science 7(2):99–108. https://doi.org/10.1177/1745691611434210

  36. Greenwald, A. G., Pratkanis, A. R., Leippe, M. R., Baumgardner, M. H. (1986). Under what conditions does theory obstruct research progress? Psychological Review 93(2):216–229. https://doi.org/10.1037/0033-295X.93.2.216

  37. Klahr, D. & Dunbar, K. (1988). Dual space search during scientific reasoning. Cognitive Science 12(1):1–48. https://doi.org/10.1207/s15516709cog1201_1

  38. Merton, R. K. (1948/1968). Social Theory and Social Structure (enlarged ed.). Free Press. https://archive.org/details/socialtheorysoci0000mert (The original pages depend on the edition [primary-source verification needed].)

  39. Dunbar, K. (1995). How scientists really reason: scientific reasoning in real-world laboratories. In Sternberg & Davidson (eds.), The Nature of Insight, MIT Press, 365–395. https://direct.mit.edu/books/edited-volume/3958/chapter/165254/

  40. Alvesson, M. & Kärreman, D. (2007). Constructing mystery: empirical matters in theory development. Academy of Management Review 32(4):1265–1281. https://doi.org/10.5465/amr.2007.26586822

  41. Foster, J. G., Rzhetsky, A., Evans, J. A. (2015). Tradition and innovation in scientists’ research strategies. American Sociological Review 80(5):875–908. https://doi.org/10.1177/0003122415601618

  42. Rzhetsky, A., Foster, J. G., Foster, I. T., Evans, J. A. (2015). Choosing experiments to accelerate collective discovery. PNAS 112(47):14569–14574. https://doi.org/10.1073/pnas.1509757112

  43. Hornbæk, K., Sander, S. S., Bargas-Avila, J., Simonsen, J. G. (2014). Is once enough? On the extent and content of replications in human-computer interaction. Proc. CHI 2014. https://doi.org/10.1145/2556288.2557004

  44. Wilson, M. L., Mackay, W. E., Chi, E. H., Bernstein, M., Russell, D., Thimbleby, H. (2011). RepliCHI: CHI should be replicating and validating results more. CHI EA ‘11, 463–466. https://doi.org/10.1145/1979742.1979491

  45. Uzzi, B., Mukherjee, S., Stringer, M., Jones, B. (2013). Atypical combinations and scientific impact. Science 342(6157):468–472. https://doi.org/10.1126/science.1240474

  46. Shi, F. & Evans, J. A. (2023). Surprising combinations of research contents and contexts are related to impact and emerge with scientific outsiders from distant disciplines. Nature Communications 14:1641. https://doi.org/10.1038/s41467-023-36741-4

  47. Arts, S. & Veugelers, R. (2015). Technology familiarity, recombinant novelty, and breakthrough invention. Industrial and Corporate Change 24(6):1215–1246. https://doi.org/10.1093/icc/dtu029

  48. Weitzman, M. L. (1998). Recombinant growth. Quarterly Journal of Economics 113(2):331–360. https://doi.org/10.1162/003355398555595

  49. Fleming, L. (2001). Recombinant uncertainty in technological search. Management Science 47(1):117–132. https://doi.org/10.1287/mnsc.47.1.117.10671

  50. Boudreau, K. J., Guinan, E. C., Lakhani, K. R., Riedl, C. (2016). Looking across and looking beyond the knowledge frontier: intellectual distance, novelty, and resource allocation in science. Management Science 62(10):2765–2783. https://doi.org/10.1287/mnsc.2015.2285 (The specific effect size has not been confirmed in the peer-reviewed text [primary-source verification needed]; only the direction is adopted.)

  51. Criscuolo, P., Dahlander, L., Grohsjean, T., Salter, A. (2017). Evaluating novelty: the role of panels in the selection of R&D projects. Academy of Management Journal 60(2). DOI [primary-source verification needed]

  52. Johnson, W. & Proudfoot, D. (2024). Greater variability in judgements of the value of novel ideas. Nature Human Behaviour 8:88–102. https://doi.org/10.1038/s41562-023-01794-4

  53. Wang, J., Veugelers, R., Stephan, P. (2017). Bias against novelty in science: a cautionary tale for users of bibliometric indicators. Research Policy 46(8):1416–1436. https://doi.org/10.1016/j.respol.2017.06.006

  54. Trapido, D. (2015). How novelty in knowledge earns recognition: the role of consistent identities. Research Policy 44(8):1488–1500. https://doi.org/10.1016/j.respol.2015.05.007

  55. Zhao, Y., Yang, C., Wang, Y., Bao, T., Zhang, H., Zhang, C. (2026). Beyond single-dimension novelty: how combinations of theory, method, and results-based novelty shape scientific impact. arXiv:2604.12471. https://arxiv.org/abs/2604.12471 (Peer-review status unconfirmed [primary-source verification needed].)

  56. Cortes, C. & Lawrence, N. D. (2021). Inconsistency in conference peer review: revisiting the 2014 NeurIPS experiment. arXiv:2109.09774. https://arxiv.org/abs/2109.09774

  57. NeurIPS Program Chairs (2021). The NeurIPS 2021 consistency experiment. https://blog.neurips.cc/2021/12/08/the-neurips-2021-consistency-experiment/

  58. Frayling, C. (1993/94). Research in art and design. Royal College of Art Research Papers 1(1). https://researchonline.rca.ac.uk/384/

  59. Zimmerman, J., Forlizzi, J., Evenson, S. (2007). Research through design as a method for interaction design research in HCI. Proc. CHI 2007, 493–502. https://doi.org/10.1145/1240624.1240704

  60. Gaver, W. (2012). What should we expect from research through design? Proc. CHI 2012, 937–946. https://doi.org/10.1145/2207676.2208538

  61. Höök, K. & Löwgren, J. (2012). Strong concepts: intermediate-level knowledge in interaction design research. ACM TOCHI 19(3):Article 23. https://doi.org/10.1145/2362364.2362371

  62. Gaver, W. & Bowers, J. (2012). Annotated portfolios. Interactions 19(4):40–49. https://doi.org/10.1145/2212877.2212889

  63. Cross, N. (1982). Designerly ways of knowing. Design Studies 3(4):221–227. https://doi.org/10.1016/0142-694X(82)90040-0

  64. Stolterman, E. (2008). The nature of design practice and implications for interaction design research. International Journal of Design 2(1):55–65. https://www.ijdesign.org/index.php/IJDesign/article/view/240/148

  65. Cash, P. (2020). Where next for design research? Understanding research impact and theory building. Design Studies 68:113–141. https://doi.org/10.1016/j.destud.2020.03.001

  66. Davis, M., Feast, L., Forlizzi, J., Friedman, K., Ilhan, A., Ju, W. et al. (2023). Responding to the indeterminacy of doctoral research in design. She Ji 9(2):283–307. https://doi.org/10.1016/j.sheji.2023.05.005

  67. Si, C., Yang, D., Hashimoto, T. (2025). Can LLMs generate novel research ideas? A large-scale human study with 100+ NLP researchers. ICLR 2025. https://arxiv.org/abs/2409.04109

  68. Wu, W., Zhang, C., Zhao, Y. (2025). Automated novelty evaluation of academic paper: a collaborative approach integrating human and large language model knowledge. JASIST. https://doi.org/10.1002/asi.70005

  69. Wu, W., Zhao, Y., Wang, Y., Li, S., Shao, J., Long, Y., Zhang, C. (2026). NovBench: evaluating large language models on academic paper novelty assessment. ACL 2026 Findings. https://arxiv.org/abs/2604.11543

  70. Zhao, Y. & Zhang, C. (2025). A review on the novelty measurements of academic papers. Scientometrics 130:727–753. https://doi.org/10.1007/s11192-025-05234-0


← All Notes · Home