Shuichiro Ogawa
日本語

Notes · updated 2026-07-28

How Many Moves Are Available When You Are Told to Add Novelty

Supervisors and reviewers alike demand novelty in the form of a noun. Insufficient, weak, no visible difference from prior work. What the person on the receiving end needs, though, is a verb. Tomorrow morning, at the desk, what should one do to increase novelty?

Advice usually stops at question your assumptions. How to find the assumption that should be questioned does not come with it. Descriptions that close this distance in fact exist in quantity on the literature’s side. They are scattered, though, across management studies, the philosophy of science, scientometrics, applied linguistics, and HCI, and they barely cite one another. This note gathers those scattered strategies into a single ledger and holds them against each other.

Strategies for Making and Strategies for Claiming

What becomes clear as soon as they are gathered is that what has been called a strategy of novelty splits into two kinds.

Strategies of generation: operations that actually produce a finding that does not yet exist. Inverting assumptions, selecting materials, decomposing mechanisms, repurposing tools, and transferring from distant fields belong here.

Strategies of claiming: operations that make the finding produced read as novel. Arranging prior work, pointing to a gap, calibrating certainty, and choosing hype words belong here.

This distinction is not drawn explicitly on the literature’s side. When Medawar wrote in 1963 that the form of the scientific paper is a fraud that conceals the process of discovery, what he was pointing at was the divergence between these two layers. Latour and Woolgar traced, in a laboratory ethnography, the five degrees of modality by which a claim moves from conjecture to fact, and Knorr-Cetina described the paper as a reconstructed product of the experimental process. Claiming is not a record of generation. Even so, methodology textbooks line the two up within the same account of how research proceeds.

Below, the generation side comes first and the claiming side second. The order is not the order of writing but an arrangement meant to keep the two apart.

To Question Assumptions Is to Question Which Assumptions?

The most cited typology of how research questions are constructed is Sandberg and Alvesson’s contrast. Analyzing a sample of articles in management studies, they showed that the dominant way of constructing a question is gap-spotting. Gap-spotting divides further into three subtypes. Confusion spotting (prior work disagrees), neglect spotting (no one has looked), and application spotting (it has not been applied to another object). Each is an operation that points at a white patch on the map the existing literature has laid down.

Set against it is problematization, which takes the assumptions of the map itself as the object of challenge. Alvesson and Sandberg’s contribution is to have divided the assumptions available for challenge into five layers.

  • in-house assumptions: particular assumptions shared by a specific body of research
  • root metaphor assumptions: the underlying metaphors that support a domain
  • paradigmatic assumptions: assumptions at the level of ontology and epistemology
  • ideological assumptions: political and moral premises
  • field assumptions: premises that different schools have come to share

The higher the layer, the larger the return on a challenge and the harder it is to get through. These five layers were the first thing to make the advice to question assumptions operable.

The one who formulated the inversion of assumptions as a general-purpose generator is the older Davis. He collected social theories judged interesting and extracted twelve types of inversion pattern that deny the reader’s assumptions. It is a list of pairs: what seems disordered is in fact ordered, what seems local is in fact global. Abbott’s fractal distinctions are of the same form, exploiting the way the oppositions that have divided a field (quantitative and qualitative, positivist and interpretive) recur within its subdivisions, and creating a new position by applying the axis one level deeper.

Locke and Golden-Biddle described the same operation from the side of how prior work is arranged. There are three types of how prior work is bundled (intertextual coherence)—synthesized, progressive, and noncoherent—and crossing them with the three types of problematization (incomplete, inadequate, incommensurate) constructs an opportunity for contribution. Incomplete is the claim that existing research is unfinished, inadequate that it is unsuitable, and incommensurate that a fundamentally different view is required; the claims grow stronger toward the end.

Naming What Is Not Known

In a 1987 paper Merton named two strategies. One is specified ignorance. Rather than vaguely conceding that something is not understood, locating precisely what is not known within the structure of prior knowledge itself prescribes the next question. Specifying ignorance resembles gap-spotting on the surface and differs inside. Gap-spotting points at a gap in the literature, while specified ignorance points at a place where the theory demands an answer and there is none.

The other strategy Merton put forward is strategic research materials. One chooses the object of research as the material in which the theoretical question appears most clearly. The claim is that the selection of what to look at is itself a contribution, and it connects directly to the Materials locus discussed later.

A path that generates novelty from the side of materials exists independently in the philosophy of science as well. Rheinberger’s epistemic things is the account that keeping an object whose identity is not yet fixed inside an experimental system produces novelty. Hacking’s argument about intervention likewise holds that experiment can produce new objects independently of theory. Neither presupposes the order in which a theory is set up first and then checked.

Decomposing a Mechanism

The philosophy of the life sciences and of cognitive science has formulated discovery as the search for mechanisms. Bechtel and Richardson drew from the history of science the finding that understanding a complex system reduces to two operations, decomposition and localization. One divides the system into subordinate functions and maps each function onto a part of the structure. Changing how the decomposition is cut becomes, in itself, a new theory.

Darden wrote the same line more procedurally. The three strategies of mechanism discovery are schema instantiation (filling in specifics on a known mechanism schema), modular subassembly (building part by part), and forward-backward chaining (working in the middle from a known start and end). She further places the maturity of an explanation on three stages, how-possibly → how-plausibly → how-actually, requiring that the stage of one’s claim be made explicit. Dealing with an anomaly is strategized as well: localize which module is at fault, then redesign it.

Darden and Maull’s interfield theory is the first systematic description of a cross-field theory as a device for generating novelty. When separate fields treat different aspects of the same phenomenon, as genetics and cytology do, a theory that bridges them produces new explanatory power. The boundary-crossing strategies discussed later have in part been rediscoveries of this formulation within individual fields.

As a form of inference, abduction corresponds. Schurz classified into eleven patterns the inference Peirce set up as the logic of discovery. Factual abduction, abduction that posits a law hypothesis, abduction that posits a theoretical model, and abduction that posits a second-order existential hypothesis are included. Magnani divided it broadly into theoretical and manipulative (forming a hypothesis through manipulation). Being able to say which type of abduction one is using comes close to being able to say what kind of novelty one has.

How Far Does Atypicality of Combination Get Tolerated?

The lineage that measures novelty as a new combination of existing elements is the furthest along in quantitative verification. Uzzi et al. measured, for 17.9 million papers, how atypical the pairs of cited journals were. The result is not monotonic. The papers that become most highly cited are those whose references are conventional for the most part while containing a small number of atypical combinations in the tail. Papers that raise atypicality alone do not take off.

Foster et al. sorted research strategies in chemistry into types and measured their frequency and outcomes across 6.45 million MEDLINE abstracts. Conservative strategies account for 85.8%, while innovative strategies have a low rate of success and a large impact when they succeed. Those who take the gamble skew toward the prize-winning stratum, who would not lose their jobs by losing. Rzhetsky et al. showed with a computational model that this individually optimal choice diverges from the collective optimum. When each person chooses the safe experiment, the collective search narrows.

Care is needed on the side of the indicators as well. Fontana et al. showed that several novelty indicators—atypicality, new combinations, and interdisciplinarity—agree with one another only weakly. Which indicator one measures with swaps out the set of highly novel papers. Funk and Owen-Smith’s CD index (which measures disruptiveness by the degree to which subsequent research stops citing the prior work) is a lineage of its own, and Wu et al. applied it to 65 million papers, patents, and software products, showing that large teams develop existing knowledge while small teams disrupt it. Park et al., with the same indicator, report that the disruptiveness of papers and patents has continued to decline over the long run.

When a Tool Becomes a Theory

Gigerenzer’s tools-to-theories is the most concrete of the generative strategies, and also the most overlooked. Tools researchers use daily, statistical methods or computers for instance, are first accepted as method and then theorized as models of the mind itself. The course by which the theory that humans are intuitive statisticians appeared after significance testing had taken hold is one example. Once this path is in view, the operation of repurposing what one currently uses as a means into a model of the object can be taken deliberately. Gigerenzer and Sturm later examined the conditions under which this transfer falls into a circle, in which the tool makes the theory and that theory in turn guarantees the validity of the tool.

Research on analogy deals with the same kind of transfer, but as a mapping from a distant field. Gentner’s structure-mapping formulated that only a transfer that maps relational structure, not attributes, works scientifically. Gentner and Markman showed that comparison yields not only similarities but alignable differences. The yield of comparing is not that these are alike but that this is where they differ, and that difference becomes the question.

The actual practice, though, is far from the ideal. Dunbar observed molecular biology laboratories in the field and found that about 98% of the analogies used in actual reasoning come from near domains. Distant analogies are not retrieved. Chan and Schunn showed that a distant analogy does not work in a single shot but raises the novelty of an idea when it passes through a chain of transformations. It is not that bringing something from far away is enough; the work of transforming the distant thing by stages is required.

The quantitative consequences of crossing fields are known as well. Larivière and Gingras showed that the relation between interdisciplinarity and citation is an inverted U. Impact is greatest at moderate distance, and at too great a distance the work is not valued. Leahey et al. showed that interdisciplinary researchers gain in prominence while losing in productivity. Boundary crossing is a strategy that carries a cost.

Star and Griesemer’s boundary objects and Galison’s trading zones amount to descriptions of devices that lower this cost. They are arrangements by which different fields collaborate on a local trading vocabulary alone, without building a common language. Oswick et al. divided the importation of concepts from other fields into management studies into three stages—borrowing, transplanting, and blending—and held blending to be the most innovative. Whetten et al., as an axis for inspecting the validity of borrowing, distinguish horizontal borrowing within the same level of analysis from vertical borrowing across levels.

Moving the Edges of a Theory

As a path that produces novelty without denying an existing theory, there is the manipulation of boundary conditions. Of Whetten’s four elements of theoretical contribution (What, How, Why, Who-Where-When), the last, Who-Where-When, points to the range of application. Changing for whom, where, and when a theory holds is a contribution on its own. Busse et al. formulated the procedure for exploring, setting, and testing a boundary condition, and Gonzalez-Mulé and Aguinis showed a method for testing boundary conditions quantitatively with meta-regression.

There is also the strategy of raising a theory’s precision. Edwards and Berry took issue with predictions in management science being so broad that they cannot be falsified, and positioned the narrowing of a prediction as a contribution in itself. Meehl, contrasting with physics, pointed out the paradox that in psychology’s null hypothesis testing an increase in power protects the theory. The greater the power, the easier it is to reach significance, and the harder the theory is to falsify.

Tsang and Kwan are the ones who organized replication as a type of contribution. They divided replication into six types and matched each with what it establishes. Checking an analysis again, reanalysis of data, exact replication, conceptual extension, empirical generalization, and generalization and extension. The last three are replications and yet carry novelty. Köhler and Cortina analyzed the state of constructive replication in management studies at scale, and Bettis et al. argue for its necessity in strategy research.

Strategies on the Breaking Side

The operation of breaking a theory has itself been systematized. Platt’s strong inference is a three-stage procedure: set up multiple alternative hypotheses, design a crucial experiment that excludes them, and repeat. Its precursor is Chamberlin’s method of multiple working hypotheses, which asks that several hypotheses be held at once in order to prevent attachment to a single ruling hypothesis. Gray and Cooper presented the active search for a theory’s failures as a tactic.

Adversarial collaboration, in which opposed disputants jointly design a decisive experiment, is known through the practice of Mellers, Hertwig, and Kahneman. Under the arbitration of a third party, which theory is right is judged by a procedure fixed in advance. Clark and Tetlock argue that this should be placed at the core of scientific reform.

The points where breaking strategies collide with institutions are recorded as well. The Open Science Collaboration showed, in replications of 100 psychology studies, that significant replication stops at 36%. Simmons et al. showed by simulation that researcher degrees of freedom easily push the false positive rate above 60%, and Ioannidis derived the conditions under which prior probability, power, and competition raise the false positive rate. HARKing (hypothesizing after the results are known), which Kerr named, looks like a strategy of generation and is in fact an operation of claiming. Scheel et al. showed that whereas positive results in ordinary psychology papers run at 96%, in Registered Reports they run at 44%. Merely inserting a mechanism that reviews before the results are seen halves the distribution of claiming.

Claiming Collapses into Move 2

Now from the generation side to the claiming side. The most cited work in the genre analysis of academic English is Swales’s CARS model (Create a Research Space). An introduction proceeds in three moves: Move 1 (establishing a territory), Move 2 (establishing a niche), Move 3 (occupying the niche). The claiming of novelty happens in Move 2, and its steps are exhausted by four.

  • counter-claiming: contest a claim made in prior work
  • gap-indicating: point out an area not treated
  • question-raising: raise a question prior work has not answered
  • continuing a tradition: position the work as continuing an existing lineage

Where the strategies of generation split into dozens of kinds, the ways of speaking a claim fold into these four. Problematization and gap-spotting alike appear in an introduction as either counter-claiming or gap-indicating. The variety of strategies does not surface in the text.

The calibration of certainty has been systematized as well. Hyland contrasted hedging (weakening) and boosting (strengthening) as the two poles of the negotiation of knowledge and showed, across 80 articles in 8 fields, that the strength of a claim differs by field. The hard sciences are restrained, the soft sciences explicit. Salager-Meyer divided hedges into five types: shields, approximators, the author’s personal doubt, emotionally charged intensifiers, and compound hedges.

The placement of citations is part of claiming too. Gilbert analyzed referencing as an act of persuasion, and Cozzens proposed the rhetoric-first model, in which a citation is first of all a rhetorical resource. Myers traced the revision of papers and grant applications and described the process by which claims of novelty are revised downward in negotiation with review. A claim is not something an author places unilaterally; it is settled in the exchanges with review.

The Going Rate of Hype Keeps Rising

The vocabulary of claiming has clearly inflated over the past fifty years. Hyland and Jiang defined about 400 hype words and showed in a diachronic corpus that their use has roughly doubled over fifty years. Vinkers et al. report that positive words in PubMed abstracts rose from 2.0% in 1974 to 17.5% in 2014, and that the frequency of words such as novel increased by 2,500% to 15,000%. Millar et al. examined the abstracts of 901,717 successful NIH applications and showed that hype words appear in about 97% of them.

Claiming novelty, in other words, no longer differentiates. In an environment where everyone is claiming, something other than the strength of the claim is used to tell them apart. Trapido showed that researchers with a consistent identity have their novelty recognized more readily. Teplitskiy et al. showed in a randomized experiment that, for the same paper, signals of status govern how it is read and cited. What Merton’s Matthew effect and Bourdieu’s field theory pointed at is this structure of distribution. Bourdieu divided strategies of accumulation in the scientific field into two, succession and subversion, and held that which of them one can choose is governed by the amount of capital one holds. Subversion is not a strategy available to everyone.

The one who described the drawing of boundaries itself as a strategy is Gieryn, with boundary-work, whose purposes divide into three. Expansion of a domain, expulsion of rivals, and protection of autonomy. Collins’s law of the attention space (the positions that succeed in a given period are limited to between three and six) shows that claiming is a competition with a limited number of seats.

What Has Design Research Called a Contribution?

Design research has built its own taxonomy of contribution, separately from every lineage above. The starting point is Frayling’s three types: research into, through, and for art and design. Cross argued that there are ways of knowing peculiar to design, giving grounds for design research to hold criteria different from those of science.

The most cited on the HCI side is Wobbrock and Kientz’s seven types. Empirical, artifact, methodological, theoretical, dataset, survey, opinion. An evaluation criterion corresponds to each type, which makes it possible to avoid mismatches such as demanding empirical evaluation of an artifact contribution. Greenberg and Buxton argued that forcing evaluation at an early stage kills novelty. Their point is still cited as a case of the mismatch between contribution types and evaluation demands. Oulasvirta and Hornbæk redefined HCI research as three kinds of problem—empirical, conceptual, and constructive—and matched each with its own criterion for the quality of a solution.

In the Research through Design (RtD) lineage, what to call a contribution has long been contested. Zimmerman et al. presented four criteria for evaluating RtD (process, invention, relevance, extensibility), but later criticized the weakness of its knowledge claims themselves. Gaver argued that what should be expected of RtD is not replicable theory but intermediate generalization through the annotated portfolio (annotation across several works). Gaver and Bowers’s annotated portfolio is a device for claiming generality without losing particularity. Bardzell et al., conversely, criticized RtD’s knowledge claims as too modest and demanded bold proposals as a form of contribution.

The concept of intermediate-level knowledge functions as the settlement of this dispute. Höök and Löwgren’s strong concept is knowledge positioned between theory and the particular instance, and it must satisfy five requirements: being generative, cutting across several particular instances, bearing on both design elements and use practices, grounding vertically in theory, and grounding horizontally in instances. Dalsgaard and Dindler’s bridging concept is composed of three elements: a theoretical grounding, exemplification in practice, and design implications. Löwgren compares annotated portfolios, strong concepts, design patterns, experiential qualities, and methods as forms of intermediate-level knowledge.

There is also a lineage that defines contribution from the side of the thing made. Odom et al.’s research product carries four qualities that distinguish it from a prototype (situatedness, autonomy, wholeness, subjectification), and its condition of standing is that it withstand long-term operation in real life. Koskinen et al. divided constructive design research into three types, Lab, Field, and Showroom, and Krogh and Koskinen extended this to four traditions by adding Street, then formulated drifting, the deliberate deviation of research from its initial plan, into five types (accumulative, comparative, serial, expansive, probing). The claim is that deviating becomes a strategy rather than a failure. Redström described design theory as an oscillation between program and experiment, giving a procedure for the making side of theory.

Design research’s taxonomy of contribution does not map cleanly onto the generative strategies of other fields. Strong concepts and bridging concepts are evaluated by generative power, not by the truth or falsity of a proposition. Of Bacharach’s criteria for evaluating theory (falsifiability and utility), the former does not apply from the outset. This asymmetry becomes the largest point of friction when design research imports other fields’ accounts of novelty.

Are LLMs Standing In for the Strategies of Generation?

From 2024 on, attempts to hand idea generation itself to machines have come before peer review. Si, Yang, and Hashimoto ran a blind comparison with more than 100 NLP researchers and showed that research ideas generated by an LLM exceeded those of humans in novelty scores and fell below them in feasibility. They also report that the LLM’s own self-evaluation of novelty does not agree with human evaluation. In follow-up work, when 43 experts actually implemented ideas from both sides over more than 100 hours, the ranking on novelty reversed. Novelty measured at the idea stage is not preserved after execution.

Systems on the generation side are moving toward optimizing similarity to the existing literature as a proxy indicator of novelty. Wang et al.’s SciMON sets novelty as an explicit optimization target, and Lu et al.’s AI Scientist automated everything from conception through experiment, writing, and review at about 15 dollars per paper. Zhou et al. present a framework that generates hypotheses from data and updates them iteratively.

The lineage of computational discovery runs from before LLMs. Swanson’s ABC model (A and B are known and B and C are known, but A and C are unconnected) is the origin of literature-based discovery, and Thilakaratne et al. systematically reviewed 409 items in this lineage. Krenn and Zeilinger predicted unconnected concept pairs from a concept network in quantum physics, and in Krenn et al.’s link prediction competition, hand-built features beat end-to-end machine learning. Hope et al. built a method that decomposes product descriptions into purpose and mechanism and retrieves cases with the same purpose but a different mechanism, showing that inspiration retrieval by fine-grained aspects increases the generation of good ideas by 50% to 60%.

Sourati and Evans’s result occupies a singular position within this lineage. They modeled human tendencies of search and then generated alien hypotheses that deliberately avoid them. The strategy is to predict the direction humans want to go and step off it. The machine’s role becomes not thinking in place of humans but pointing at the complement of human bias.

The price has been measured as well. Doshi and Hauser showed, in a randomized experiment with 293 writers and 600 evaluators, that generative AI raises individual creativity while lowering the diversity of content at the collective level. Anderson et al. confirmed experimentally the homogenization by which ideas produced with LLM assistance come to resemble one another. A choice rational as an individual strategy narrows the collective search. The same structure as the externality Rzhetsky et al. showed with a computational model arises from the side of the tools as well.

Which Strategies Have Their Effects Demonstrated?

Make a list of the strategies and the number of strategies that have names diverges sharply from the number whose effects have been verified quantitatively.

StrategySource (representative)Number of typesDemonstration of effect
Problematization of assumptionsAlvesson & Sandberg 20115 layersNone (normative)
Gap-spottingSandberg & Alvesson 20113 subtypesIts dominance is demonstrated. Effect untested
Inversion for interestingnessDavis 197112 typesNone (post hoc content analysis)
Recursive distinctionAbbott 20041 core operationNone
Specification of ignoranceMerton 19871None
Strategic selection of materialsMerton 19871None
Decomposition and localizationBechtel & Richardson 19932Induction from the history of science
Procedures of mechanism discoveryDarden 20063 strategies × 3 stagesInduction from the history of science
Typology of abductionSchurz 200811 patternsNone (formal logic)
Atypical combinationUzzi et al. 20131 indicatorYes (17.9 million items, associated with high citation)
Surprise of contents and contextsShi & Evans 20231 indicatorYes (arises more readily from outsiders)
Indicator of disruptivenessFunk & Owen-Smith 20171 indicatorYes (CD index, verified on 65 million items)
Tools to theoriesGigerenzer 19911Induction from the history of science
Structure-mappingGentner 19831Yes (though distant ones are used only about 2% of the time)
Chained transformation of distant analogyChan & Schunn 20151Yes (sequential analysis)
Boundary crossingLarivière & Gingras 20101Yes (inverted U; too far and it is not valued)
Extension of boundary conditionsBusse et al. 20173 proceduresA method of verification by meta-regression exists
Contribution through replicationTsang & Kwan 19996 typesReplication rate 36% (OSC 2015)
Multiple working hypotheses and the crucial experimentPlatt 19643 stagesNone (normative)
Adversarial collaborationMellers et al. 20011Instances of practice exist. No aggregate of effects
CARS Move 2Swales 19904 stepsFrequency demonstrated in corpora. Effect untested
Hedging and boostingHyland 19982 polesDisciplinary differences demonstrated
Hype wordsHyland & Jiang 2021About 400 wordsThe increase is demonstrated. Effect on acceptance untested
Intermediate-level knowledgeHöök & Löwgren 20125 requirementsNone (normative)
Evaluation criteria for RtDZimmerman et al. 20074 criteriaNone (normative)
Contribution taxonomyWobbrock & Kientz 20167 typesNone (normative)
Idea generation by LLMSi et al. 20241Yes (higher in novelty, lower in feasibility, reversed after implementation)

Almost every row with a verified effect is either a scientometric indicator or a psychological experiment. And the verified results do not line up neatly in the direction of recommending a strategy. Boudreau et al. showed, across 2,130 evaluator-and-proposal pairs, that the greater the intellectual distance from the evaluator, and the higher the novelty of the proposal, the lower the evaluation. Criscuolo et al. showed that selection panels most prefer moderate novelty, extreme novelty being pared away in the course of deliberation. Wang, Veugelers, and Stephan demonstrate that highly novel papers become highly cited over the long run but are not valued in the short term and appear in journals with lower impact factors.

The harder one pushes a strategy of generation, in other words, the lower the pass rate at the stage of claiming. Azoulay et al.’s two quasi-experiments suggest that this conflict can be eased by institutions. Long-term funding tolerant of failure produces more novel output than funding of the short-term-evaluation type, and after the sudden death of a star scientist, new entry by outsiders into that field increases. Part of the strategy for raising novelty lies not in individual technique but on the side of the structure of funding and authority.

Connecting to the Four-Locus Hypothesis

The sister note novel-research-questions-literature examined a practitioner’s hypothesis that divides the loci where novelty resides into four: Problem Definition, Materials, Methods, and Results. Where that is a classification of where to place novelty, this note is a classification of how to make it. The two are orthogonal.

Setting them against each other makes the skew of the collected strategies visible.

  • The Problem Definition locus has the richest set of corresponding strategies. Problematization’s five layers, gap-spotting’s three subtypes, Davis’s twelve types, Abbott’s recursive distinction, and Merton’s specified ignorance all fall here.
  • The Materials locus has few corresponding strategies. Merton’s strategic research materials, Rheinberger’s epistemic things, and Hacking’s argument about intervention are nearly all of it. The literature that makes the selection of materials explicit as a strategy is thin.
  • The Methods locus is matched by Gigerenzer’s tools-to-theories and Darden’s procedures of mechanism discovery. The former, though, is a path that repurposes a method into a theory, which points in a different direction from novelty of the method itself.
  • The Results locus has no dedicated corresponding strategy within the range of this collection. A design that pushes by scale connects to Simonton’s equal-odds rule (output volume determines outcome), but this is a description of probability rather than a strategy.

This skew is consistent with the judgment the four-locus hypothesis set up as practitioner knowledge, that competing on Results is to be avoided because the cost is enormous. That strategies are scarce at the Results locus can be read not as strategies being unnecessary there, but as novelty there being determined by resources rather than by technique.

Gaps in the Collection

  • Discussion outside the Anglophone world: scientific methodology in Japanese, Chinese, and German, and contribution taxonomies in non-Anglophone design research. Only Lillis and Curry, treating the inequity of English-centrism, is included; no primary literature from the parties themselves is held.
  • Novelty in the humanities and in arts research: the norms for claiming that an interpretation is new. Accounts of contribution in art history and literary studies are not in this corpus.
  • Mathematics and theoretical computer science: a typology of novelty in proof (a new proof technique, an alternative proof of a known theorem, generalization). This corpus is skewed toward the empirical and design sciences.
  • Evidence from the reviewer’s side: studies that analyze how claims of novelty are read in review, on a corpus of review reports themselves. Research on claiming is skewed toward the rhetoric of the submitting side.
  • Non-obviousness in patent examination: the similarities and differences between the industrial determination of non-obviousness and the academic concept of novelty.
  • Direct comparison of the effects of strategies: a design that sets a body of research that adopted a given strategy against one that did not. The column for demonstration of effect in the table is mostly empty because this design barely exists.

Unverified Items

Listed here are the claims used as grounds in the body whose bibliographic details or numerical conditions could not be confirmed at first hand.

  • The figure of about 98% for near analogies in Dunbar 1995: the pages of the containing volume (The Nature of Insight, MIT Press) have not been confirmed.
  • The magnitude of the improvement in discovery efficiency in Sourati and Evans 2023 changes with the conditions of application, so no multiplier is stated in the body.
  • The DOI of Krenn et al. 2023 (Nature Machine Intelligence) has not been confirmed.
  • The arXiv number of the follow-up by Si et al. reporting the reversal after execution (The Ideation-Execution Gap) has not been confirmed.
  • The ICLR 2025 acceptance of Si et al. 2024, and the conditions of application for the effect sizes in Zhou et al. 2024 (synthetic data +31.7%, real data +3.3% to +24.9%), have not been confirmed.
  • The DOIs for Busse et al. 2017, Gonzalez-Mulé and Aguinis 2018, Tsang and Kwan 1999, Bettis et al. 2016, Edwards and Berry 2010, Gray and Cooper 2010, Meehl 1990, and Chambers 2013 were inferred from the bibliographic records and have not been checked against the primary sources.
  • For Frayling 1993, the primary URL of RCA Research Papers 1(1) has not been confirmed.
  • Medawar 1963 is an essay published in The Listener, and no DOI exists.
  • Some of the ISBNs of the books (Bechtel & Richardson, Boden, Holyoak & Thagard, Hesse, Galison, Koestler, Simonton, Klein, Hofstadter & Sander, and others) are unconfirmed. A full list is in the ## 未検証事項 section of the corpus.
  • Recorded in the corpus but not identified bibliographically and not used as grounds in the body: Nova, IdeaBench, NovBench, LLMs Hallucinate Alike, On the Limits of LLM-as-Judge, and Bailer-Jones / Knuuttila on models.

References

Problematization and the Construction of Contribution (Management and Organization Studies)

Heuristics of Discovery and the Philosophy of Science

  • Bechtel, W., & Richardson, R. C. (1993). Discovering Complexity. MIT Press. ISBN 978-0-262-52528-8
  • Campbell, D. T. (1960). Blind variation and selective retention in creative thought. Psychological Review, 67(6). https://doi.org/10.1037/h0040373
  • Darden, L. (2006). Reasoning in Biological Discoveries. Cambridge University Press. https://doi.org/10.1017/CBO9780511760730
  • Darden, L., & Maull, N. (1977). Interfield theories. Philosophy of Science, 44(1). https://doi.org/10.1086/288761
  • Hacking, I. (1983). Representing and Intervening. Cambridge University Press. https://doi.org/10.1017/CBO9780511814563
  • Hanson, N. R. (1958). Patterns of Discovery. Cambridge University Press. ISBN 978-0-521-09261-6
  • Kuhn, T. S. (1977). The Essential Tension. University of Chicago Press. ISBN 978-0-226-45806-0
  • Lakatos, I. (1970). Falsification and the methodology of scientific research programmes. In Criticism and the Growth of Knowledge. https://doi.org/10.1017/CBO9781139171472.009
  • Magnani, L. (2009). Abductive Cognition. Springer. https://doi.org/10.1007/978-3-642-03631-6
  • Merton, R. K. (1987). Three fragments from a sociologist’s notebooks. Annual Review of Sociology, 13. https://doi.org/10.1146/annurev.so.13.080187.000245
  • Popper, K. R. (1963). Conjectures and Refutations. Routledge. ISBN 978-0-415-28594-0
  • Rheinberger, H.-J. (1997). Toward a History of Epistemic Things. Stanford University Press. ISBN 978-0-8047-2785-6
  • Schickore, J. (2018). Scientific discovery. Stanford Encyclopedia of Philosophy. https://plato.stanford.edu/entries/scientific-discovery/
  • Schurz, G. (2008). Patterns of abduction. Synthese, 164(2). https://doi.org/10.1007/s11229-007-9223-4
  • Simonton, D. K. (1988). Scientific Genius: A Psychology of Science. Cambridge University Press. ISBN 978-0-521-35892-0
  • Wimsatt, W. C. (2007). Re-Engineering Philosophy for Limited Beings. Harvard University Press. ISBN 978-0-674-01545-3

Measuring the Effects of Strategies (Scientometrics and the Sociology of Science)

The Rhetoric and Genre of Claiming

  • Cozzens, S. E. (1989). What do citations count? The rhetoric-first model. Scientometrics, 15. https://doi.org/10.1007/BF02020002
  • Gilbert, G. N. (1977). Referencing as persuasion. Social Studies of Science, 7(1). https://doi.org/10.1177/030631277700700102
  • Gilbert, G. N., & Mulkay, M. (1984). Opening Pandora’s Box. Cambridge University Press. ISBN 978-0-521-27430-5
  • Hyland, K. (1998). Boosting, hedging and the negotiation of academic knowledge. TEXT, 18(3). https://doi.org/10.1515/text.1.1998.18.3.349
  • Hyland, K. (2000). Disciplinary Discourses. Longman. ISBN 978-0-582-41637-7
  • Hyland, K., & Jiang, F. (2021). “Our striking results demonstrate…”: Persuasion and the growth of academic hype. Journal of Pragmatics, 182. https://doi.org/10.1016/j.pragma.2021.01.025
  • Knorr-Cetina, K. (1981). The Manufacture of Knowledge. Pergamon Press. ISBN 978-0-08-025753-8
  • Latour, B., & Woolgar, S. (1986). Laboratory Life (2nd ed.). Princeton University Press. ISBN 978-0-691-02832-3
  • Medawar, P. B. (1963). Is the scientific paper a fraud? The Listener, 70. (no DOI)
  • Millar, N., Mohammadi, M., & Budgell, B. (2022). Trends in the use of promotional language (hype) in abstracts of successful NIH grant applications, 1985-2020. JAMA Network Open, 5(8). https://doi.org/10.1001/jamanetworkopen.2022.28676
  • Millar, N., Salager-Meyer, F., & Budgell, B. (2019). “It is important to reinforce the importance of…”. English for Specific Purposes, 54. https://doi.org/10.1016/j.esp.2018.11.001
  • Myers, G. (1990). Writing Biology. University of Wisconsin Press. ISBN 978-0-299-12484-2
  • Salager-Meyer, F. (1994). Hedges and textual communicative function in medical English written discourse. English for Specific Purposes, 13(2). https://doi.org/10.1016/0889-4906(94)90013-2
  • Swales, J. M. (1990). Genre Analysis: English in Academic and Research Settings. Cambridge University Press. ISBN 978-0-521-33813-7
  • Swales, J. M. (2004). Research Genres: Explorations and Applications. Cambridge University Press. ISBN 978-0-521-53833-0
  • Vinkers, C. H., Tijdink, J. K., & Otte, W. M. (2015). Use of positive and negative words in scientific PubMed abstracts between 1974 and 2014. BMJ, 351. https://doi.org/10.1136/bmj.h6467

Extension, Replication, and Falsification of Theory

Transfer, Analogy, and Boundary Crossing

Contribution Taxonomies in Design Research and HCI

LLMs and Algorithmic Discovery


← All Notes · Home