Shuichiro Ogawa
日本語

Notes · updated 2026-08-09

Writing Material for a Review Article: The Assumption Ledger and Nearest-Neighbour Literature Left by 17 Novelty Audits

What This Note Records

The work of selecting a research topic proceeded in four stages. First, 117 notes accumulated in the wiki were read across, and 20 candidates at a granularity that could stand as peer-reviewed papers were extracted (record of the candidate extraction). Second, these were narrowed to 5 that could be carried out as qualitative studies, and each candidate was subjected to a novelty audit that identified the nearest prior research through external literature searching. Third, because the audits showed that “claims of the type that point to a gap in the literature cannot be defended in peer review,” the mode of question generation was switched to assumption reversal (problematization, the procedure of identifying and challenging the assumptions a field shares), and 12 challenges were generated from 45 assumptions across four fields (record of the assumption ledger; the academic lineage of this procedure is set out in Where Novel Research Questions Come From and How Novelty Is Made and Claimed). Fourth, those 12 were subjected to the same kind of audit.

This note reworks the material left by these 17 audits into raw material for writing a review article.

The One Fact the Audits Established

Across the 17 audits, the number of candidates whose novelty remained intact was zero.

Verdict5 qualitative studies12 assumption challengesTotal
Intact000
Survives with qualifications3912
Substantively already published235

The reasons behind the verdicts matter more than their distribution. Among the empirical candidates, the constituent elements had already been filled in, each in a different field. The jurisdictional analysis of emerging design occupations, for instance, already had discourse studies of professionalization, job analysis based on job postings, and applications of the sociology of professions to AI, each existing separately, with only the combination left open. The same happened with the assumption challenges. The exposure itself, of the form “the field takes X for granted,” had already been put in print by someone else in nearly all 12 cases. The reversal that treats judgement as an allocation of responsibility rather than a capability had been claimed by responsibility research; the criticism of removal tests by the assessment-reform school; the reactivity of evaluators by the sociology of quantification and by benchmark criticism.

Novelty that can withstand peer review therefore remains not on the side of the claim but only on the side of empirical execution in a specific field and unit of analysis. And empirical execution demands field access, participants, and longitudinal time, resources that are not currently at hand.

Why the Centre of Gravity Moves to a Review Article

The novelty of a review article is placed somewhere other than that of an empirical paper. Whereas an empirical paper demands “a fact no one has reported,” a review demands coverage, an integrative framework, and the identification of disconnections between fields. The by-products left by the 17 audits correspond to these precisely.

The audits actually identified the nearest literature for each of the four fields. There are 88 items in total, each confirmed down to a DOI or URL. This is exactly what the related-work section of a review article demands. The 45-item assumption ledger can also serve as the analytical framework that keeps the review from ending as a descriptive bibliography. And the map of “which claims are already occupied and where the openings are,” confirmed for each field, provides the grounds for assembling the review’s contribution claim.

In other words, the material that was insufficient for an empirical paper is very nearly sufficient for a review article.

Three Writable Proposals

ProposalTypeMain corpusMaturityFirst target venue
AScoping review (PRISMA-ScR)53 academic + 39 industry + 35–50 additionalOutline finalizedIJDCI
BCritical literature review (identification of assumptions)45 assumptions across four fields + 88 nearest-neighbour itemsNewly established in this sessionDesign Studies / She Ji
CIntegrative review (adjudicating the conditions of harm)About 70 items (triple corpus)Outline existsEducational Psychology Review

Proposal A: LLM-as-a-Judge for Creativity Research

This is the most advanced proposal: its research questions, chapter structure, target venue, and mapping onto the 22 PRISMA-ScR items are already settled (record of the outline; the main corpus comes from the literature map of LLM-as-a-Judge and its genesis). There are four research questions (what can be evaluated at what level of human agreement, what types of validity evidence are reported, where it connects with and departs from the lineage of design critique, and what research questions remain open). The contribution lies in mapping the empirical work scattered across NLP and HCI proceedings using the vocabulary of evaluation theory in design research (consensual assessment, studio crit, novelty metrics in engineering design), and in placing a matrix of artifact types against validity-evidence types at the centre of the Results.

The audits in this session added two pieces of material to this proposal. First, as a point that can be placed in the Discussion, there is the possibility that the evaluator operates as a selection environment for generation (assumption 3-2, discussed below). The judge is a reward model from reinforcement learning taken out of the training loop, and it already exerts selection pressure on the generation side. If, the more the evaluator’s accuracy is improved, the distribution of the very creativity being measured shifts, then validity can degrade independently of calibration accuracy. For this point, a set of nearest-neighbour literature is available: reactivity research in the sociology of quantification, the argument on algorithmic monoculture, and empirical work and reviews on homogenization by large language models. Second, as an unexplored question that can be placed in the Research agenda section, there is the position that treats inter-rater disagreement as signal rather than error (assumption 3-1). This position, however, has already been institutionalized in natural language processing as perspectivism, and there are two negative results in the field of research evaluation showing that the amount of disagreement does not predict novelty. To handle it in a review, the form would be to state these negative results explicitly and then present the untested residue, the structure of disagreement, as a research question.

Proposal B: A Critical Review of the Assumptions in Research on AI and the Creative Professions

This is the proposal newly established in this session. There are precedents for critical literature reviews of this kind: a critical review of AI discourse in higher education appeared in 2023. Reviews whose subject is “what a field’s literature takes for granted” therefore exist as an established genre.

The research questions of this proposal are: what assumptions are shared by the body of research on AI and the creative professions (design practice, learning, creativity evaluation); by what evidence are those assumptions supported; and where do they generate internal contradictions. Three contributions can be placed here. First, arranging three bodies of literature that are usually read separately along a single axis, the assumptions they share. Second, naming the unconfronted contradictions inside each field. The clearest is the example from learning research: harm research defines true learning as “the performance that remains once AI is removed,” whereas competency-measurement research adopts the exactly opposite assumption that “the era of measuring the individual without AI is over,” and the two coexist in the same field without ever having confronted each other over the choice of reference environment. Third, presenting, for each assumption, a pairing of an alternative assumptive ground and an observable quantity that could test it.

The weakness of this proposal is equally clear. Because the exposure of each individual assumption has precedents, the review’s contribution has to be placed not in “each individual exposure” but in “the arrangement that cuts across three fields, and the identification of contradictions.” The related-work section will need to cite responsibility-allocation research, the assessment-reform school, the sociology of quantification, and the criteriology debate directly, and to state explicitly what this review has newly placed side by side.

Proposal C: What Kinds of Offloading Break Learning

This is an integrative review confined to learning research, and the material, a triple corpus (the original harm studies, peer-reviewed refutations and replications, and the debate over failure-based design; roughly 70 items in total), is in place (it comes from Cognitive Offloading and Learning, Patterns of Technological Harm in Education, and Failure-Driven Learning). The claim converges on a single axis of adjudication. Harm that survives robustly after replication testing is concentrated where the cognitive processing being delegated is the learning objective itself. The structure would be to classify fifty years of harm discourse, from calculators, search, laptops, and screen time to generative AI, by mechanism, and to sort it along this axis.

The audits in this session changed one thing about how this proposal is positioned. The assessment-reform school has been repeatedly asserting, at the level of norms, the question of whether it is right to measure under AI-removed conditions, and this has entered a regulator’s discussion paper as well. Proposal C therefore connects better to the existing debate if it is framed not merely as adjudicating the presence or absence of harm, but as a problem of measurement theory in which “the meaning of harm changes according to which reference environment is chosen.”

The Analytical Framework: An Assumption Ledger for Four Fields

This is the core of Proposal B, and can also be used as the frame for the Discussion in Proposals A and C. The assumptions were excavated in five layers (in-theory assumptions, the field’s shared image of what counts as evidence, paradigm, ideology, and field assumptions). The 12 that passed the assessment of being worth challenging are presented here together with the audit verdicts.

Field 1: AI, Design Practice, and the Professions

  • Judgement is a capability internal to the individual: the residue not replaceable by AI is taken to be judgement and taste, and these are held to be definable, teachable, and measurable. This permeates all 9 notes. The alternative ground is the view that judgement is the warranting authority and answerability that an organization confers on a particular position, supported by the sociology of professions and workplace studies. The audit found that the reversal in which the human remains as the addressee of responsibility already exists in responsibility research. What remains is a description of the interaction through which the entitlement to be recognized as “having judged” is allocated in review meetings.
  • AI is an exogenous technological shock: the schema in which organizations introduce it and individuals accept or resist it. The alternative ground is the view that the design profession is itself the supplier of AI’s conditions of execution (design tokens, specification writing, evaluation criteria). The audit found that the theory of knowledge codification and the labour discourse on self-automation are prior work, but no study was found that observed the machine-readability of design systems as an endogenous laying of infrastructure.
  • The editor role is an end state: the one-directional transition from maker to editor is taken to be a durable form of work (the structural analysis of this transition is From Maker to Editor). The alternative ground is the view that monitoring and editing are temporary scaffolding filling in the model’s deficiencies, but this has already been formulated as the paradox of automation’s last mile, and it was judged substantively already published.

Field 2: Generative AI and Learning

  • True learning is the performance that remains once AI is removed: this uniformly specifies the dependent variable of the harm-research literature. The alternative ground is the view that judging transfer requires specifying a reference environment, and that the post-graduation environment is AI-saturated. The audit found that the assessment-reform school has already made this claim at the level of norms. What remains is the empirical question of what AI-removed performance predicts about subsequent performance in AI-saturated environments.
  • Morally loaded vocabulary is neutral description: labels such as laziness, debt, and dependence are used as analytical concepts. The alternative ground is the view that these are verdicts that diffuse into policy ahead of any settling of the evidence. The audit found that discourse analysis, metaphor analysis, the application of the technology-panic cycle argument, and criticism of the concept of dependence are all already occupied, and it was judged substantively already published.
  • Exposure to AI is assigned: it is taken to be assignable as a condition and to be openable and closable as a curricular stage. The alternative ground is the view that AI access is environmental infrastructure, and that learners compose their own ecology of use across sanctioned and unsanctioned channels. The audit found that a theoretical frame in which learners build their own learning ecology through unsanctioned practices when the sanctioned route fails exists in research on surgical training, and that its transfer to generative-AI education is open.

Field 3: The Evaluation and Measurement of Creativity

  • The higher the agreement the better: maximizing inter-rater agreement and agreement with human ratings is taken to be an improvement in evaluation. The alternative ground is the view that structured disagreement among raters who share existing norms is a signal that the artifact is touching the boundary of those norms. The audit found that frameworks treating disagreement as information are already institutionalized in natural language processing, and that there are two negative results showing that the amount of disagreement does not predict novelty.
  • Evaluation is a downstream process that does not change generation: the evaluator is taken to be a measuring instrument that can be calibrated independently of its object. The alternative ground is the view that the evaluation apparatus operates as a selection environment for generation and changes the very distribution it seeks to measure. The audit found that the application of reactivity, the monoculture argument, and empirical work and reviews on homogenization are all already published. What remains is the connection that bundles these into a validity theory for creativity evaluation.
  • Sensitivity to provenance is a bias: evaluation is taken to be properly independent of the artifact’s provenance. The alternative ground is the view that provenance is a constituent of aesthetic judgement, and that evaluation is the evaluation of achievement. The audit found that this philosophical thesis has already been applied to AI, not only in the canon but in the latest issue of an aesthetics journal. What remains is decomposing expert judgement functions empirically and translating them into construct validity for evaluator design.

Field 4: Design Process Research and Methodology

  • Design knowledge is a distinct third culture: this grounds the field’s methodological autonomy. The alternative ground is the observation that this claim of distinctiveness rests on an asymmetric comparison, in which the design side was described through observation of practice while the contrast term, science, was represented by a textbook self-description. The audit found that this problematization has already been established as a controversy in the pages of design research journals. What remains is the empirical adjudication of that controversy.
  • The project is the unit of observation: a container with a start and an end point is taken to be the natural unit of process. The alternative ground is the view that project boundaries derive from contracting and billing conventions and are accounting containers, and that some of the findings reported as properties of process may be properties of the container. The audit found that a methodological precedent, the biography of artifacts, exists in information systems research, but that criticism targeting the validity apparatus of design cognition research was not found.
  • Methodological criteria are not applied to oneself: a field that applies evaluation theory to its objects does not turn it on its own peer review and quality criteria. The alternative ground is the view that qualitative methodological criteria are themselves an evaluation infrastructure with reactivity. The audit found that performativity research on reporting checklists and theories of constitutive effects that do not presuppose commensuration are prior work, and it was judged substantively already published.

Where to Place the Novelty of the Review

Translated into the design of a review article, the lessons of the audits come to four.

First, do not make claims of absence the pillar of the contribution. “There is no research addressing X” is merely a declaration of search scope, and collapses at a single counterexample from a reviewer. Instead, cite the nearest prior work by name and write the difference from it.

Second, make coverage itself the contribution. In a review, how many items were picked up, by what criteria, and how far, is the claim. This is why Proposal A treats compliance with PRISMA-ScR as a requirement.

Third, treat disconnections between fields as a principal finding. The audits in this session found several concrete instances of disconnection: that the assessment-reform school and harm research do not cite each other; that evaluation frameworks from computational creativity are not connected to post-LLM empirical work; and that design cognition research and infrastructure research address the same problem in different vocabularies. Disconnection is a finding that only a review can write.

Fourth, identifying an internal contradiction is stronger than a claim of absence. Absence collapses at a counterexample, whereas for a contradiction the field’s own literature is the evidence.

Methodological Conventions Common to All Proposals

The procedures already fixed in the outline of Proposal A can be carried over to Proposals B and C.

  • Preregistration: register the protocol (RQs, inclusion and exclusion criteria, search strings, charting form) with OSF or similar. In Proposal A, this is the highest-priority unmet item among the 22 PRISMA-ScR items.
  • Two-layer corpus: restrict the subject of the flow diagram to empirical studies, and treat genealogical literature and industry material as a separate layer of evidence. A merged flow makes the denominator and the reasons for exclusion ambiguous and invites methodological criticism.
  • Reporting the screening arrangement: humans make the final judgement on all items, and language models are restricted to disagreement detection and sensitivity analysis. Disclose the model version, dates, prompts, thresholds, and the handling of disagreement, and state single screening explicitly as a limitation. Because Proposal A’s subject is the evaluative capability of language models, it does not call a language model a second reviewer.
  • Place the synthesis in the Results: so as not to end as a descriptive inventory, place a cross-tabulation of axes (artifact type against validity-evidence type in Proposal A; field against assumption layer in Proposal B) at the centre of the Results, and treat empty cells as a source of research questions.

Writing Order and Remaining Work

Proposal A goes out first. All design decisions are settled, and the remaining work has been reduced to execution alone: preregistration, an additional 35 to 50 items of collection, charting, and writing the text. No competing review has been confirmed on the design research journal side, and given the growth rate of this area the window for going first is not long.

Proposal B comes next. The assumption ledger and the nearest-neighbour literature are in place, but the coverage required of a review (the work of systematically re-collecting the literature supporting each assumption) remains. The audits in this session only identified nearest neighbours; they were not exhaustive collection.

Proposal C is to be started after either Proposal A or Proposal B has been submitted. The material is in place, but the work of rewriting its position relative to the assessment-reform school within the frame of measurement theory will be required.

Unverified Items

The items requiring primary verification before submission are carried over from the audits in this session.

  • The scope of the homogenization review (Trends in Cognitive Sciences 2026) and of the reactivity application paper (DIS 2024). If either explicitly addresses evaluators for creativity evaluation, the positioning of Proposal A’s Discussion changes.
  • The full text of the assessment-rethinking paper (Assessment & Evaluation in Higher Education 2026). It bears on the positioning of Proposal C.
  • The full text of the workplace jurisdiction paper (Academy of Management Journal). The full text has not been reached.
  • Forward citations of the design-and-science controversy (Farrell and Hooker, 2013), the methodological review of the biography of artifacts (2019), and the performativity study of reporting checklists (2025). These are needed to confirm coverage of the three assumptions handled in Proposal B.
  • Settling the bibliographic details of the division-of-labour reconfiguration study (CSCW 2025).
  • The publication record of the target venues. Precedents for scoping reviews and word limits at IJDCI, Design Science, and Design Studies.
  • Bibliographic details (author, year, journal, DOI) for roughly 30 additional collection candidates.

The audits relied on general web searching in English, and systematic searches of Scopus, Web of Science, and the ACM Digital Library, non-English literature, and paywalled full texts remain unchecked. A further audit will be required when work begins.

References

The nearest-neighbour literature identified in this session’s novelty audits is collected here by field. Items whose full text has not been confirmed as reached are annotated.

Methodology of Question Generation

AI, Design Practice, the Professions, and the Allocation of Responsibility

AI Adoption, Attitudes, and Creative Labour

Knowledge Codification, Infrastructure, and the Last Mile of Automation

Learning, Assessment, and Cognitive Offloading

Design Education and Generative AI

The Evaluation of Creativity, Disagreement, and the Reactivity of Evaluators

Provenance, Achievement, and Aesthetic Judgement

Design Process Research and Methodology

Criteria for Qualitative Research and Their Reactivity


← All Notes · Home