Shuichiro Ogawa
日本語

Notes · updated 2026-07-20

Returning to the Same Object a Third Time

Returning to the same object a third time is warranted not because the object has changed but because the instrument has.

ai-research-gaps-abduction inverted AI research’s taken-for-granted premises one by one and searched for gaps through seven frames. ai-research-gaps-revisited lowered that inversion into the A-through-G worksheet and sorted the seven into symmetry-completion frames and the rest. Both instruments took as their starting point the question of which premise to invert.

This time the instrument starts somewhere else. What it asks is what has been counted as a “gap” in the first place, and what board decides that. The typology of Problem Definition organized in problem-definition-metagame takes as its unit of operation not the inversion of a premise but the rewriting of the certification system a field tacitly holds fixed (what counts as a phenomenon, what serves as a unit, what stands as evidence of success, and where the endpoint lies). Research on how research questions are constructed has distinguished gap-spotting, which fills blanks in the literature, from problematization, which challenges the premises themselves12. The two sister notes performed the latter, but what certification criteria the performance itself stood on had not yet been asked. Rereading the same object with this generator brings into view three things the earlier two did not show. The identity of the seven frames, a retrial of the two demoted candidates, and the division of labor among the generators themselves.

Writing Out the Game Board of AI Research (Step 1)

The generation protocol’s first job is to transcribe the board without questioning it. Within the range that sources can corroborate, the board of AI research can be written as follows.

  • Purpose and the certification of success: the most frequent values are performance, generalization, quantitative evidence, efficiency, and novelty. In a value analysis of 100 highly cited ML papers, 15% justified a connection to a societal need and 1% discussed negative potential impacts3.
  • What is measured: a few “general” benchmarks are venerated as proxies for general-purpose capability, yet that framing lacks construct validity4. The tendency to look only at what can be measured has long been warned against from inside the field5.
  • The unit of comparison: one model, one user, one task, one point in time. The datasets in use themselves concentrate on those originating from a few elite institutions6.
  • The endpoint: the moment of testing. The passage of time after deployment lies outside evaluation.
  • The position of observation: a researcher standing outside the system. Even as LLMs take on literature review and refereeing and the instrument enters the inside of the research process, this assumption stays in place7.
  • Noise: negative results and failure. Their receptacle stays in sporadic workshops rather than a main-conference track8, and the systematization of failure modes like data leakage belongs to the exceptions910.
  • The record-keeping institution: papers and leaderboards are the field’s memory, and exploration that does not appear in them does not remain.

The evidence for this board is the same the sister notes used. The difference lies in the arrangement. The earlier two read these as “premises to be inverted”; here they are read as “constants of the board.” Arranged as constants, the next move (what lies off the board) can be played mechanically.

What Was Pushed Off the Board, and the Most Unnatural Asymmetries (Steps 2 and 3)

Each entry on the board has an off-board counterpart. Failed experiments, designs not adopted, negative results, post-deployment degradation and adaptation, the provenance of the evaluation apparatus itself, the reflux of model products into data, capabilities that no benchmark carries. None of these were discarded because they are unobservable. They go unrecorded because the board’s certification criteria classified them as “things not counted.”

From this off-board list, select the three most unnatural asymmetries.

First, a model’s success is recorded, but failed exploration is not. That the publishing culture’s valuation skews toward performance and novelty3, and that the receptacle for negative results stays at the margin8, are the institutional supports of this asymmetry.

Second, models are evaluated, but the institution that certifies an evaluation apparatus as a success-certifying apparatus is not. Critiques of benchmark validity exist4. But research that takes as its object of explanation the provenance of authority itself, when and how a measure acquires the authority to certify success and when it loses it, is a different job from critique.

Third, training data are measured, but the pathway by which model outputs reflux into data does not enter the unit of evaluation. The consequence of the reflux (model collapse) has been demonstrated11, yet the reflux pathway itself remains off the board.

Which Operation on the Board Each of the Seven Frames Was

With the board’s vocabulary in hand, the sister notes’ seven frames can be restated at a different resolution. Which entry on the board was each frame an operation to rewrite?

FrameOperation on the board
1 Time horizonChanging the time span and the endpoint
2 Reflexivity of researchChanging the position of observation (putting the researcher on the board)
3 AI as an ecologyChanging the unit of comparison (from individual to system)
4 Asymmetry of non-knowledgeRecertifying noise (failure into a diagnostic instrument)
5 Cognitive diversityChanging the legitimate position of observation and the certification of success (whose viewpoint)
6 Monism of epistemologyQuestioning the cultural provenance of the success-certifying apparatus
7 Self-mapping of blind spotsMaking the record-keeping institution itself the object

All seven fall onto some row of the meta-conversion table. The two generators, then, are not separate things. Symmetry inversion is a rewriting specialized to one entry of the board (a pair placed asymmetrically), and Problem Definition is the superordinate family of operations above it. That AI’s own problem-setting as a field is bound by premises embedded in its technical practice was pointed out early, and from inside the field12. The seven frames can be read as executions of that observation as individual operations on the board.

Being superordinate, however, does not mean being stronger. Inversion came with vetting criteria, D and G (the structural correspondence across categories, and the minimal observable). The rewriting of the board as such has no vetting criteria. This difference does its work in the retrial that follows.

Retrying the Two Demoted Candidates in the Board’s Vocabulary

The revisit note demoted cognitive diversity and non-Western epistemologies (frames 5 and 6) as “an extension of coverage rather than a displacement of property,” and the self-mapping of blind spots (frame 7) because “the observable for G will not be written.” As a determination of kind, the judgment still holds. An extension of coverage is not a completion of symmetry, and a candidate whose observable cannot be written does not become a falsifiable hypothesis.

Restated in the board’s vocabulary, though, the meaning of the demotion comes one step clearer. The A-through-G worksheet had installed falsifiability and the minimal observable as the success-certifying apparatus for “blind spots.” The worksheet itself, that is, was a game board, and it certified as “countable blind spots” only what could write a G. The two candidates that fell did so not because they are weak as questions but because they did not fit that worksheet’s certification criteria. Frames 5 and 6 are legitimately strong as a board rewriting of “whose viewpoint counts as a legitimate position of observation.” Frame 7 is legitimately strong as a rewriting that “makes the record-keeping institution itself the object.” The retrial’s conclusion is to append this one line to the revisit note’s judgment. Not a symmetry-completion blind spot: that stands. On top of it is added: a strong question as a different type of Problem Definition.

This retrial does not overturn the revisit note’s selection. It only made explicit, from outside the apparatus, what the selecting apparatus counts and what it does not. Standing outside the apparatus is possible because this generator carries the type that “questions the success-certifying apparatus” (type 2), and this section is the result of applying that type to the A-through-G worksheet itself.

Applying the Conversion Table Comes Out at the Same Hole

Against the suspicion that a different generator should open a different hole, there is one piece of evidence pointing the other way.

Take three rows of the meta-conversion table, “snapshot evaluation → temporal chain,” “data → the record-keeping institution that generated the data,” and “in-game behavior → the conditions that constitute the game,” and apply them to research on scaling laws. Scaling laws have taken as their unit a one-time performance measurement at each scale, treated training data as given, and left off the board where the data come from. Every one of the three conversions says the same thing: put on the board the pathway by which a model’s outputs beget the next data. This is the same place as the endogenous scaling law the revisit note dug out through symmetry inversion (invert the exogeneity of the input distribution and write the closed loop with the generated-product fraction as an endogenous state variable). Scaling-law research that treats the synthetic-data fraction as an exogenous mixing parameter13, the demonstration of the reflux’s consequence11, and its counterexample14 are all prior work, and that the remaining void lies at the single point of endogenizing the fraction also agrees with the revisit note’s limitation.

Two independent generators converged on the same void. This is weak but non-negligible evidence that the void is not an artifact made by a generator’s quirks. Weak, because the same author ran both generators, and a shared vantage cannot be ruled out.

Generating Two New Problem Statements (Steps 4 and 5)

A board asymmetry is not left at confirmation; it is converted into a problem statement. The form is fixed. Existing research has presupposed X and explained Y. But under that presupposition, Z cannot be identified. The two below are candidates generated mechanically from the board’s asymmetries in this form, and this note has not checked their coverage by prior work. They are placed not as confirmed gaps but as gap candidates to be verified in the next collection.

The first problem statement comes out of the second asymmetry (the evaluation apparatus is not evaluated).

Existing research has presupposed benchmark scores as evidence of capability and explained the ranking of models. But under that presupposition, when a measure acquires the authority to certify success, and when it loses it, cannot be identified.

RQ: from which records (leaderboard adoption and retirement, paper citations, corporate adoption criteria) can the process by which a benchmark acquires, renews, and abandons its authority as a success-certifying apparatus be reconstructed? How much meta-science research takes the provenance of benchmarks as its object is unverified. [primary verification needed]

The second problem statement comes out of the first asymmetry (failed exploration is not recorded).

Existing research has presupposed published successes (papers, models, leaderboard entries) as the field’s output and explained progress. But under that presupposition, where rejected exploration was discarded, and what the field failed to learn, cannot be identified.

RQ: which exploration does the field’s record-keeping institution make visible, and which does it make invisible? On what convention for discarding things as noise does that selection depend? The extent of overlap with research on publication bias and negative results (frame 4’s body of evidence) is unverified. [primary verification needed]

The two candidates may look like restatements of frames 4 and 7. But the difference lies in the unit. Frame 4 pointed to the absence of a systematization of non-knowledge; frame 7, to the absence of a method of self-mapping. The two here install as the object of explanation neither models nor methods but the institutions of certification and record-keeping themselves. The conversion table’s rows “data → the record-keeping institution that generated the data” and “subject → the apparatus that forms the subject” drive this move of the unit.

The Division of Labor Between Generators

The three notes’ instruments differ in role.

The Problem Definition typology generates questions. Generation is fast; once the board can be written, candidates come out mechanically. But it carries no criterion that guarantees the strength of the candidates it produces. The A-through-G worksheet vets questions. Vetting is strict, and a candidate that passes becomes a falsifiable hypothesis. But it cannot count questions of the types that cannot write a G (coverage, fairness, institutional provenance).

Hence the division of labor. Make questions by rewriting the board, and take the candidates that fall into the symmetry-completion type down to falsifiability on the worksheet. For candidates of types the worksheet does not carry, erect separate vetting criteria proper to the type (reachability of archival sources, reconstructability from institutional data, and the like). Generation without vetting drifts into mass-producing candidates; vetting without generation spins idle.

The question left open sits on the side of the vetting apparatus. The symmetry-completion type had F and G. What apparatus gives falsifying conditions of the same strength to questions of the type that asks after institutional provenance? Until that stands, the questions this third generator makes will pile up as candidates that look strong and cannot be vetted. The design of that apparatus is left to the next collection and examination.

References

Sources of the method, and works whose existence was confirmed as evidence for the description of the board, are given with DOI/URL. Board evidence inherited from the sister notes carries over its verification status there (some items with [primary verification needed]).

Method (Problem Definition, Problematization, Critical Technical Practice)

  • Alvesson, M. & Sandberg, J. (2011). Generating Research Questions Through Problematization. Academy of Management Review 36(2):247–271. https://doi.org/10.5465/amr.2009.0188
  • Sandberg, J. & Alvesson, M. (2011). Ways of constructing research questions: gap-spotting or problematization? Organization 18(1):23–44. https://doi.org/10.1177/1350508410372151
  • Agre, P. E. (1997). Toward a critical technical practice: Lessons learned in trying to reform AI. In G. C. Bowker, S. L. Star, W. Turner & L. Gasser (Eds.), Social Science, Technical Systems, and Cooperative Work: Beyond the Great Divide (pp. 131–157). Lawrence Erlbaum Associates. https://pages.gseis.ucla.edu/faculty/agre/critical.html
  • Dorst, K. (2011). The Core of ‘Design Thinking’ and Its Application. Design Studies 32(6):521–532. https://doi.org/10.1016/j.destud.2011.07.006
  • Dorst, K. (2015). Frame Innovation: Create New Thinking by Design. MIT Press. ISBN 9780262324311.
  • Kuhn, T. S. (1962). The Structure of Scientific Revolutions. University of Chicago Press. ISBN 0-226-45808-3.

Evidence for the Description of the Board

Scaling Laws and the Reflux of Generated Products (Confirming Convergence with the Endogenous Scaling Law)

  • Shumailov, I. et al. (2024). AI models collapse when trained on recursively generated data. Nature 631(8022):755–759. https://doi.org/10.1038/s41586-024-07566-y
  • Dohmatob, E., Feng, Y., Yang, P., Charton, F., Kempe, J. (2024). A Tale of Tails: Model Collapse as a Change of Scaling Laws. ICML 2024. arXiv:2402.07043. https://arxiv.org/abs/2402.07043
  • Gerstgrasser, M. et al. (2024). Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data. arXiv:2404.01413. https://arxiv.org/abs/2404.01413

Items Flagged for Primary Verification ([primary verification needed])

  • For the first RQ candidate (the provenance of benchmark authority): the extent of coverage by prior work in the sociology of benchmarks and on the meta-science side.
  • For the second RQ candidate (the selection of exploration by the record-keeping institution): the extent of overlap with existing research on publication bias and negative results.
  • The quantitative values in Koch et al. (2021) (the degree of usage concentration on a few institutions). The sister notes’ verification status carries over.

Footnotes

  1. Sandberg, J. & Alvesson, M. (2011). Ways of constructing research questions: gap-spotting or problematization? Organization 18(1):23–44. https://doi.org/10.1177/1350508410372151

  2. Alvesson, M. & Sandberg, J. (2011). Generating Research Questions Through Problematization. Academy of Management Review 36(2):247–271. https://doi.org/10.5465/amr.2009.0188

  3. Birhane, A., Kalluri, P., Card, D., Agnew, W., Dotan, R., Bao, M. (2022). The Values Encoded in Machine Learning Research. FAccT ‘22. https://doi.org/10.1145/3531146.3533083 2

  4. Raji, I. D., Bender, E. M., Paullada, A., Denton, E., Hanna, A. (2021). AI and the Everything in the Whole Wide World Benchmark. NeurIPS 2021 D&B. https://arxiv.org/abs/2111.15366 2

  5. Wagstaff, K. L. (2012). Machine Learning that Matters. ICML 2012. https://arxiv.org/abs/1206.4656

  6. Koch, B., Denton, E., Hanna, A., Foster, J. G. (2021). Reduced, Reused and Recycled: The Life of a Dataset in Machine Learning Research. NeurIPS 2021 D&B. https://arxiv.org/abs/2112.01716

  7. Messeri, L. & Crockett, M. J. (2024). Artificial intelligence and illusions of understanding in scientific research. Nature 627(8002):49–58. https://doi.org/10.1038/s41586-024-07146-0

  8. ICBINB (“I Can’t Believe It’s Not Better!”) NeurIPS Workshop (2020, 2021, 2023). 2020 proceedings: https://proceedings.mlr.press/v137/ 2

  9. Kapoor, S. & Narayanan, A. (2023). Leakage and the reproducibility crisis in machine-learning-based science. Patterns 4(9):100804. https://doi.org/10.1016/j.patter.2023.100804

  10. Pineau, J. et al. (2021). Improving Reproducibility in Machine Learning Research. JMLR 22(164):1–20. https://www.jmlr.org/papers/v22/20-303.html

  11. Shumailov, I. et al. (2024). AI models collapse when trained on recursively generated data. Nature 631(8022):755–759. https://doi.org/10.1038/s41586-024-07566-y 2

  12. Agre, P. E. (1997). Toward a critical technical practice: Lessons learned in trying to reform AI. In G. C. Bowker, S. L. Star, W. Turner & L. Gasser (Eds.), Social Science, Technical Systems, and Cooperative Work: Beyond the Great Divide (pp. 131–157). Lawrence Erlbaum Associates. Full text: https://pages.gseis.ucla.edu/faculty/agre/critical.html

  13. Dohmatob, E., Feng, Y., Yang, P., Charton, F., Kempe, J. (2024). A Tale of Tails: Model Collapse as a Change of Scaling Laws. ICML 2024. arXiv:2402.07043. https://arxiv.org/abs/2402.07043

  14. Gerstgrasser, M. et al. (2024). Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data. arXiv:2404.01413. https://arxiv.org/abs/2404.01413


← All Notes · Home