Shuichiro Ogawa
日本語

Notes · updated 2026-07-19

The Gaps We See, and the Voids We Don’t

A field usually has only one way to count its unexplored territory. It traces the outline of its known problems and extends a line beyond them. Unsolved benchmarks, the open problems listed at the end of a survey, the next hole in performance to be filled. The AAAI “Future of AI Research” report, published in March 2025, is the archetype. Layering 475 survey responses onto 25 researchers, its seventeen chapters each enumerate challenges along axes like reasoning, factuality, alignment, agents, and interpretability1. This is a fine act of a field mapping itself, but the way the map is drawn is never questioned.

Such enumeration can be called, in Dorst’s terms, abduction-1. It is reasoning that searches for an unknown answer while keeping the known working principle (what to measure, what to count as good) fixed2. When an alignment survey derives sub-problems through the RICE principles, or an interpretability survey counts the actionability gap, it stays within this type3. These ask what should be measured, but not what the measuring framework itself fails to see.

What this note attempts is the other type. What Dorst called abduction-2 is reasoning in which both the answer and the working principle are unknown at once, finding the “what” and the “how” together2. At work there is a frame: the invention of a new way of seeing, “if I look at the situation through this working principle, such a value comes into view”4. It is a species of the leap Peirce named abduction (also called retroduction), “the reasoning that forms a hypothesis to explain a surprising fact”5. The method runs like this. Take one premise that an existing lens treats as self-evident and never questions, and invert it. Describe the gap that comes into view only when seen through the inverted lens.

This leap is a hypothesis, not a confirmation. So this note writes each gap in the posture of a conditional: “if we look through this lens, such a void should come into view.” And, to keep the leap from drifting into fantasy, it imposes two constraints. First, immediately after each gap proposition it appends an evidence-level tag [evidence: strong|circumstantial|author's conjecture], making explicit how strong the grounds are that the gap is “currently thin.” Second, in every frame it names the existing partial research, narrowing the gap precisely from “wholly unexplored” to “the absence of an integrative methodology,” “an asymmetry of institutionalization,” or “insufficient depth.” It does not exaggerate. The seven frames were selected, under these two constraints, only if they satisfied four conditions: non-obviousness, generativity, falsifiability, and a structural reason for being overlooked.

Frame 1: See AI as a Presence That Settles In Over Time

AI is usually measured as a snapshot. A benchmark turns the capability at one point in time into a number, and the human subject too is treated as a fixed entity. Most experiments are one-shot, closing within minutes to a few days. Invert it. What if we see AI as a presence that settles persistently into humans and institutions over months to years, with each transforming the other?

Through this lens, a gap in the timeline of skill rises first. Lee et al.’s study, presented at CHI 2025, had 319 knowledge workers self-report 936 instances, finding that higher trust in generative AI accompanied lower critical thinking6. Gerlich’s survey of 666 people, in Societies the same year, likewise found a negative correlation between frequent AI use and critical thinking, mediated by cognitive offloading7. Both are important observations, but the method is cross-sectional self-report, and the authors themselves concede they can specify neither causation nor temporal development. That is, how sustained use reshapes the trajectory of mastery over years is invisible from these cross-sections. That a five-year longitudinal study of the relationship between generative AI and creators positions itself as a “rare case that supplements prior cross-sectional work along the time axis” shows, in reverse, that longitudinal work is exceptional8. The longitudinal dynamics of how sustained use drives the de-formation and re-formation of skill are thin as an integrative research program. [evidence: circumstantial]

Yet calling this void “untrodden” would stumble. Budzyń et al.’s multicentre observational study, published in The Lancet Gastroenterology & Hepatology in August 2025, captures exactly the timeline of skill9. Across 19 experienced endoscopists at four Polish centres, with the introduction and removal of AI assistance bracketing the period, the adenoma detection rate for non-AI-assisted colonoscopy fell from 28.4% to 22.4% (a relative 20% drop). It is the first empirical record of deskilling by a clinical AI. Because it exists, Frame 1’s gap is not “no longitudinal work at all.” Precisely, it is the limitation that a longitudinal design tracking the de-formation and re-formation of cognition in general, beyond a single domain’s observation, has not been established. Dell’Acqua et al.’s 2023 field experiment with 758 BCG consultants likewise showed large gains in productivity and quality on within-frontier tasks, but this was a one-shot measurement of effect—the very snapshot evaluation Frame 1 targets10.

The same lens illuminates the temporal decay of capability itself. Model collapse, shown by Shumailov et al. in Nature in 2024, is the phenomenon where, when generations recursively flow back into future training, the tails of the distribution vanish and diversity and quality degrade irreversibly11. It is a representative primary study of system-level temporal decay. But this conclusion, too, is not monolithic. Gerstgrasser et al. proved in a linear model that if synthetic data is accumulated onto real data rather than replacing it, collapse is avoided and test error has a finite bound independent of the number of iterations, confirming this in language models, diffusion models, and VAEs12. Borji reread collapse as a statistical inevitability of repeated sampling13. That whether decay occurs is still contested is itself circumstantial evidence that this void is unsettled. A static view of AI overlooks these because benchmark culture rewards static, reproducible, comparable numbers, and longitudinal design is costly and hard to fit onto the annual peer-review cycle.

ai-psychology-cognitive-science-trends, which maps the currents of psychology and cognitive science, has vigorously tracked AI’s effects on cognition. What that map presupposed was a framework that measures effects as a difference at one point in time. The process by which AI settles into human cognition, and humans in turn are reshaped over years, is left outside that cross-section.

Frame 2: The Instrument of Observation Is Made of the Same Stuff as What It Observes

AI researchers have been assumed to stand outside the system they study. AI is the object; the researcher is a neutral observer. Invert it. AI is now inside the process of research itself. Literature review, ideation, coding, data analysis, and peer review have begun to be handled by LLMs. The instrument of observation is made of the same stuff as what it observes.

Through this lens, a gap in the homogenization of ideation rises. Messeri and Crockett’s 2024 essay in Nature argued that AI draws scientists in with promises of productivity and objectivity while exploiting cognitive limits to breed an “illusion of understanding,” and renders invisible, to the community itself, a scientific monoculture in which certain methods, questions, and perspectives come to dominate14. It is a warning that asks not about hallucination but about what happens when AI works exactly as intended. A caution is due here. This essay is a theoretical warning; measurement that the exploration space has actually contracted is yet to come, as the authors themselves position it. So the precise gap is that empirical evidence for how AI assistance reshapes the epistemology of discovery is thin. [evidence: circumstantial]

This void, too, is not “untrodden.” Homogenization shows opposite faces at the individual and the collective level. Si et al.’s 2024 study (accepted at ICLR 2025), which had 100+ NLP researchers give blind evaluations, showed that LLM-generated ideas were higher in novelty than human experts’ (p<0.05)15. At the individual level this is a counterexample that weakens Frame 2’s homogenization hypothesis. Yet the same study explicitly flags, as an open problem, that at scale ideas converge and duplicate, and generative diversity is lacking. At the collective level, it supports Frame 2. Creative homogeneity, in which LLM outputs resemble one another in meaning, structure, and style more than a human population does, is also reported by several empirical studies16. That is, the gap lies not on the side of individual novelty but on the side of a reflexive meta-science that measures the process by which the exploration space narrows at collective scale.

The instrument’s entry into its object reaches peer review, too. Sakana AI’s The AI Scientist automates the whole cycle from ideation through experiment, writing, and review, with v2 claiming workshop-level automated discovery17. It is a concrete case of AI beginning to carry the research process itself. Meanwhile, a randomized study of 20,000 reviews at ICLR 2025 reports that LLM-assisted feedback significantly improved clarity and actionability18. This is a scope-limiting counterexample that the influx of LLMs into review is not uniformly harmful. Even so, the systematic measurement of the reflexive effect that AI-written AI research exerts on the field is not yet in place. That reflexivity is not native to computer science, having remained a habit of STS and the humanities, is the structural reason for this blind spot. [evidence: circumstantial]

ai-social-science-research-trends, which maps the currents of the social sciences, carefully tracked the methods of using LLMs as research tools and the critiques of that use’s validity. What that map presupposed was a separation between the researcher who uses the tool and the tool being used. The reflexive loop, in which the tool rewrites the field’s own output and that effect never registers in the field’s own eyes, keeps turning behind that separation.

Frame 3: See AI Not as an Individual but as a Member of an Ecology

Evaluation is done in isolation: one model, one user, one task. Invert it. See AI as a member of an ecology. Many models interact, humans and AI and environment form one system, and markets and institutions condition its behavior.

Through this lens, a gap in collective-scale emergence rises. But this is a domain where existing research is thick, so the scope must be narrowed with care. What Ashery et al. showed in Science Advances in 2025 was that a population of LLM agents without central coordination (N=24, up to N=200 in robustness checks) spontaneously forms shared naming conventions, and collective-level bias emerges even when individuals are unbiased19. What Calvano et al. showed in American Economic Review in 2020 was that Q-learning pricing algorithms learn supracompetitive prices without explicit communication, forming a stable tacit collusion20. Both norm emergence and algorithmic collusion have already been demonstrated. So “AI as an ecology is wholly unexplored” would be an exaggeration.

The precise gap lies on the side of a unit that cuts across all these cross-sections. Norm emergence is an experiment within the naming-game task, not behavior embedded in real-world institutions and markets. Algorithmic collusion is a model from a simple duopoly to an oligopoly, not an ecology in which many LLM agents entangle across many tasks. System-level feedback (generations flowing back into training) was demonstrated as model collapse (cited above11), but that was mainly a single model’s closed loop. That is, no integrative evaluation unit that treats many models, humans, institutions, and markets as one system has been established. [evidence: circumstantial] That a technical report by 51 authors of the Cooperative AI Foundation declares, as insiders, the risks of concurrently deploying multiple agents to be “novel and under-explored,” organizing failure modes into three types (Miscoordination, Conflict, Collusion), supports the absence of this integrative unit from inside the field21. In finance, Gensler—later SEC chair—argued in 2020 that dependence on a shared foundation model could heighten homogeneity and concentration risk22. It is a precedent for the ecological view, but it stays a policy discussion in a single domain. The benchmark paradigm itself, which cuts the unit of evaluation down to the individual, structurally renders system-level phenomena invisible.

ai-economics-research-trends, which maps the currents of economics, has treated AI’s effects on markets and labor with precision. What that map presupposed was a construction in which AI is injected into the market as an exogenous shock. The dynamics by which many AIs, humans, and institutions condition one another and co-evolve as one system cannot be raised as a unit within the exogenous-shock framework.

Frame 4: Treat Failure and Non-Knowledge as First-Class Knowledge

The field publishes successes. SOTA, positive results, methods that worked. Invert it. Treat failure, negative results, and “why it does not work” as first-class knowledge. Make what cannot be measured an object of research.

Through this lens, a gap in the asymmetry of failure’s systematization rises. Here too there is genuine existing research. The data leakage reported by Kapoor and Narayanan in Patterns in 2023 is a widespread failure mode in machine-learning-based science, affecting papers on the order of hundreds across 17 fields, organized into 8 types23. It is a rare example of a systematic archive of failure modes. NeurIPS 2019’s reproducibility program institutionalized a code-submission policy, a reproducibility challenge, and a checklist24. The ICBINB workshop, which champions the sharing of negative results, has been held repeatedly since 202025. So “the systematization of non-knowledge is nonexistent” would be wrong.

The precise gap lies on the side of the asymmetry. The receptacle for reproducibility stays at the stage of checklists and challenges, and the venue for negative results sits not in a main-conference track but in sporadic workshops. Non-knowledge is not institutionalized as first-class knowledge; it is marginalized. [evidence: circumstantial] This asymmetry is rooted in the very valuation of the publishing culture. Birhane et al.’s 2022 FAccT paper, which analyzed 100 highly cited ML papers, showed that the six most frequent values were performance, generalization, quantitative evidence, efficiency, building on past work, and novelty; that only 15% justified a connection to a societal need; and that a mere 1% discussed negative potential impacts26. As for the object of measurement itself being fixed onto a few, Koch et al.’s dataset analysis reports that usage concentrates on datasets originating from a few elite institutions, with concentration rising to a Gini above 0.80 in recent years27.

The same lens illuminates the concealment of the unmeasurable. As Raji et al. argued at NeurIPS 2021, a few “general” benchmarks are venerated as proxies for general-purpose ability, yet that framing lacks construct validity and severely misrepresents capability28. The classic critique of the streetlight effect—looking only where the light falls—which Wagstaff formulated in 2012, lies at its source29. Further, there is the asymmetry in which one can show that something works but the mechanistic theory of “why it works” lags. Zhang et al. showed that SOTA convolutional nets can easily memorize random labels, and that classical generalization theory cannot explain why deep nets generalize30. Mechanistic interpretability is an active object of research, but Williams et al.’s 2025 position paper points out, as insiders, that MI research is at a pre-paradigmatic stage, unable to define “what counts as a valid explanation”31. The streetlight effect, publication bias, and the incentive to chase SOTA keep illuminating only what can be measured. The lag on the side of non-knowledge is not “absence” but a “speed difference” relative to the side of success. [evidence: circumstantial]

ai-humanities-digital-humanities-trends, which maps the currents of the humanities and digital humanities, tracked how AI handles meaning, interpretation, and text. What that map presupposed was the optimism that the object can be read, that meaning can be extracted. The work of actively thematizing what is structurally unreadable, what remains outside measurement, grows thin behind that optimism.

Frame 5: Place Cognitive Diversity at the Origin of Design, Not as an Exception

The implicit user is a literate adult of standard ability. Often a WEIRD population: from a society that is Western, educated, industrialized, rich, and democratic. Invert it. Place cognitive and bodily diversity at the center. Make disabled people, learners in progress, older people, low-literacy users, and neurological minorities the origin of design, not exceptions.

Through this lens, a gap in the encoding of the default user rises. Its scholarly source is Henrich et al.’s 2010 critique that most subjects of behavioral science derive from a WEIRD population that is a fraction of the world’s people32. The same bias is inscribed in AI research. When Septiandri et al. analyzed 128 human-subject studies in fairness research at FAccT 2023, 84% relied on Western-country participants only, and 63% on U.S. participants only33. The very research that discusses fairness stands on an atypical sample of the world’s population.

This void, as individual studies, does exist. Carik et al. qualitatively analyzed the discussions of 61 neurodivergent communities, showing a construction in which LLM responses are excessively neurotypical and users fill the gap with community-driven workarounds like prompt sharing34. In Jang et al.’s CHI 2024 study, 11 autistic workers strongly preferred an LLM for workplace communication, but professional coaches rated its advice “questionable”35. User satisfaction and expert standards diverge, and a residue remains that satisfaction-based evaluation cannot capture. That these studies fill the gap not with AI-side adaptation but with adaptation on the users’ side, and that design proceeds with data on disabled users missing, is reported repeatedly.

So the precise gap lies on the side of integration and institutionalization. Sporadic individual studies exist, but there is no integrative program that places cognitive diversity at the origin of design and the default of evaluation. [evidence: strong] That Moharana et al. showed, at AIES 2025, from interviews with 25 AI practitioners, that the intersection of responsible AI and accessibility is siloed within organizations, with volunteers and grassroots filling the absence of formal structure, is primary-source evidence of this construction36. That benchmarks encode the default user in the first place, and that reward models depend on the preferences of a few annotators, support this bias as a mechanism. Kirk et al.’s PRISM dataset (1,500 people across 75 countries, NeurIPS 2024) is a fine example that squarely thematizes which humans provide which alignment data37, but the gap lies precisely in that such efforts are exceptional and not standardized.

ai-education-learning-sciences-trends, which maps the currents of education and the learning sciences, has eagerly tracked AI’s effects on learners. What that map presupposed was an average image of the standard learner. The view that places learners with non-standard cognition at the starting point of design, rather than as deviations from the average, is left in the shadow of that average image.

Frame 6: See AI’s Knowledge as Culturally Situated

Training data, evaluation, and “knowledge” are assembled around the Anglophone and Western center, and the concepts that hold there are assumed universal. Invert it. See AI’s knowledge as culturally situated. Raise non-Western epistemologies and ethical frameworks as objects to be studied in their own right, not objects to be translated.

This frame differs from the others. It is a frame of “existing research exists, but shallow,” not of “unexplored.” The Western-centric encoding has been demonstrated at scale. Tao et al. showed in PNAS Nexus in 2024 that GPT-family outputs align most closely, on the World Values Survey scales, with the Anglophone world and Protestant Europe, diverging greatly from Jordan, Libya, and Ghana38. BLEnD (NeurIPS 2024) measured, over 52.6k QA, that the more a culture is represented online the higher the LLM’s performance, with a gap of up to 57.34% between cultures even for GPT-439. The side of thought is established, too. Mohamed et al. theorized decolonial AI in 2020 as a tool of sociotechnical foresight40. Sambasivan et al. argued, in the Indian context, that mainstream fairness research is West-centric in its subgroups, values, and methods, going as far as a “re-imagining” of the metrics41. So writing “unexplored gap” here would be immediately refuted by this research.

The gap lies on the side of depth and reach, not quantity. First, a unified definition of the concept of “culture” is absent. Liu et al.’s systematic survey of culturally aware NLP (TACL 2024/2025) declares that a shared understanding of “culture” remains unclear, proposing a fine-grained taxonomy while enumerating the gaps to be filled42. Second, the engineering operationalization of non-Western ethical frameworks is missing. Ubuntu, Confucian virtue ethics, and Buddhist compassion may be mentioned within essays, but their operationalization—embedding them into evaluation and training—is almost untouched43. Third, and most vexing, is the trade-off between raising cultural diversity and universal human rights. When Zhou et al. measured five LLMs on the WVS at AIES 2025, weakening alignment to Western values increased cultural diversity but raised outputs violating human rights (especially gender equality) by 2–4%44. Fact presses a tension onto the naive prescription “make the non-West an object in its own right.” The gap narrows to these three points: the absence of a unified definition of the concept of culture, the missing engineering operationalization of non-Western ethics, and the unresolved trade-off between diversity and universal human rights. [evidence: strong] Data availability, the language used for evaluation, and the institution of publishing in English support this insufficiency of depth.

ai-arts-culture-studies-trends, which maps the currents of the arts and cultural studies, tracked how AI engages the generation and interpretation of culture. What that map presupposed was the viewpoint of an observer who regards culture as an object. Which culture’s values are embedded in the evaluation metrics themselves, whose “common sense” is encoded, is laid at that observer’s feet and is hard to question.

Frame 7: A Field Maps Its Own Negative

A field maps its own progress. Leaderboards, surveys, roadmaps. But it has no method to map its own negative, its structurally invisible. Invert it. Apply abduction-2 recursively to the field itself, and raise, as a research program, a reflective method that systematically draws the blind spots.

This frame is the most conjectural of the seven. The tools one can cite in support are all general philosophy of science, not specific to AI. Kuhn argued that during normal science, researchers assume the paradigm is correct and refine details, never questioning the framework itself45. Blind spots become structurally invisible from inside the paradigm. Wimsatt showed that we, as limited beings, build knowledge on heuristics, that those heuristics carry detectable systematic biases, and that those biases expose the heuristics’ operation46. The streetlight effect—illuminating only what can be measured—was formulated by Kaplan in 1964 as “the drunkard’s search”47. Applying these recursively to AI research is Frame 7’s leap. That leap is a hypothesis, not a confirmation. [evidence: author's conjecture]

There is one piece of recent empirical evidence that the leap is not fantasy. Traberg et al. argued in Communications Psychology in 2026 that the rush into generative-AI research produces a convergence feedback of topic, method, and language, flattening scientific imagination48. Diverse questions are reframed into “AI and ~,” LLMs become the standard analytic tool displacing experimental, qualitative, and ethnographic traditions, and phrases like “trustworthy AI” recur. That the field is contracting its own exploration space—one cross-section of it—is being measured. This shows that Frame 7’s reflexive meta-science does partially exist. So Frame 7’s gap, too, is not “absence.” It is the limitation that a reflective method to systematically map “unknown unknowns,” beyond sporadic empirical work, does not stand as a research program. [evidence: author's conjecture]

A survey is an extrapolation from what exists. It is a product of abduction-1, and it does not see what is structurally excluded. This note itself does not escape the limit of that extrapolation. The seven frames laid out here are no more than inversions of the premises the author’s vantage could illuminate; the premises it could not illuminate go uncounted. The instrument that measures progress cannot see the very frame of that instrument.

ai-law-governance-research-trends, which maps the currents of law and governance, tracked by what one regulates AI and by what one measures to govern it. What that map presupposed was the assumption that the contour of what must be governed is already visible. A method to map what the net of governance structurally cannot scoop, what the framework of measurement excludes as negative, lies outside that contour.

So, without closing, one question is left. Suppose a field acquires a method to map its own blind spots—who, then, and how, maps the blind spots of that method itself? Following Wimsatt, an instrument’s bias should be detectable and should expose its own operation. But the next instrument charged with that detection also has its own negative. Where this regress halts, or whether it halts at all, is beyond this note’s reach.

  • ai-capability-relational-ontology — A sequel that deepens one of the gaps named here through a de Broglie-style completion of symmetry: a relational ontology in which measurement does not reveal but enacts capability, and a test for its contextuality
  • ai-research-gaps-revisited — A sequel that revisits this catalog with a sharpened discipline (assumption removal → symmetry inversion → operationalization, the A–G sheet), sorts symmetry-completion blind spots from coverage gaps, and surfaces a new one: the endogenous scaling law

References

The methodological sources, and the literature whose existence was verified as evidence for each frame, are given with DOIs/URLs. Items whose bibliography is partly unconfirmed or whose full text was not reached carry [primary-source verification needed]. The internal working ledger, including provenance, confidence, and scope-limiting material, is source/review/ai-research-gaps-abduction/papers.md (internal to the repository, not published).

Method (abduction-2, frames, hypothesis formation)

  • Dorst, K. (2011). The Core of ‘Design Thinking’ and Its Application. Design Studies 32(6), 521–532. https://doi.org/10.1016/j.destud.2011.07.006
  • Dorst, K. (2015). Frame Innovation: Create New Thinking by Design. MIT Press. ISBN 9780262324311.
  • Peirce, C. S. (1933–1958). Collected Papers of Charles S. Peirce (8 vols.), eds. C. Hartshorne, P. Weiss & A. W. Burks. Harvard University Press. (Specific volume/paragraph numbers are [primary-source verification needed].)

Deductive baseline, meta-science, philosophy of science

  • AAAI (2025). Presidential Panel on the Future of AI Research. https://aaai.org/about-aaai/presidential-panel-on-the-future-of-ai-research/ (chapter bodies are [primary-source verification needed])
  • Ji, J. et al. (2023). AI Alignment: A Comprehensive Survey. arXiv:2310.19852. https://arxiv.org/abs/2310.19852
  • Birhane, A., Kalluri, P., Card, D., Agnew, W., Dotan, R., Bao, M. (2022). The Values Encoded in Machine Learning Research. FAccT ‘22. https://doi.org/10.1145/3531146.3533083
  • Koch, B., Denton, E., Hanna, A., Foster, J. G. (2021). Reduced, Reused and Recycled: The Life of a Dataset in Machine Learning Research. NeurIPS 2021 D&B. https://arxiv.org/abs/2112.01716 (quantitative values are [primary-source verification needed])
  • Kuhn, T. S. (1962). The Structure of Scientific Revolutions. University of Chicago Press. ISBN 0-226-45808-3.
  • Wimsatt, W. C. (2007). Re-Engineering Philosophy for Limited Beings. Harvard University Press. ISBN 9780674015456.
  • Kaplan, A. (1964). The Conduct of Inquiry: Methodology for Behavioral Science. Chandler. (streetlight effect, p. 11)
  • Traberg, C. S., Roozenbeek, J., van der Linden, S. (2026). AI is turning research into a scientific monoculture. Communications Psychology 4:37. https://doi.org/10.1038/s44271-026-00428-5

Frame 1 (timeline) and Frame 3 (ecology)

Frame 2 (research reflexivity)

Frame 4 (the absence of systematized non-knowledge and failure)

Frame 5 (cognitive diversity) and Frame 6 (the monism of epistemology)

Unverified items ([primary-source verification needed])

  • The specific volume/paragraph numbers of Peirce’s Collected Papers (citation numbers vary across sources).
  • The chapter bodies of the AAAI 2025 report (titles and scale confirmed via the landing page; body is an unextracted binary PDF).
  • The quantitative values in Koch et al. (2021) (Gini above 0.80; concentration of usage onto a few institutions).
  • The count of affected papers in Kapoor & Narayanan (2023) (preprint 329 vs. Patterns 294; described in-text as “on the order of hundreds across 17 fields”).
  • The body details of Gensler & Bailey (2020) (metadata triangulated; body PDF unextractable).
  • The final publication / peer-review status of The AI Scientist v2 (arXiv:2504.08066) and the LLM peer-review RCT (arXiv:2504.09737).
  • The journals/volumes of the two creative-homogeneity papers (ScienceDirect S294988212500091X / arXiv:2501.19361).
  • Direct retrieval of the original DOI of Henrich et al. (2010) (content consistent across multiple secondary sources).
  • The primary figures such as “low-resource languages are 6.2% of evaluation benchmarks” in Liu et al. (2024, TACL).
  • The original sources of the body of essays on the uptake of non-Western ethics (ubuntu / Confucian / Buddhist) into AI ethics.
  • The peer-review status of Frame 1’s rare longitudinal case, arXiv:2511.03117.

Footnotes

  1. AAAI (2025). Presidential Panel on the Future of AI Research. Association for the Advancement of Artificial Intelligence. https://aaai.org/about-aaai/presidential-panel-on-the-future-of-ai-research/ — chapter titles and scale confirmed via the landing page; chapter bodies remain [primary-source verification needed].

  2. Dorst, K. (2011). The Core of ‘Design Thinking’ and Its Application. Design Studies 32(6), 521–532. https://doi.org/10.1016/j.destud.2011.07.006 2

  3. Ji, J. et al. (2023). AI Alignment: A Comprehensive Survey. arXiv:2310.19852. https://arxiv.org/abs/2310.19852

  4. Dorst, K. (2015). Frame Innovation: Create New Thinking by Design. MIT Press. ISBN 9780262324311.

  5. Peirce, C. S. Collected Papers of Charles S. Peirce (8 vols.), eds. Hartshorne, Weiss & Burks. Harvard University Press, 1933–1958. Abduction (later also retroduction / presumption). Specific volume and paragraph numbers are [primary-source verification needed].

  6. Lee, H.-P. et al. (2025). The Impact of Generative AI on Critical Thinking. CHI 2025. https://doi.org/10.1145/3706598.3713778

  7. Gerlich, M. (2025). AI Tools in Society: Impacts on Cognitive Offloading and the Future of Critical Thinking. Societies 15(1):6. https://doi.org/10.3390/soc15010006

  8. As a rare five-year longitudinal case, arXiv:2511.03117 (peer-review status is [primary-source verification needed]).

  9. Budzyń, K. et al. (2025). Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study. The Lancet Gastroenterology & Hepatology. https://www.thelancet.com/journals/langas/article/PIIS2468-1253(25)00133-5/abstract

  10. Dell’Acqua, F. et al. (2023). Navigating the Jagged Technological Frontier. Harvard Business School WP 24-013 / SSRN 4573321. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4573321

  11. Shumailov, I. et al. (2024). AI models collapse when trained on recursively generated data. Nature 631(8022):755–759. https://doi.org/10.1038/s41586-024-07566-y 2

  12. Gerstgrasser, M. et al. (2024). Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data. arXiv:2404.01413. https://arxiv.org/abs/2404.01413

  13. Borji, A. (2024). A Note on Shumailov et al. (2024). arXiv:2410.12954. https://arxiv.org/abs/2410.12954

  14. Messeri, L. & Crockett, M. J. (2024). Artificial intelligence and illusions of understanding in scientific research. Nature 627(8002):49–58. https://doi.org/10.1038/s41586-024-07146-0

  15. Si, C., Yang, D., Hashimoto, T. (2024). Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers. arXiv:2409.04109 (ICLR 2025). https://arxiv.org/abs/2409.04109

  16. For example, “Homogenizing effect of LLMs on creative diversity” (ScienceDirect S294988212500091X) and “We’re Different, We’re the Same” (arXiv:2501.19361). Journal and volume are [primary-source verification needed].

  17. Sakana AI, The AI Scientist v2. arXiv:2504.08066. https://arxiv.org/pdf/2504.08066 — third-party evaluation Beel et al., arXiv:2502.14297. Peer-review status is [primary-source verification needed].

  18. Can LLM feedback enhance review quality? arXiv:2504.09737 (ICLR 2025, RCT over 20,000 reviews). https://arxiv.org/pdf/2504.09737 — final publication status is [primary-source verification needed].

  19. Ashery, A. F., Aiello, L. M., Baronchelli, A. (2025). Emergent social conventions and collective bias in LLM populations. Science Advances 11(20):eadu9368. https://doi.org/10.1126/sciadv.adu9368

  20. Calvano, E., Calzolari, G., Denicolò, V., Pastorello, S. (2020). Artificial Intelligence, Algorithmic Pricing, and Collusion. American Economic Review 110(10):3267–3297. https://doi.org/10.1257/aer.20190623

  21. Hammond, L., Chan, A., Clifton, J. et al. (2025). Multi-Agent Risks from Advanced AI. Cooperative AI Foundation. arXiv:2502.14143. https://arxiv.org/abs/2502.14143

  22. Gensler, G. & Bailey, L. (2020). Deep Learning and Financial Stability. MIT / SSRN 3723132. https://papers.ssrn.com/sol3/papers.cfm?abstract_id=3723132 — body details are [primary-source verification needed].

  23. Kapoor, S. & Narayanan, A. (2023). Leakage and the reproducibility crisis in machine-learning-based science. Patterns 4(9):100804. https://doi.org/10.1016/j.patter.2023.100804 — the count of affected papers differs from the preprint version; described here as “on the order of hundreds across 17 fields.”

  24. Pineau, J. et al. (2021). Improving Reproducibility in Machine Learning Research. JMLR 22(164):1–20. https://www.jmlr.org/papers/v22/20-303.html

  25. ICBINB (“I Can’t Believe It’s Not Better!”) NeurIPS Workshop (2020, 2021, 2023). 2020 proceedings: https://proceedings.mlr.press/v137/

  26. Birhane, A. et al. (2022). The Values Encoded in Machine Learning Research. FAccT ‘22. https://doi.org/10.1145/3531146.3533083https://arxiv.org/abs/2106.15590

  27. Koch, B., Denton, E., Hanna, A., Foster, J. G. (2021). Reduced, Reused and Recycled: The Life of a Dataset in Machine Learning Research. NeurIPS 2021 D&B. https://arxiv.org/abs/2112.01716 — quantitative values are [primary-source verification needed].

  28. Raji, I. D., Bender, E. M., Paullada, A., Denton, E., Hanna, A. (2021). AI and the Everything in the Whole Wide World Benchmark. NeurIPS 2021 D&B. https://arxiv.org/abs/2111.15366

  29. Wagstaff, K. L. (2012). Machine Learning that Matters. ICML 2012. https://arxiv.org/abs/1206.4656

  30. Zhang, C., Bengio, S., Hardt, M., Recht, B., Vinyals, O. (2021). Understanding deep learning (still) requires rethinking generalization. Communications of the ACM 64(3):107–115. https://doi.org/10.1145/3446776

  31. Williams, I. et al. (2025). Mechanistic Interpretability Needs Philosophy. arXiv:2506.18852. https://arxiv.org/abs/2506.18852

  32. Henrich, J., Heine, S. J., Norenzayan, A. (2010). The weirdest people in the world? Behavioral and Brain Sciences 33(2–3):61–83. https://doi.org/10.1017/S0140525X0999152X — direct retrieval of the original DOI is [primary-source verification needed]; content is consistent across multiple secondary sources.

  33. Septiandri, A. A., Constantinides, M., Tahaei, M., Quercia, D. (2023). WEIRD FAccTs: How Western, Educated, Industrialized, Rich, and Democratic is FAccT? ACM FAccT 2023. https://arxiv.org/abs/2305.06415

  34. Carik, B., Ping, K., Ding, X., Rho, E. H. R. (2024). Exploring Large Language Models Through a Neurodivergent Lens. arXiv:2410.06336. https://doi.org/10.48550/arXiv.2410.06336

  35. Jang, J. Y., Moharana, S., Carrington, P., Begel, A. (2024). “It’s the only thing I can trust”: Envisioning Large Language Model Use by Autistic Workers for Communication Assistance. CHI ‘24. https://doi.org/10.1145/3613904.3642894

  36. Moharana, S. et al. (2025). “Accessibility people, you go work on that thing of yours over there”: Addressing Disability Inclusion in AI Product Organizations. AIES 2025. https://doi.org/10.1609/aies.v8i2.36669https://arxiv.org/abs/2508.16607

  37. Kirk, H. R. et al. (2024). The PRISM Alignment Dataset. NeurIPS 2024 D&B. https://doi.org/10.48550/arXiv.2404.16019

  38. Tao, Y., Viberg, O., Baker, R. S., Kizilcec, R. F. (2024). Cultural bias and cultural alignment of large language models. PNAS Nexus 3(9):pgae346. https://doi.org/10.1093/pnasnexus/pgae346

  39. Myung, J. et al. (2024). BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages. NeurIPS 2024 D&B. https://arxiv.org/abs/2406.09948

  40. Mohamed, S., Png, M.-T., Isaac, W. (2020). Decolonial AI: Decolonial Theory as Sociotechnical Foresight in Artificial Intelligence. Philosophy & Technology 33(4):659–684. https://doi.org/10.1007/s13347-020-00405-8

  41. Sambasivan, N. et al. (2021). Re-imagining Algorithmic Fairness in India and Beyond. FAccT 2021. https://arxiv.org/abs/2101.09995

  42. Liu, C., Gurevych, I., Korhonen, A. (2024). Culturally Aware and Adapted NLP: A Taxonomy and a Survey of the State of the Art. TACL. https://doi.org/10.48550/arXiv.2406.03930 — some statistics are [primary-source verification needed].

  43. The uptake of ubuntu / Confucian / Buddhist ethics into AI ethics remains fragmentary, per a body of essays (original sources are [primary-source verification needed]).

  44. Zhou, Y., Constantinides, M., Quercia, D. (2025). Should LLMs be WEIRD? Exploring WEIRDness and Human Rights in Large Language Models. AIES 2025. https://arxiv.org/abs/2508.19269

  45. Kuhn, T. S. (1962). The Structure of Scientific Revolutions. University of Chicago Press. ISBN 0-226-45808-3.

  46. Wimsatt, W. C. (2007). Re-Engineering Philosophy for Limited Beings: Piecewise Approximations to Reality. Harvard University Press. ISBN 9780674015456.

  47. Kaplan, A. (1964). The Conduct of Inquiry: Methodology for Behavioral Science, p. 11 (the formulation of the streetlight effect / drunkard’s search).

  48. Traberg, C. S., Roozenbeek, J., van der Linden, S. (2026). AI is turning research into a scientific monoculture. Communications Psychology 4:37. https://doi.org/10.1038/s44271-026-00428-5


← All Notes · Home