Notes · updated 2026-07-21
Gaps in AI Research, a Fourth Time: Seen from Outside the Inversion Family
An essay that rereads the three sibling notes that searched for gaps in AI research (the seven frames, the A–G operational sheet, the game board of Problem Definition) from the vantage of the full set of problem-setting methods now assembled.
Contents (10)
- The Fourth Time, We Change Not the Generator but the Family of Generators
- The Three Generators Were One Family of Operations
- Morphological Analysis: The Family Cannot Reach the Intersection Cells
- TRIZ: Incompatibilities That Cannot Be Inverted Are Solved by Separation
- PSM: The Two Demoted Candidates Were Another Board
- Cynefin and Representational Change: The Infinite Regress Was a Misclassification of Type
- The Voids Seen from Outside the Family
- Coverage Check: All Three Candidates Were Partial
- The One Side That Stays Open in the Fourth Note Too
- Footnotes
The Fourth Time, We Change Not the Generator but the Family of Generators
We return to the same object a fourth time. The previous three returned each time with a different tool. Where Are the Gaps in AI Research? Seven Voids Found by Reframing searched for gaps in AI research with seven frames that invert one premise at a time; AI Research Gaps, Revisited: Sorting Symmetry-Completion from Coverage, and Digging Out an Endogenous Scaling Law cast that inversion onto an operational sheet from A to G; AI Research Gaps, a Third Time: Rereading Them as Rewritings of the Game Board searched with operations that rewrite the constants of the board. The three tools all looked different.
What we change this time is not the tool. It is the family of the tool. With a map of problem-setting methods gathered from both the scholarly and the practical side in hand (An Academic Map of Methods for Reframing Problems: From Abduction-2 to Problem Structuring and A Genealogy of Practitioner Methods for Reframing Problems: Who Made Them, Traced to Their Origins), it becomes visible that the previous three had been turning within one and the same family. Outside that family, several operations remain that we have not yet applied.
The Three Generators Were One Family of Operations
The three claimed to be independent tools. But lay the operations side by side and they are isomorphic.
Abduction-2 inverts a hidden binary. The A–G sheet inverts an asymmetric binary and checks it against a completion test. The board of Problem Definition rewrites the board’s constants one term at a time.
The problem-definition note itself half concedes this isomorphism. It wrote that “inverting a symmetry is a rewriting specialized to a single term of the board, and Problem Definition is the higher family of operations above it.” Whatever the difference between higher and lower, both share the single move of taking out one fixed binary and inverting it. The three were not separate lineages but one family, binary inversion.
This diagnosis bites against the argument the later two notes set up. The revisited and problem-definition notes took the fact that two generators independently converged on the same void (an endogenous scaling law, the institution of recording) as “weak evidence that the void is not an artifact of the generator’s idiosyncrasy.” But if the generators are of the same family, convergence may be a fixed point of the family of operations more than a property of the object. That applying the same inversion three times lands in the same place is no astonishing fact. The author had honestly reserved that “the same author ran both generators, so a shared vantage cannot be ruled out”; but the reservation runs one level deeper. Not only the vantage but the type of operation was shared.
So the fourth note’s task is settled. Apply operations outside the family, and see whether a void the family structurally could not reach comes into view, or whether the existing voids recombine into a different type.
Morphological Analysis: The Family Cannot Reach the Intersection Cells
The inversion family knocked over the board’s constants one term at a time. Frame 1 time, Frame 2 reflexivity, Frame 3 ecology, and so on, one dimension at a time. All three left the same confession at the end. “The premises left unilluminated are not counted.” An operation that topples one term at a time cannot, in principle, count what remains untoppled.
General Morphological Analysis is precisely a tool for this untoppled remainder. Zwicky structured multidimensional problems exhaustively as a matrix of parameters and options1. Ritchey applied this to wicked problems, making non-quantitative problem complexes visible through decomposition and combination2. The seven terms of the board that the problem-definition note wrote out (recognition of success, measurement, unit, endpoint, position of observation, noise, the institution of recording) become, as they stand, the dimensions of a morphological matrix.
Rather than one term at a time, cross the dimensions and search for unvisited cells.
The most vacant is the intersection of Frame 1 (time), Frame 2 (reflexivity), and Frame 3 (ecology).
Written not as three separate inversions to be summed but as one phenomenon, it comes to this.
It is the dynamics in which a field’s own collective use of AI co-evolves, over the scale of years, as a multi-agent ecology.
Researchers’ ideation is homogenized by LLMs, that population of LLMs emerges collective norms, and over years these remake the field’s exploration space.
The three notes placed these three in separate sections and never crossed them.
Only morphological analysis mechanically points at this cell.
Insofar as the inversion family moves one term at a time, it cannot reach this intersection. [coverage check: partial. Wu et al. (2026) come closest but is an ODE simulation. See the "coverage check" section below.]
TRIZ: Incompatibilities That Cannot Be Inverted Are Solved by Separation
The inversion family had terms that cannot be inverted. Two incompatible demands.
When Zhou et al. measured five LLMs on the World Values Survey, the more they weakened alignment to Western values the more cultural diversity increased, but outputs violating human rights (especially gender equality) rose by 2–4%3. In Si et al.’s blind evaluation, individual LLM ideation was higher in novelty than human experts, yet at scale the ideas converged and lost diversity4. In both, raising one lowers the other. This is not a premise to be inverted but a technical contradiction.
The operation of TRIZ is not inversion. From hundreds of thousands of patents, Altshuller placed the essence of invention in the removal of contradiction, deriving principles that separate the improving from the worsening parameter5. Separation assigns two conflicting demands to different places, by time or space or condition or scale. For diversity and human rights, the axis of separation can be layers. It is a design that places a layer rewarding diversity (surface style, vocabulary, breadth of topic) and a layer fixing human rights as an invariant constraint (a constitutional lower bound) on separate layers.
This bites where Frame 6 ran up against fact. The naive prescription “invert WEIRD and make the non-West the default” stalled before Zhou et al.’s tension. TRIZ turns that spot from a failure of inversion into a design problem. The gap narrows not to “the unexplored territory of non-Western epistemologies” but to “the absence of a procedure to operationalize cultural diversity and universal constraints by separating them into layers.” This is a formulation that a family toppling a single binary structurally cannot produce.
PSM: The Two Demoted Candidates Were Another Board
The revisited note demoted cognitive diversity and non-Western epistemologies (Frames 5 and 6) as “an extension of coverage, not a transposition of a property.” The problem-definition note restored them as “strong questions under a different type.” Both judgments are handed down on a single board.
Problem Structuring Methods begin from a different premise. The core of Soft OR is that a problem is relative to a worldview (Weltanschauung)67. Checkland’s SSM, through the CATWOE frame, externalizes whose view of the customer, whose view of transformation, and whose worldview a given problem definition stands on. Seen through this frame, the board of AI research reveals that “recognition of success” is merely that of a particular worldview.
A neurological minority or a subject of a non-Western epistemology does not hold a different gap on the same board.
They hold another board.
Then the true gap of Frames 5 and 6 is neither an extension of coverage nor a question of a different type.
It is that the field has no method for adjudicating among competing plural problem frames.
AI research has no higher-order procedure that decides which boards to count.
The discussion that separated gap-spotting from problematization recommended a type that reopens the premise itself8, but which to adopt when plural sets of premises compete is a question that discussion too does not answer.
This is the only operation that pierces, from within method, the limit the three notes acknowledged of a single author’s vantage. [coverage check: partial. Reaches Burden et al. / Jo and Wilson, and the adjudication procedure is absent. See the "coverage check" section below.]
Cynefin and Representational Change: The Infinite Regress Was a Misclassification of Type
The three notes get stuck at the same place. Frame 7’s “who maps the blind spot of the very tool that maps blind spots.” The revisited note’s “if the critical point moves, the test that measures it also changes its object.” The problem-definition note’s “what device gives the type that questions the provenance of institutions an equally strong refutation condition.” All are infinite regresses, ending unclosed.
Cynefin diagnoses this as a misclassification of the type of problem situation. Kurtz and Snowden divided situations into ordered and unordered systems, arguing that the legitimate mode of problem perception differs by domain9. Because Frame 7 is treated as a “solvable given a good method” problem, that is, as Complicated, method cannot be written and the regress ensues. But self-mapping is essentially Complex. Complex is the domain that can be understood only after the fact and has no repeatable method. The correct practice for that domain is not a search for method but probe-sense-respond, that is, scattering probes rather than method.
The revisited note docked Frame 7 as “G (the minimal observable) cannot be written.” Cynefin gives an affirmative reason for the same conclusion. That G cannot be written is not a defect but a mistake of type, having demanded a minimal observable in the Complex domain.
The theory of representational change restates this from the analyst’s side. Ohlsson formulated insight as arising not from continuous search but from the restructuring of the problem representation after an impasse10. That the three notes fall into the same regress three times is not because the object is hard. It is because the analyst’s representation is fixated on “binary inversion.” An impasse is escaped by changing the representation. The moment one abandons the inversion family and moves the representation to morphological analysis’s intersection or PSM’s plural boards, the regress ceases to be a regress. When de Bono described lateral thinking as a strategy of deliberately interrupting an existing pattern and moving to a new one11, it was this switching of representation he meant.
The Voids Seen from Outside the Family
Four operations pointed at places the inversion family cannot reach. Let us bring them together.
Morphological analysis pointed at an intersection cell that raises recursive, ecological, and longitudinal dynamics as a single system. TRIZ pointed at the absence of a procedure to operationalize cultural diversity and universal constraints by separating them into layers. PSM pointed at the absence of a method for adjudicating among competing problem frames. Cynefin and representational change dissolved the infinite regress the three notes fell into as a misclassification of type for the Complex domain, and recombined Frame 7 from a deduction into an object of probing.
All four are formulations that a family toppling a single binary structurally cannot produce. And all four remain candidates, not confirmed gaps. In the next section we check their coverage against recent prior work.
Coverage Check: All Three Candidates Were Partial
We do not leave candidates as candidates. We checked the three voids that the operations outside the family pointed to against recent (2024 to 2026) prior work. All three came out the same type. Not untrodden, but a partial coverage in which the residue can be precisely narrowed against a dense body of prior work.
For the intersection cell (recursive and ecological and longitudinal), the three viewpoints were split across separate papers. Recursion was theorized by Messeri and Crockett as a mechanism from the illusion of understanding to a scientific monoculture12; ecology was formulated by Wu et al. as a three-variable coupled system of human cognition, data quality, and model capability13; the longitudinal axis was tracked by Shen and Wang, who followed the AI-generated ratio of peer reviews year by year14. Even Wu et al., closest to the intersection, stay with a simulation of ordinary differential equations rather than a longitudinal study of an actual literature corpus. The residue narrowed to the absence of research that observes the three axes as one closed dynamical system on real data.
For the layer separation of diversity and human rights, the asymmetry was a dense diagnosis and a thin prescription. That LLMs treat human rights as a negotiable principle was demonstrated by Zhou et al., Samway et al., and Javed et al.31516, and the pluralization and steering of diversity has been thickly operationalized by the lineage of pluralistic alignment1718. But no procedure was found that separates, as distinct components of the same frame, a variable layer that moves diversity by reward from an invariant layer that fixes human rights as a hard constraint, and optimizes them at once. Even CuMA’s architectural separation is between cultures, not on the axis of human rights18. The residue narrowed to the absence of an operationalization that translates the diagnosed two-layer trade-off into a prescription by layer separation.
For the adjudication of problem-framing, the map was drawn and only the adjudication was left vacant. The validity of a single frame has been rapidly assembled by importing measurement theory192021; the coexistence of plural frames was mapped by Burden et al. as six paradigms22; and the origin of the difference among frames was pinned down by Jo and Wilson as a difference in implicit theoretical assumptions23. But once the assumptions are made explicit, the standard for adjudicating which frame to adopt is not treated. Jo and Wilson’s Evaluation Card is a tool for disclosing assumptions, not a tool for adjudicating23. The residue narrowed to the absence of an adjudication standard among competing frames.
The three narrowings rather strengthen the fourth note’s claim.
Where the operations the inversion family cannot reach pointed was not a deductive fantasy but exactly the gap in the dense body of research from 2024 to 2026.
That said, fixing the gap still requires a close reading of each cluster’s unreached full text (the corpus’s [primary-source verification needed]).
Insofar as this is not an exhaustive survey of coverage, it is not a declaration of untrodden ground but the residue that could be confirmed at present.
The One Side That Stays Open in the Fourth Note Too
Even having stepped outside the inversion family, the same limit returns to the operations outside it. Which dimensions morphological analysis raises is chosen by the author. Which worldviews PSM lays side by side is also chosen by the author. This very section, which pointed at the absence of a method for adjudicating the plurality of boards, is itself written from a single worldview, the author’s. The inversion family’s limit of “the same author” does not vanish by changing the family.
So the earlier criticism that convergence is weak evidence returns to this fourth note itself. That the four new operations all lean toward the same “absence of institution and adjudication” may also be a fixed point of the author’s vantage. Increasing the problem-setting methods is not the same as increasing the vantages. Even changing the family of tools, the hand that selects that family remained one. A method to pluralize this hand is left to the next round of collection and examination.
Related Notes
- Where Are the Gaps in AI Research? Seven Voids Found by Reframing — The catalog of seven frames by the first generator (inversion of premises)
- AI Research Gaps, Revisited: Sorting Symmetry-Completion from Coverage, and Digging Out an Endogenous Scaling Law — Sorting and the generation of the endogenous scaling law by the second tool (the A–G operational sheet)
- AI Research Gaps, a Third Time: Rereading Them as Rewritings of the Game Board — A rereading by the third generator (rewriting the game board)
- An Academic Map of Methods for Reframing Problems: From Abduction-2 to Problem Structuring — The scholarly map of the methods this note applies (GMA, TRIZ, C-K, co-evolution, PSM, Cynefin, representational change)
- A Genealogy of Practitioner Methods for Reframing Problems: Who Made Them, Traced to Their Origins — The practical lineage of the same methods
References
The methodological sources, and the literature whose existence was verified as evidence for the void each operation pointed to, are given with DOIs/URLs. Bibliographic detail and retrieval confirmation inherit from the confirmation status of the sibling note An Academic Map of Methods for Reframing Problems: From Abduction-2 to Problem Structuring.
Method (the inversion family and the operations outside it)
- Dorst, K. (2011). The Core of ‘Design Thinking’ and Its Application. Design Studies 32(6):521–532. https://doi.org/10.1016/j.destud.2011.07.006
- Zwicky, F. (1969). Discovery, Invention, Research Through the Morphological Approach. Macmillan. https://archive.org/details/discoveryinventi0000zwic
- Ritchey, T. (2013). Wicked problems: modelling social messes with morphological analysis. Acta Morphologica Generalis 2(1). ISSN 2001-2241. (self-published journal, non-peer-reviewed) https://www.swemorph.com/pdf/wp.pdf
- Altshuller, G. S. (1984). Creativity as an Exact Science (A. Williams, Trans.). Gordon & Breach. ISBN 9780677212302. https://searchworks.stanford.edu/view/1223823
- Hatchuel, A. & Weil, B. (2009). C-K design theory: an advanced formulation. Research in Engineering Design 19(4):181–192. https://doi.org/10.1007/s00163-008-0043-4
- Maher, M. L. & Tang, H. H. (2003). Co-evolution as a computational and cognitive model of design. Research in Engineering Design 14(1):47–64. https://doi.org/10.1007/s00163-002-0016-y
- Checkland, P. B. (2000). Soft systems methodology: a thirty year retrospective. Systems Research and Behavioral Science 17(S1):S11–S58. https://doi.org/10.1002/1099-1743(200011)17:1+<::AID-SRES374>3.0.CO;2-O
- Rosenhead, J. & Mingers, J. (Eds.) (2001). Rational Analysis for a Problematic World Revisited (2nd ed.). Wiley. ISBN 9780471495239.
- Kurtz, C. F. & Snowden, D. J. (2003). The new dynamics of strategy: sense-making in a complex and complicated world. IBM Systems Journal 42(3):462–483. https://doi.org/10.1147/sj.423.0462
- Ohlsson, S. (1992). Information-processing explanations of insight and related phenomena. In M. T. Keane & K. J. Gilhooly (Eds.), Advances in the Psychology of Thinking, Vol. 1, 1–44. Harvester Wheatsheaf.
- de Bono, E. (1969). Information processing and new ideas: lateral and vertical thinking. Journal of Creative Behavior 3(3):159–171. https://doi.org/10.1002/j.2162-6057.1969.tb00124.x
- Sandberg, J. & Alvesson, M. (2011). Ways of constructing research questions: gap-spotting or problematization? Organization 18(1):23–44. https://doi.org/10.1177/1350508410372151
Evidence for the voids (on the AI-research side)
- Zhou, Y., Constantinides, M., Quercia, D. (2025). Should LLMs be WEIRD? Exploring WEIRDness and Human Rights in Large Language Models. AIES 2025. https://arxiv.org/abs/2508.19269
- Si, C., Yang, D., Hashimoto, T. (2024). Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers. arXiv:2409.04109 (ICLR 2025). https://arxiv.org/abs/2409.04109
- Shumailov, I. et al. (2024). AI models collapse when trained on recursively generated data. Nature 631(8022):755–759. https://doi.org/10.1038/s41586-024-07566-y
Recent work referenced in the coverage check (2024–2026; all rows are in the corpus above)
- Messeri, L. & Crockett, M. J. (2024). Artificial intelligence and illusions of understanding in scientific research. Nature 627(8002):49–58. https://doi.org/10.1038/s41586-024-07146-0
- Wu, X. et al. (2026). Human-AI Co-Evolution and Epistemic Collapse: A Dynamical Systems Perspective. arXiv:2605.06347. https://arxiv.org/abs/2605.06347
- Shen, S. & Wang, K. (2026). Detecting AI-Generated Content in Academic Peer Reviews. arXiv:2602.00319. https://arxiv.org/abs/2602.00319
- Samway, K. et al. (2026). When Do Language Models Endorse Limitations on Human Rights Principles? EACL 2026 Findings. https://arxiv.org/abs/2603.04217
- Javed, R. et al. (2025). Hedging and Non-Affirmation: Quantifying LLM Alignment on Questions of Human Rights. arXiv:2502.19463. https://arxiv.org/abs/2502.19463
- Sorensen, T. et al. (2024). A Roadmap to Pluralistic Alignment. ICML 2024. https://arxiv.org/abs/2402.05070
- Sun, A. et al. (2026). CuMA: Aligning LLMs with Sparse Cultural Values via Demographic-Aware Mixture of Adapters. ACL 2026. https://arxiv.org/abs/2601.04885
- Jacobs, A. Z. & Wallach, H. (2021). Measurement and Fairness. FAccT 2021. https://arxiv.org/abs/1912.05511
- Wallach, H. et al. (2025). Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge. ICML 2025. https://proceedings.mlr.press/v267/wallach25a.html
- Bean, A. M. et al. (2025). Measuring what Matters: Construct Validity in LLM Benchmarks. NeurIPS 2025 D&B. https://arxiv.org/abs/2511.04703
- Burden, J. et al. (2025). Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture. IJCAI 2025 Survey. https://arxiv.org/abs/2502.15620
- Jo, N. & Wilson, A. (2026). Position: AI Evaluations Should be Grounded on a Theory of Capability. ICML 2026. https://arxiv.org/abs/2509.19590
Unverified Items ([primary-source verification needed])
The three void candidates were checked against the coverage of recent (2024–2026) prior work and each was judged partial (a residue can be narrowed against a dense body of prior work) (the “coverage check” section). This is not an exhaustive survey of coverage, and the remaining reservations are as follows.
- Intersection cell: the closest, Wu et al. (2026), is an ODE simulation, not a longitudinal study of an actual literature corpus. No research observing the three axes as a closed dynamical system on real data could be confirmed, but the search was not exhaustive.
- Layer separation: in part of the pluralistic-alignment lineage (VALUEFLOW, VISPA, HiVaP), the full-text detail of the relationship between the cultural-diversity layer and the universal-constraint layer is unconfirmed
[primary-source verification needed]. - Adjudication: the formal ACM DL version of Jacobs & Wallach (2021) returns 403 and is substituted by the arXiv version. The original Messick (1989) was confirmed only indirectly.
- All rows and provenance of the collected corpus are in
source/review/ai-research-gaps-method-pluralism/papers.md. - Bibliographic detail of the method sources inherits from the unverified items of An Academic Map of Methods for Reframing Problems: From Abduction-2 to Problem Structuring (the ISBN of Ohlsson’s host volume, the non-peer-reviewed status of Ritchey, the reprint year of Altshuller).
Footnotes
-
Zwicky, F. (1969). Discovery, Invention, Research Through the Morphological Approach. Macmillan. https://archive.org/details/discoveryinventi0000zwic ↩
-
Ritchey, T. (2013). Wicked problems: modelling social messes with morphological analysis. Acta Morphologica Generalis 2(1). ISSN 2001-2241. (self-published journal, non-peer-reviewed) https://www.swemorph.com/pdf/wp.pdf ↩
-
Zhou, K., Constantinides, M., Quercia, D. (2025). Should LLMs Be WEIRD? Exploring WEIRDness and Human Rights in Large Language Models. AIES 2025. https://doi.org/10.1609/aies.v8i3.36761 / https://arxiv.org/abs/2508.19269 ↩ ↩2
-
Si, C., Yang, D., Hashimoto, T. (2024). Can LLMs Generate Novel Research Ideas? arXiv:2409.04109 (ICLR 2025). https://arxiv.org/abs/2409.04109 ↩
-
Altshuller, G. S. (1984). Creativity as an Exact Science (A. Williams, Trans.). Gordon & Breach. ISBN 9780677212302. https://searchworks.stanford.edu/view/1223823 ↩
-
Checkland, P. B. (2000). Soft systems methodology: a thirty year retrospective. Systems Research and Behavioral Science 17(S1):S11–S58. https://doi.org/10.1002/1099-1743(200011)17:1+<::AID-SRES374>3.0.CO;2-O ↩
-
Rosenhead, J. & Mingers, J. (Eds.) (2001). Rational Analysis for a Problematic World Revisited (2nd ed.). Wiley. ISBN 9780471495239. ↩
-
Sandberg, J. & Alvesson, M. (2011). Ways of constructing research questions: gap-spotting or problematization? Organization 18(1):23–44. https://doi.org/10.1177/1350508410372151 ↩
-
Kurtz, C. F. & Snowden, D. J. (2003). The new dynamics of strategy: sense-making in a complex and complicated world. IBM Systems Journal 42(3):462–483. https://doi.org/10.1147/sj.423.0462 ↩
-
Ohlsson, S. (1992). Information-processing explanations of insight and related phenomena. In M. T. Keane & K. J. Gilhooly (Eds.), Advances in the Psychology of Thinking, Vol. 1, 1–44. Harvester Wheatsheaf. ↩
-
de Bono, E. (1969). Information processing and new ideas: lateral and vertical thinking. Journal of Creative Behavior 3(3):159–171. https://doi.org/10.1002/j.2162-6057.1969.tb00124.x ↩
-
Messeri, L. & Crockett, M. J. (2024). Artificial intelligence and illusions of understanding in scientific research. Nature 627(8002):49–58. https://doi.org/10.1038/s41586-024-07146-0 ↩
-
Wu, X. et al. (2026). Human-AI Co-Evolution and Epistemic Collapse: A Dynamical Systems Perspective. arXiv:2605.06347. https://arxiv.org/abs/2605.06347 ↩
-
Shen, S. & Wang, K. (2026). Detecting AI-Generated Content in Academic Peer Reviews. arXiv:2602.00319. https://arxiv.org/abs/2602.00319 ↩
-
Samway, K. et al. (2026). When Do Language Models Endorse Limitations on Human Rights Principles? EACL 2026 Findings. https://arxiv.org/abs/2603.04217 ↩
-
Javed, R. et al. (2025). Hedging and Non-Affirmation: Quantifying LLM Alignment on Questions of Human Rights. arXiv:2502.19463. https://arxiv.org/abs/2502.19463 ↩
-
Sorensen, T. et al. (2024). A Roadmap to Pluralistic Alignment. ICML 2024. https://arxiv.org/abs/2402.05070 ↩
-
Sun, A. et al. (2026). CuMA: Aligning LLMs with Sparse Cultural Values via Demographic-Aware Mixture of Adapters. ACL 2026. https://arxiv.org/abs/2601.04885 ↩ ↩2
-
Jacobs, A. Z. & Wallach, H. (2021). Measurement and Fairness. FAccT 2021. https://arxiv.org/abs/1912.05511 ↩
-
Wallach, H. et al. (2025). Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge. ICML 2025. https://proceedings.mlr.press/v267/wallach25a.html ↩
-
Bean, A. M. et al. (2025). Measuring what Matters: Construct Validity in LLM Benchmarks. NeurIPS 2025 D&B. https://arxiv.org/abs/2511.04703 ↩
-
Burden, J. et al. (2025). Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture. IJCAI 2025 Survey. https://arxiv.org/abs/2502.15620 ↩
-
Jo, N. & Wilson, A. (2026). Position: AI Evaluations Should be Grounded on a Theory of Capability. ICML 2026. https://arxiv.org/abs/2509.19590 ↩ ↩2
Author: Shuichiro Ogawa (Design Researcher / Consultant) About me →