Notes · updated 2026-07-20
AI Research Gaps, Revisited: Sorting Symmetry-Completion from Coverage, and Digging Out an Endogenous Scaling Law
A revisit of the seven frames catalogued earlier, using a sharpened discipline (a procedure that lowers a premise, inverts a symmetry, and operationalizes it across an A-through-G worksheet). The same discipline is turned on its own past product to select, strengthen, and generate blind spots.
Contents (9)
- Turning a Sharpened Discipline on the Map You Drew Yourself
- First, List the Common Sense to Be Inverted (Worksheet A)
- The Single Strongest Candidate Whose Every Slot Fills (Group 1)
- The Four That Survive as Symmetry-Completion Frames (Group 2)
- The Two That Fail the Symmetry Test and Are Demoted (Group 3)
- Generating a New Blind Spot: A Scaling Law’s Open-Loop Assumption (Filling One Worksheet Through)
- Where the Novelty Is, and Where It Is Not
- The Limit of Turning the Discipline on Oneself
- Footnotes
Turning a Sharpened Discipline on the Map You Drew Yourself
When you have sharpened a method, the first thing to do is turn it on your own past product.
The sister note Where Are the Gaps in AI Research? Seven Voids Found by Reframing searched for unexplored territory in AI research through seven frames that invert a premise. That catalogue was the first, coarse application of de Broglie’s discovery heuristic of completing a symmetry. Invert the asymmetry admitted on only one side, and read the gap on the unclosed opposite side. The move was single, but the seven frames were uneven in strength and in kind.
So this note lowers that heuristic into a worksheet and revisits the earlier catalogue under its discipline. The worksheet demands that a candidate blind spot be written out fully across seven slots, A through G. A is the common-sense premise to be inverted; B is the swap of whether the model or the measurement is taken as subject; C is the pair that anchors a cross-domain structural correspondence (like particle and wave); D is whether that correspondence can be written one-to-one across categories; E is the empirical condition the inversion predicts; F is the test that would falsify it; G is the minimal observable that exhibits non-classicality or a critical point. The more slots a candidate fills, the stronger it stands as a symmetry-completion blind spot. If a slot stays empty, the candidate is demoted to a different kind of concern.
Under the discipline, the catalogue splits into three. The single strongest candidate whose every slot fills, the few that survive as symmetry-completion frames, and the few that fail the symmetry test and are demoted. Demotion is not disqualification. It is a determination of kind: a legitimate fairness-and-coverage concern, but not the kind of blind spot that inverts a symmetry and closes it. On that basis, the discipline then generates one new blind spot.
This revisit is also a hypothesis, not a confirmation. Both the selection and the generation are judgments within the range the author’s vantage could illuminate.
First, List the Common Sense to Be Inverted (Worksheet A)
The first slot of the discipline was to make explicit the common sense to be inverted. List, without questioning, the premises AI research wears together with its measuring apparatus.
- A model has a capability, and evaluation merely measures it from outside.
- Data gives rise to a model in one direction; that a model produces data in return may be ignored.
- The researcher stands outside the system, a neutral observer of the AI.
- A snapshot at one time suffices for evaluation.
- A model can be evaluated as an individual, in isolation, with one user and one task.
- The field publishes successes and SOTA; failure and non-knowledge are secondary.
- The user present at evaluation is a literate, standard adult.
- Knowledge is assembled around the Anglophone West, and the concepts current there are universal.
- Learning (training) and inference (testing) are separate processes.
- A scaling law is a monotone map from inputs (data and parameters) to performance over an exogenously fixed distribution.
There are ten. Each is inscribed in the apparatus’s design spec, and hard to question from inside the apparatus. The discipline goes on to fill B through G for each item. How they fill decides the kind of blind spot.
The Single Strongest Candidate Whose Every Slot Fills (Group 1)
The earlier catalogue had exactly one candidate whose worksheet fills completely.
It is the relational ontology of capability (developed in Is Capability Inside the Model? A Relational Ontology of AI Evaluation and a Test for Contextuality). A, the premise to invert, is “a model has a capability and evaluation merely measures it.” B, the swap, is the inversion of the asymmetry that sorts a model’s weights and the measurement context into internal and external. Since both are inputs that condition the output, there is no ground to privilege the weights alone as the seat of capability. C, the anchoring pair, is the classical particle and the quantum wave, and, correspondingly, capability and eliciting context. D, the one-to-one correspondence across categories, is writable as the structural correspondence between classical and quantum measurement. Classical measurement assumes a fixed true value for each object; quantum measurement cannot reconcile correlations across incompatible contexts with any single value assignment. Mapping capability measurement onto this structure, the contrast between an evaluation view that assumes a true value and one in which context co-constitutes capability corresponds one-to-one. E, the empirical condition, is the prediction that capability assignments turn non-classical across incompatible eliciting contexts (prompt design, order, scaffolding). F, the falsifying test, is a test for non-classical contextuality in capability measurement. G, the minimal observable, is the degree of non-classicality across incompatible contexts, the extent of structure that no single definite capability assignment can reproduce.
The candidate whose A through G all fill was, in the earlier catalogue, this one alone. [evidence: author’s conjecture] That is why this single candidate became the object of a standalone deep dive in the continuation note. This note’s job is to sort the rest against that standard.
The Four That Survive as Symmetry-Completion Frames (Group 2)
Four candidates do not fill every slot, but their B and C inversions are clear and they stand as symmetry-completion frames. Each is brief, giving only the reason it survives.
First, the time horizon. A is “a snapshot suffices for evaluation.” Inverted, AI takes up residence with people and institutions over months to years, and each remakes the other. F, the test, becomes the measurement of longitudinal drift in capability itself. Model collapse, in which recursive reflux of outputs into future training erases the distribution’s tails and irreversibly degrades diversity, has already been measured as one system-level form of this drift1. The time horizon survives as a frame that inverts the asymmetry of static versus persistent. [evidence: circumstantial]
Second, the reflexivity of research. A is “the researcher is an observer outside the system.” Inverted, the instrument of observation (the LLM that now handles ideation, review, and refereeing) is made of the same material as the object observed. It survives as a frame that inverts the asymmetry of observer and observed. [evidence: author’s conjecture]
Third, AI as an ecology. A is “a model can be evaluated as an individual.” Inverted, many models, people, and institutions form one system. It is a frame that inverts the asymmetry of individual and ecosystem. [evidence: author’s conjecture]
Fourth, the duality of learning and inference. A is “training and testing are separate processes.” The footing for the inversion already exists. Results across several studies show in-context learning to be mechanistically equivalent to gradient descent within the forward pass2. Even with the mechanism the same, the claim that the seam between train and test is ontologically the same thing is distinct, and that step is not yet taken. The inversion that the train/test seam is a frame-dependent artifact stands as a symmetry-completion frame. [evidence: author’s conjecture]
The Two That Fail the Symmetry Test and Are Demoted (Group 3)
This is the largest addition of the revisit. Two of the candidates the earlier catalogue admitted fail at slot D. They are valid concerns, but not symmetry-completion blind spots. They are honestly demoted, and the reason recorded.
The first is cognitively diverse users and non-Western epistemologies. The earlier note set this up as a frame inverting the encoding of a default user. But when one tries to write D (the one-to-one structural correspondence across categories), the correspondence will not be written. In the relational ontology, the particle’s property was “displaced” into the wave’s property, as the internal capability was displaced into a relation. It was a structural correspondence in which the property itself changes place. The concern of cognitive diversity and non-Western epistemology has a different shape. Widening an evaluation built around the standard adult and the Anglophone West toward non-standard cognition and non-Western epistemologies is not an operation that displaces the seat of capability. It is an operation that widens the default sample or coverage. Because it is an extension of coverage rather than a displacement of property, no one-to-one structural correspondence like classical to quantum can be written. So it is not a symmetry-completion blind spot. [evidence: author’s conjecture]
This is not a judgment that lowers the concern’s importance. The encoding of a default user is a legitimate fairness problem, demonstrated at scale3, and the engineering operationalization of non-Western ethical frameworks is genuinely thin. What this note fixes is only the kind. These are coverage-and-fairness concerns, not gaps of the kind that inverts a symmetry and closes it.
The second is the self-mapping of blind spots. The earlier note set this up as a frame in which the field reflexively maps its own negative image. But G (the minimal observable that exhibits non-classicality or a critical point) will not be written. The regress that the very instrument mapping the field’s blind spots has its own negative image does not reduce to a minimal falsifying test. What one would observe to say “the mapping of blind spots succeeded” cannot be defined. Since G stays empty, this is demoted from a worksheet blind-spot candidate to a concern of meta-reflection. [evidence: author’s conjecture]
What the two demoted candidates share is that each is a valid problem yet lies outside the reach of the single discovery heuristic of inverting a symmetry. This selection itself is the largest thing the discipline adds to the earlier catalogue.
Generating a New Blind Spot: A Scaling Law’s Open-Loop Assumption (Filling One Worksheet Through)
The discipline serves not only to select among existing candidates but to generate a new blind spot. Take the last of the ten common-sense premises on the worksheet and fill it through.
A, the premise to invert, was this. A scaling law is a monotone map from inputs (data and parameters) to performance over an exogenously fixed distribution. It tacitly places a one-way direction: a model learns from a given data distribution, and the model does not make that distribution in return.
B, the swap, is the operation that reverses cause and effect of this one-way arrow. “A is the effect of B” is inverted into “A is the condition of B.” In the common sense, capability was the effect of data. Inverted, the product of capability becomes the condition of future data. If the text and images a model generates are published and mix into the next generation’s training corpus in a non-negligible fraction, the output of capability rewrites the input distribution. The open arrow from input to performance folds back from performance to input and closes. The exogenously fixed distribution becomes an endogenously movable one.
C (the anchoring pair) and D (the one-to-one correspondence across categories) are borrowed from two adjacent fields. The validity of the borrowing is warranted by Gentner’s structure-mapping theory. Its criterion is that analogy is sound when it maps a system of relations (systematicity), not surface attributes4. The first borrowing is niche construction from evolutionary biology. Odling-Smee, Laland, and Feldman formalized, as an evolutionary process, the mutual constitution of organism and environment, in which an organism modifies its environment and that modification changes the selection pressures on itself and its descendants5. The correspondence of relational systems runs thus. In the original frame, input is exogenous data, output is capability, and the distribution is fixed. After inversion, the condition is capability, the product is data, and the distribution is endogenously movable. As an organism makes back its niche, a model makes back its data environment. The second borrowing is the positive-feedback loop from control engineering. When output returns to input with the same sign, the system departs from a fixed point and diverges or destabilizes. The loop in which the output of capability returns to the data input, and that input conditions the next capability, can be written not as an open-loop monotone map but as a closed-loop system containing positive feedback. Both correspondences map one-to-one at the level of relational systems, not the surface. [evidence: author’s conjecture]
E, the empirical condition, names the condition under which the inversion predicts a break. When AI-generated products occupy a non-negligible fraction of the training distribution and later generations ingest them, the exogeneity assumption breaks. This condition is already observed. Model autophagy disorder (MAD), in which a generative model continually ingesting its own output degrades quality and diversity, has been demonstrated in image generation6.
F, the falsifying test, strikes where an endogenous scaling law diverges from the classical one. A classical scaling law assuming an exogenously fixed distribution predicts that, however the fraction of generated products rises, performance grows monotonically so long as data volume increases. An endogenous scaling law predicts that, once the endogenous variable of the generated-product fraction crosses a threshold, the coefficients change and monotonicity breaks. So if raising the generated-product fraction leaves the scaling-law coefficients and capability unchanged, the endogenous scaling law is refuted.
G, the minimal observable exhibiting the critical point, is the critical point of this divergence. One side’s sign already exists. Recursion that replaces the distribution produces collapse1. The counterexample already exists too. If generated data are accumulated with, rather than substituted for, real data, collapse is avoided and test error has a finite upper bound independent of the number of iterations7. Since replacement collapses and accumulation does not, a critical region lies between the collapsing regime and the non-collapsing one. The minimal observable is the threshold in the generated-product fraction that divides these two regimes, the point at which the scaling coefficient bends. [evidence: circumstantial]
The last of the ten common-sense premises filled from A through G. This is the blind spot the discipline newly generated.
Where the Novelty Is, and Where It Is Not
The more attractive the generated blind spot looks, the more strictly the boundary with prior work must be drawn. Care is especially required here. The intersection of scaling laws and synthetic data is an extremely active area of recent research.
Model collapse is prior work (cited above1). The self-consuming loop and MAD are prior work (cited above6). Avoidance by accumulation is prior work (cited above7). Closest of all is the body of work that rewrote the scaling law itself under synthetic data. Dohmatob et al. formalized that when synthetic data enter the training corpus the scaling law deteriorates and the coefficients change with each generation8. In a follow-up they showed, within the scaling-law paradigm, a strong collapse in which even a tiny synthetic-data fraction produces collapse and increasing data volume does not lift performance9. So “a scaling law parameterized by the synthetic-data fraction” is already written. The niche-construction analogy, too, is scattered through discussions of AI and data-environment co-evolution.
So the novelty is limited strictly to this single point. In Dohmatob et al.’s scaling law, the fraction of synthetic data is a mixing parameter given from outside. The experimenter exogenously sets “how much synthetic to mix,” and measures how the scaling law changes at each setting. The fraction stays on the input side of the open arrow. What remains new is to treat this fraction not as an externally given mixing parameter but as an endogenous state variable produced by the model’s own capability, and to write the loop that folds from capability to data and from data to the next capability as closed. The exogeneity assumption itself is made the object of inversion, and the scaling law’s input distribution is endogenized as a function of capability. This single step, moving the fraction from the experimenter’s dial to the system’s state variable, corresponds to the difference between open loop and closed loop. Only the shift from “a scaling law parameterized by the synthetic-data fraction” to “a closed-loop scaling law in which capability endogenously drives the fraction” is the unfilled side. [evidence: author’s conjecture]
Whether “endogenous scaling law” is already coined was checked. No established term was found, but the term being available and the underlying dynamic being unexplored are two different things. So the novelty is not placed on the presence of a word. Of the underlying dynamic, only the step of inverting exogeneity and closing the loop is taken as new; everything else (collapse, fraction-dependent scaling laws, avoidance by accumulation) is foregrounded as prior work.
This one point, too, is not a confirmed fact. It is a falsifiable hypothesis awaiting a confirming experiment that explicitly drives the generated-product fraction as a system state variable and measures the critical point where the coefficient bends. [evidence: author’s conjecture]
The Limit of Turning the Discipline on Oneself
This revisit added three things. It sorted Group 1 and Group 2 apart from Group 3, it fixed the reason for demotion as a determination of kind, and from the tenth common-sense premise it generated the blind spot of an endogenous scaling law. Selection raises falsifiability; generation adds one new test.
But this move of turning the discipline on one’s own past product also leaves an unclosed side. The very articulation of the worksheet into seven slots, A through G, is a frame the author’s vantage chose. Choose a different articulation, and a candidate dropped into Group 3 might survive while a survivor might drop. The instrument that measures by the discipline cannot measure the discipline’s own articulation.
So one question is left open. Suppose the critical point of the endogenous scaling law is measured and the fraction at which the coefficient bends is found: is that critical point a fixed constant, or a movable point that shifts each time the system remakes its niche? If the niche-construction correspondence holds literally, the threshold itself should move with the product of capability. If the critical point moves, the test that measures it also changes its object as it measures. Where this self-reference stops, or whether it stops, is beyond this note’s reach.
Related Notes
- Where Are the Gaps in AI Research? Seven Voids Found by Reframing — the catalogue of seven frames that this note revisits
- Is Capability Inside the Model? A Relational Ontology of AI Evaluation and a Test for Contextuality — the standalone deep dive, via de Broglie-style symmetry completion, of the Group 1 candidate whose every slot fills
References
Sources of the method, and works whose existence was confirmed as evidence for each claim, are given with DOI/URL.
Items whose bibliography is partly unconfirmed, or whose full text was not reached, are marked [primary verification needed].
Method (Symmetry Completion and Structure Mapping as a Discovery Heuristic)
- de Broglie, L. (1924). Recherches sur la théorie des quanta. Doctoral thesis, University of Paris. (Direct access to the original PDF is
[primary verification needed]; the argument and submission date are consistently confirmed across multiple secondary sources.) - Davisson, C. & Germer, L. H. (1927). The Diffraction of Electrons by a Crystal of Nickel. Physical Review 30(6):705–740. https://doi.org/10.1103/PhysRev.30.705
- Stanford Encyclopedia of Philosophy. Symmetry and Symmetry Breaking. https://plato.stanford.edu/entries/symmetry-breaking/
- Gentner, D. (1983). Structure-Mapping: A Theoretical Framework for Analogy. Cognitive Science 7(2):155–170. https://doi.org/10.1207/s15516709cog0702_3
- Popper, K. (1959). The Logic of Scientific Discovery. Hutchinson. (Falsifiability as the demarcation criterion of science; English translation of the 1934 German Logik der Forschung.)
Selected Existing Candidates (Groups 2 and 3)
- Shumailov, I. et al. (2024). AI models collapse when trained on recursively generated data. Nature 631(8022):755–759. https://doi.org/10.1038/s41586-024-07566-y
- von Oswald, J. et al. (2023). Transformers Learn In-Context by Gradient Descent. arXiv:2212.07677. https://arxiv.org/abs/2212.07677
- Garg, S. et al. (2022). What Can Transformers Learn In-Context? arXiv:2208.01066. https://arxiv.org/abs/2208.01066
- Septiandri, A. A., Constantinides, M., Tahaei, M., Quercia, D. (2023). WEIRD FAccTs. ACM FAccT 2023. https://arxiv.org/abs/2305.06415
The Generated Blind Spot (Endogenous Scaling Law) and Its Boundary with Prior Work
- Odling-Smee, F. J., Laland, K. N., Feldman, M. W. (2003). Niche Construction: The Neglected Process in Evolution. Princeton University Press (Monographs in Population Biology 37). ISBN 9780691044378.
- Alemohammad, S. et al. (2023). Self-Consuming Generative Models Go MAD. arXiv:2307.01850 (ICLR 2024). https://arxiv.org/abs/2307.01850
- Gerstgrasser, M. et al. (2024). Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data. arXiv:2404.01413. https://arxiv.org/abs/2404.01413
- Dohmatob, E., Feng, Y., Yang, P., Charton, F., Kempe, J. (2024). A Tale of Tails: Model Collapse as a Change of Scaling Laws. ICML 2024. arXiv:2402.07043. https://arxiv.org/abs/2402.07043
- Dohmatob, E., Feng, Y., Subramonian, A., Kempe, J. (2024). Strong Model Collapse. arXiv:2410.04840. https://arxiv.org/abs/2410.04840
Items Flagged for Primary Verification ([primary verification needed])
Collected here are items whose bibliography is unconfirmed, or whose full text was not reached.
- de Broglie (1924): direct access to the original thesis PDF (argument and submission date consistently confirmed across multiple secondary sources).
- Gentner (1983): full text (DOI 10.1207/s15516709cog0702_3 and bibliography agree across multiple secondary sources; the publisher’s full text is paywalled and the author-hosted PDF returned 403).
- Odling-Smee et al. (2003): full-text details (bibliography and ISBN agree across multiple secondary sources; the original book text was not reached).
Footnotes
-
Shumailov, I. et al. (2024). AI models collapse when trained on recursively generated data. Nature 631(8022):755–759. https://doi.org/10.1038/s41586-024-07566-y Demonstrates that recursive reflux of outputs erases distribution tails and irreversibly degrades diversity. ↩ ↩2 ↩3
-
von Oswald, J. et al. (2023). Transformers Learn In-Context by Gradient Descent. arXiv:2212.07677. https://arxiv.org/abs/2212.07677 / Garg, S. et al. (2022). What Can Transformers Learn In-Context? arXiv:2208.01066. https://arxiv.org/abs/2208.01066 Arguments that in-context learning is mechanistically equivalent to gradient descent within the forward pass. ↩
-
Septiandri, A. A. et al. (2023). WEIRD FAccTs: How Western, Educated, Industrialized, Rich, and Democratic is FAccT? ACM FAccT 2023. https://arxiv.org/abs/2305.06415 Shows that most human-subjects studies in fairness research rely on Western, and especially US, participants. ↩
-
Gentner, D. (1983). Structure-Mapping: A Theoretical Framework for Analogy. Cognitive Science 7(2):155–170. https://doi.org/10.1207/s15516709cog0702_3 Structure-mapping theory: analogy is sound when it maps a system of relations (systematicity), not surface attributes. ↩
-
Odling-Smee, F. J., Laland, K. N., Feldman, M. W. (2003). Niche Construction: The Neglected Process in Evolution. Princeton University Press (Monographs in Population Biology 37). ISBN 9780691044378. Formalizes as an evolutionary process the mutual constitution in which an organism modifies its environment and alters selection pressures. ↩
-
Alemohammad, S. et al. (2023). Self-Consuming Generative Models Go MAD. arXiv:2307.01850 (ICLR 2024). https://arxiv.org/abs/2307.01850 Demonstrates in image generation that a self-consuming loop re-ingesting output degrades quality or diversity (MAD). ↩ ↩2
-
Gerstgrasser, M. et al. (2024). Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data. arXiv:2404.01413. https://arxiv.org/abs/2404.01413 Proves that accumulating synthetic data alongside real data (rather than replacing) avoids collapse and bounds test error independent of iteration count. ↩ ↩2
-
Dohmatob, E., Feng, Y., Yang, P., Charton, F., Kempe, J. (2024). A Tale of Tails: Model Collapse as a Change of Scaling Laws. ICML 2024. arXiv:2402.07043. https://arxiv.org/abs/2402.07043 Formalizes that synthetic data entering the corpus deteriorates the scaling law and changes coefficients across generations. ↩
-
Dohmatob, E., Feng, Y., Subramonian, A., Kempe, J. (2024). Strong Model Collapse. arXiv:2410.04840. https://arxiv.org/abs/2410.04840 Shows within the scaling-law paradigm that even a tiny synthetic-data fraction produces collapse and that increasing data volume does not lift performance. ↩
Author: Shuichiro Ogawa (Design Researcher / Consultant) About me →