Notes · updated 2026-07-19
If light is a particle, matter must be a wave
Sometimes only the shape of the right question is fixed before any evidence arrives.
In 1924, de Broglie’s doctoral thesis predicted the existence of a phenomenon no one had yet measured. Light had long been believed to be a wave, but the physics of the day already accepted that it also behaves as a particle (the light quantum). Only one side was granted. So, he wrote, matter, long believed to be particulate, must also behave as a wave1. He closed the symmetry with a single equation binding a particle of momentum p to a wavelength λ. Not one experimental confirmation existed yet.
Three years later the prediction was confirmed without being aimed at. Davisson and Germer, in the middle of a different experiment scattering electrons off a nickel crystal, saw the electrons produce the same diffraction pattern as X-rays2. The electrons were interfering as waves. The symmetry, closed by inverting the granted side, was filled in afterward by measurement. This is not an isolated anecdote. Physics holds, as a method, the practice of building theory by taking the breaking of a symmetry as a clue3.
What happens if this move is carried straight into AI capability evaluation? The sister note ai-research-gaps-abduction searched AI’s unexplored territory through seven frames that invert assumptions. This note takes just one of those seven, the blind spot about the ontology of evaluation, and digs deep with de Broglie’s general heuristic of symmetry completion. The method is this. Take one asymmetry that evaluation grants on only one side, and invert it to close the symmetry. The blind spot lies on the unclosed opposite side. An assessment of this heuristic itself is the assessment of the symmetry-inversion method.
The asymmetry evaluation grants on only one side
First, write out precisely the side that is granted.
Current AI evaluation assumes that capability is intrinsic to the model. The model “has” its capability inside the physical substrate of its weights, and a benchmark merely “reveals” it from outside. This structure is isomorphic to the assumption classical physics placed on particles. That assumption: a particle has a position and momentum as pre-fixed values, and measurement merely reads those values off. In evaluation, the model’s weights correspond to intrinsic capability, and the prompt and scaffolding to a procedure added from outside. Weights count as intrinsic, prompt and context as external.
This split is carved deep into the language of practice. When one says “this model’s reasoning ability is X,” X is spoken of as a scalar value proper to the model, independent of how the prompt is written. A leaderboard can rank models in a single column only because it assumes capability is that kind of scalar property. Scaling laws can bind parameter count and capability in one curve only on the same assumption. That is, the evaluation paradigm requires an ontology of intrinsic capability before the measuring apparatus measures anything.
Here lies a side not yet inverted.
Invert it, and the true value vanishes
Invert it.
Where is the principled reason to split weights and prompt into intrinsic and external? Both are inputs that condition the same outcome, the model’s output. Weights are the parameters training wrote; prompt and scaffolding are the context supplied at inference; but in their role of determining the output distribution, the two hold the same standing. If so, there is no ground to privilege the weights alone as the seat of capability while treating context as “noise added to capability from outside.” Scaffolding and context constitute capability with the same standing as the weights.
From this single step follows a consequence that drops naive realism.
There is no such thing as a “true capability” independent of context.
What is denied here is not the proposition “the word capability is meaningless.”
What is denied is the proposition “a model has a single fixed capability value prior to every measurement context.”
Capability becomes the same kind of quantity as speed.
Just as an object’s speed takes a value only once a reference frame is specified, capability takes a value only once a measurement context (prompt design, order, scaffolding, environment) is specified.
Measurement does not reveal a capability that was already there.
Measurement brings capability into being. [evidence: author's conjecture]
This has the same form as de Broglie’s step. It inverts the “a particle has a position and measurement merely reads it” that evaluation grants on one side, and closes the symmetry. What appears on the far side is an ontology of frame-relative capability.
Why this blind spot is structurally invisible
Granting the inversion, why does the mainstream overlook it? The reason is not “researchers lack attention.” It is that the very purpose of the paradigm makes this inversion structurally invisible.
Benchmark culture aims to produce frame-invariant numbers. A leaderboard, to rank different models on one scale, demands numbers stable across measurement contexts. A scaling law, to extrapolate capability as a function of parameter count, demands that capability be a scalar property of the model. Both enterprises build in, as a precondition for holding together, the ontology of intrinsic capability. Frame-relativity denies that precondition. So from inside a paradigm whose purpose is to make frame-invariant numbers, frame-relativity is in principle invisible.
As Kuhn showed, during normal science researchers assume the paradigm is correct and refine details, leaving the framework itself unquestioned4.
A measuring apparatus wears a particular ontology before it measures.
A question that doubts that ontology lies outside the apparatus’s design specification.
This structure is exactly what makes the blind spot a blind spot. [evidence: author's conjecture]
Scattered phenomena bundle under one ontology
Whether an inversion is productive is decided by what it explains. The relational ontology bundles, under one structure, phenomena that have been treated separately.
First, prompt sensitivity. Mizrahi et al.’s large-scale study in TACL 2024 showed, across 20 LLMs and 39 tasks over 6.5 million instances, how brittle single-prompt evaluation is5. Merely paraphrasing the prompt reshuffles the model rankings. The mainstream sees this as noise to be averaged out to approach the true value. Invert it, and rankings changing with context is not noise but the very phenomenon that capability is context-relative.
Second, chain of thought. Wei et al.’s 2022 paper chose the word Elicits for its title6. The implication: chain-of-thought prompting does not teach reasoning anew but makes reasoning latent in the model manifest. The phrase “making the latent manifest” presupposes a latent capability prior to manifestation. But if capability values differ with and without CoT, which is the “true” capability? Invert it, and the question dissolves. CoT is part of context as scaffolding, and capability-with-CoT and capability-without-CoT are simply two values brought into being in two different contexts.
Third, the dispute over what post-training does to capability. Reporting LIMA, Zhou et al. proposed, from the fact that high-quality output follows from fine-tuning on just 1,000 examples, that most knowledge is acquired in pretraining and fine-tuning merely draws out existing knowledge7. Does fine-tuning “unlock” capability, or “grant” it? This dispute stands as a dispute because both camps share the premise that a hidden capability to be unlocked exists in advance. Invert it, and the very distinction between unlocking and granting reduces to whether one counts context as part of capability or as external to it.
Fourth, why contamination is feared. When evaluation data leaks into training, the evaluation community fears the benchmark number will overestimate “true capability.” But for the concept of overestimation to hold, there must be a true capability value as the reference of comparison. Contamination arises as a problem at the very moment a true value is posited. Without a true value, the number raised by leaked data and the number measured on unleaked data are both capability values brought into being in particular contexts, and one cannot be said to “overestimate” the other.
Fifth, jailbreak. A model with identical weights refuses under one prompt and responds under another. The boundary of capable versus incapable moves with prompt, a context. Greenblatt et al.’s 2024 study of password-locked models carved out this structure experimentally8. On models fine-tuned to exhibit a capability only when a specific string is present and to imitate a weaker model otherwise, methods to draw out the hidden capability were tested. This study takes the stance of positing a hidden true capability. In that sense its premise is the reverse of the present hypothesis, but by explicitly placing that reverse premise, it brings into relief where positing (or not positing) a true value takes effect.
The mainstream has handled these five individually as “noise to be removed in order to approach the true value.”
Invert it, and they are not noise but the phenomena themselves.
And the logic that bundles the five reduces to one thing.
Without a true value, the concept of noise does not hold. [evidence: circumstantial]
Build a falsifiable prediction
So far this is a matter of ontology, and as such it is not falsifiable. De Broglie’s step became science because it came with a falsifiable prediction, electron diffraction. The relational ontology, to earn the same standing, must predict something that in principle cannot occur under classical realism.
That prediction is non-classical contextuality.
Classical measurement theory holds that each object has a fixed true value and measurement merely reads it off with error added. Under this picture, whatever is measured in whatever context can be made consistent by a single underlying true-value assignment. In quantum mechanics, however, there arise correlations across mutually incompatible measurement contexts that no single deterministic value assignment can reproduce. This is Kochen–Specker and Bell-type contextuality, which Abramsky and Brandenburger formulated uniformly in the language of sheaves9.
If capability is relational, capability measurement should show the same structure.
When capability is measured across mutually incompatible eliciting contexts (prompt design, presentation order, scaffolding), there appears a structure that no single deterministic capability assignment can reproduce all at once.
Under the classical picture of “true value plus measurement error,” this structure cannot in principle arise.
So if non-classical contextuality is detected in capability measurement, the classical view of evaluation, which presupposes a true value, is falsified.
This is the test corresponding to electron diffraction. [evidence: author's conjecture]
The ground that this prediction is not fantasy lies in the fact that contextuality has already been demonstrated in a non-quantum domain. Wang, Solloway, Shiffrin, and Busemeyer showed in PNAS 2014 that the effect of question order changing responses follows the non-parametric prediction of a quantum probability model across 70 national surveys10. Human judgment shows non-classical order effects across incompatible contexts. The field of quantum cognition has treated this kind of effect systematically11. The prospect that the same test can be assembled for capability comes from here.
The one point where the metaphor becomes literal
Here is where this note grounds most firmly.
The apparatus of sheaf-theoretic contextuality has already been applied to LLMs, not as a metaphor but as a literal computation. Lo, Sadrzadeh, and Mansfield, in Proceedings of the Royal Society A in 2025, detected sheaf-contextuality at scale on probability distributions extracted from BERT’s embeddings12. Using Simple English Wikipedia, they measured 77,118 sheaf-contextual and about 36.9 million CbD-contextual instances, and showed that the degree of contextuality is best predicted by the Euclidean distance of the embedding vectors. That non-classical contextuality exists in LLMs is no longer a hypothesis.
But do not mistake the target. What Lo et al. measured is the contextuality of word meaning, not capability or evaluation. They showed how a polysemous word like “cabinet” settles into which sense in context, and that this overlap of senses has a non-classical structure, not that a model’s reasoning capability is brought into being in context.
This distinction reproduces de Broglie’s structure exactly.
Non-classicality has already been measured in one domain, meaning.
That is the “light quantum.”
What the symmetry announces is the prediction that it should therefore be shown on the capability side too.
That is the “matter wave.”
The apparatus is complete.
Only the design that moves the target from meaning to capability remains as the unfilled side. [evidence: circumstantial]
Where the novelty is, and where it is not
The more attractive an inversion looks, the more strictly one must draw the boundary with what is already known. Much of this note’s claim already exists in part.
The critique that treats capability as a measurement-dependent construct already exists. Jacobs and Wallach in 2021 offered a measurement-theory framework that captures concepts like fairness as products of measurement13. Bean et al. in 2025 analyzed 445 benchmarks with 29 experts and showed that a lack of construct validity undermines the resulting claims of evaluation14. These sharply point out that measurement shapes the image of capability. But neither takes the ontological step to the non-existence of a true value; both stay within the frame of approaching a true construct through better measurement.
The position that treats capability as a context-relative disposition also already exists. The propensity-measurement framework Romero-Alvarado et al. published in 2026 measures behavioral tendencies beyond capability across contexts15. But this framework stands on item response theory and treats propensity as a measurable construct. It does not explicitly deny the ontology of a true value. So the step of dropping classical realism is not contained here either.
The empirical measurement of non-classical contextuality in LLMs also already exists (see 12). But the target is meaning, not capability.
The claim that emergence is a mirage also already exists. Schaeffer et al. in 2023 showed that the “emergence” of capability can be a product of the non-linearity of the measurement metric16. But this argument presupposes a continuous true capability underneath, and argues that the continuous quantity merely appears to emerge through a discrete metric. In presupposing a true value, it faces the opposite way from the present hypothesis.
Capability-elicitation research also already exists (see 8). But it takes the stance of presupposing a hidden true value, the reverse premise of the present hypothesis.
Given all this, the novelty of this note is strictly limited to three points.
First, the ontological inversion. Where much of the existing measurement-dependent critique keeps the frame of approaching a true construct through better measurement, this note takes the step to the non-existence of a true value and drops realism itself. Second, the transfer of the test. It designs a move of the empirical measurement of non-classical contextuality from the domain of meaning to the domain of capability and evaluation. Third, the bundling. It bundles the scattered phenomena of prompt sensitivity, CoT, the unlock dispute over post-training, contamination, and jailbreak under one ontology, the non-existence of a true value.
And these three points are not confirmed facts.
They are a falsifiable hypothesis awaiting a confirming experiment, the capability contextuality test. [evidence: author's conjecture]
Two further asymmetries the same heuristic illuminates
The same move illuminates two more asymmetries outside evaluation. Both are touched on only briefly.
One is the asymmetry between learning and inference.
It is customary to split training as the process that writes in capability and inference as the process that uses it.
But results in several studies show that in-context learning is mechanistically equivalent to gradient descent within the forward pass17.
von Oswald et al. showed constructively that a trained transformer becomes a mesa-optimizer that learns a model by gradient descent within its forward pass.
Given that mechanistic equivalence is already established, completing the symmetry leaves an upgrade to an ontological identity: that the seam between train and test is a frame-dependent artifact.
The claim that the mechanism is the same, and the claim that they are ontologically the same thing, are distinct claims.
The latter has not yet been taken. [evidence: author's conjecture]
The other is the asymmetry of alignment.
The one direction, aligning AI to humans, has become the default.
But humans too change their cognition and behavior to fit AI.
Shen et al. in 2024 offered a bidirectional classification of the direction of aligning AI to humans and the direction of aligning humans to AI18.
That it is bidirectional has already been classified in this way.
Even so, that which side is the one being aligned is frame-dependent, and asking where the coupled system of humans and AI settles as a fixed point, have not yet been taken.
Li and Song’s 2025 framework of mutual adaptation comes closest, but it too places the optimal point of collaboration at the intersection of human and AI capability, and does not stand on an ontology that treats which side is the aligning subject as frame-relative19. [evidence: author's conjecture]
A relational ontology sees AI not as a bundle of individual properties but as a relation brought into being within a system. It resonates from a distance with the idea of affordance, which formulated perception within the relation to the environment20. The map of psychology and cognitive-science currents ai-psychology-cognitive-science-trends has tracked adjacent areas of quantum cognition and capability measurement. What that map presupposed is the structure in which the capability to be measured exists before measurement. The inversion, that measurement brings capability into being, remains placed on the far side of that structure.
Leaving one question open
De Broglie’s symmetry closed with electron diffraction. This note’s symmetry has not yet closed.
Non-classicality has been shown on the meaning side, and the apparatus is running. It should be shown on the capability side too, the symmetry announces. But announcing and showing are different. How an experiment testing capability contextuality can be assembled, how mutually incompatible eliciting contexts are to be defined, and how the structure irreproducible by any single assignment is to be detected, lie beyond this note’s reach. If confirmed, the view of evaluation that presupposes a true value is falsified. If not falsified, the classical picture, that capability after all has a context-independent true value, survives.
Which way it falls has not yet been measured. It is in that single fact of being unmeasured that this hypothesis earns its standing as a hypothesis.
References
Sources of the method, and works whose existence was confirmed as evidence for each claim, are given with DOI/URL.
Items whose bibliography is unsettled or whose body text was not reached are marked [requires primary verification].
Method (symmetry completion as heuristic)
- de Broglie, L. (1924). Recherches sur la théorie des quanta. PhD thesis, University of Paris. (Direct access to the original PDF is
[requires primary verification]; the structure and defense date are confirmed across secondary sources.) - Davisson, C. & Germer, L. H. (1927). The Diffraction of Electrons by a Crystal of Nickel. Physical Review 30(6):705–740. https://doi.org/10.1103/PhysRev.30.705
- Stanford Encyclopedia of Philosophy. Symmetry and Symmetry Breaking. https://plato.stanford.edu/entries/symmetry-breaking/
- Kuhn, T. S. (1962). The Structure of Scientific Revolutions. University of Chicago Press.
Scattered phenomena the inversion bundles (ontology of evaluation)
- Mizrahi, M. et al. (2024). State of What Art? A Call for Multi-Prompt LLM Evaluation. TACL. https://arxiv.org/abs/2401.00595
- Wei, J. et al. (2022). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. NeurIPS 2022. https://arxiv.org/abs/2201.11903
- Zhou, C. et al. (2023). LIMA: Less Is More for Alignment. arXiv:2305.11206. https://arxiv.org/abs/2305.11206
- Greenblatt, R., Roger, F., Krasheninnikov, D., Krueger, D. (2024). Stress-Testing Capability Elicitation With Password-Locked Models. arXiv:2405.19550. https://arxiv.org/abs/2405.19550
Confirming prediction (the contextuality test)
- Abramsky, S. & Brandenburger, A. (2011). The Sheaf-Theoretic Structure of Non-Locality and Contextuality. New Journal of Physics 13:113036. https://arxiv.org/abs/1102.0264
- Wang, Z., Solloway, T., Shiffrin, R. M., Busemeyer, J. R. (2014). Context effects produced by question orders reveal quantum nature of human judgments. PNAS 111(26):9431–9436. https://doi.org/10.1073/pnas.1407756111
- Busemeyer, J. R. & Bruza, P. D. (2012). Quantum Models of Cognition and Decision. Cambridge University Press.
- Lo, K. I., Sadrzadeh, M., Mansfield, S. (2025). Quantum-like contextuality in large language models. Proceedings of the Royal Society A 481(2319):20240399. https://doi.org/10.1098/rspa.2024.0399 (arXiv:2412.16806)
Limiting the novelty (boundary with prior work)
- Jacobs, A. Z. & Wallach, H. (2021). Measurement and Fairness. FAccT ‘21. https://doi.org/10.1145/3442188.3445901
- Bean, A. M. et al. (2025). Measuring what Matters: Construct Validity in Large Language Model Benchmarks. arXiv:2511.04703. https://arxiv.org/abs/2511.04703
- Romero-Alvarado, D. et al. (2026). Capabilities Ain’t All You Need: Measuring Propensities in AI. arXiv:2602.18182. https://arxiv.org/abs/2602.18182 (body reached at abstract only,
[requires primary verification]) - Schaeffer, R., Miranda, B., Koyejo, S. (2023). Are Emergent Abilities of Large Language Models a Mirage? NeurIPS 2023. https://arxiv.org/abs/2304.15004
The two further asymmetries
- von Oswald, J. et al. (2023). Transformers Learn In-Context by Gradient Descent. arXiv:2212.07677. https://arxiv.org/abs/2212.07677
- Garg, S. et al. (2022). What Can Transformers Learn In-Context? arXiv:2208.01066. https://arxiv.org/abs/2208.01066
- Akyürek, E. et al. (2023). What learning algorithm is in-context learning? ICLR 2023. https://arxiv.org/abs/2211.15661
- Shen, H. et al. (2024). Position: Towards Bidirectional Human-AI Alignment. arXiv:2406.09264. https://arxiv.org/abs/2406.09264
- Li, Y. & Song, W. (2025). Co-Alignment: Rethinking Alignment as Bidirectional Human-AI Cognitive Adaptation. arXiv:2509.12179. https://arxiv.org/abs/2509.12179 (body reached at abstract only,
[requires primary verification]) - Gibson, J. J. (1979). The Ecological Approach to Visual Perception. Houghton Mifflin.
Items requiring primary verification
Items whose bibliography is unsettled, or whose body text was not reached, are collected here.
- Direct access to the de Broglie (1924) original PDF (structure and defense date consistent across secondary sources).
- The body details of Romero-Alvarado et al. (2026) (abstract only reached; whether propensity ontologically denies a true value is unconfirmed).
- The body details of Li & Song (2025) (abstract only reached; whether it refers to a fixed point of the coupled system is unconfirmed).
Footnotes
-
de Broglie, L. (1924). Recherches sur la théorie des quanta. PhD thesis, University of Paris, defended 25 November 1924. Extended the wave-particle duality of light to matter, giving the relation between momentum and wavelength. Direct access to the original PDF is
[requires primary verification](the structure and defense date are consistent across secondary sources). ↩ -
Davisson, C. & Germer, L. H. (1927). The Diffraction of Electrons by a Crystal of Nickel. Physical Review 30(6), 705–740. https://doi.org/10.1103/PhysRev.30.705 Demonstrated electron diffraction, confirming (initially unintentionally) de Broglie’s matter-wave hypothesis. ↩
-
Stanford Encyclopedia of Philosophy, “Symmetry and Symmetry Breaking”. https://plato.stanford.edu/entries/symmetry-breaking/ An overview of symmetry and its breaking as a heuristic for theory construction. ↩
-
Kuhn, T. S. (1962). The Structure of Scientific Revolutions. University of Chicago Press. That normal science takes the paradigm as given and refines details without questioning the framework itself. ↩
-
Mizrahi, M. et al. (2024). State of What Art? A Call for Multi-Prompt LLM Evaluation. TACL. https://arxiv.org/abs/2401.00595 Demonstrated the brittleness of single-prompt evaluation across 6.5M instances, 20 LLMs, 39 tasks. ↩
-
Wei, J. et al. (2022). Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. NeurIPS 2022. https://arxiv.org/abs/2201.11903 The title’s “Elicits” (making the latent manifest) presupposes a latent capability prior to manifestation. ↩
-
Zhou, C. et al. (2023). LIMA: Less Is More for Alignment. arXiv:2305.11206. https://arxiv.org/abs/2305.11206 From high-quality output after 1,000-example fine-tuning, proposes the Superficial Alignment Hypothesis that most knowledge is from pretraining and fine-tuning merely draws out existing knowledge. ↩
-
Greenblatt, R., Roger, F., Krasheninnikov, D., Krueger, D. (2024). Stress-Testing Capability Elicitation With Password-Locked Models. arXiv:2405.19550. https://arxiv.org/abs/2405.19550 On models fine-tuned to show capability only with a specific string, tested methods to draw out the hidden capability. Takes the stance of positing a hidden true value. ↩ ↩2
-
Abramsky, S. & Brandenburger, A. (2011). The Sheaf-Theoretic Structure of Non-Locality and Contextuality. New Journal of Physics 13:113036. https://arxiv.org/abs/1102.0264 A uniform sheaf-theoretic formulation of non-locality and contextuality. ↩
-
Wang, Z., Solloway, T., Shiffrin, R. M., Busemeyer, J. R. (2014). Context effects produced by question orders reveal quantum nature of human judgments. PNAS 111(26):9431–9436. https://doi.org/10.1073/pnas.1407756111 Showed question-order effects follow a quantum probability model’s non-parametric prediction across 70 national surveys. ↩
-
Busemeyer, J. R. & Bruza, P. D. (2012). Quantum Models of Cognition and Decision. Cambridge University Press. A systematic formulation of non-classical effects in judgment and decision. ↩
-
Lo, K. I., Sadrzadeh, M., Mansfield, S. (2025). Quantum-like contextuality in large language models. Proceedings of the Royal Society A 481(2319):20240399. https://doi.org/10.1098/rspa.2024.0399 (arXiv:2412.16806, December 2024). Measured 77,118 sheaf-contextual and about 36.9M CbD-contextual instances on BERT word meaning. ↩ ↩2
-
Jacobs, A. Z. & Wallach, H. (2021). Measurement and Fairness. FAccT ‘21. https://doi.org/10.1145/3442188.3445901 (arXiv:1912.05511). A measurement-theory framework capturing fairness etc. as products of measurement. ↩
-
Bean, A. M. et al. (2025). Measuring what Matters: Construct Validity in Large Language Model Benchmarks. arXiv:2511.04703. https://arxiv.org/abs/2511.04703 Analyzed 445 benchmarks with 29 reviewers and identified the lack of construct validity. ↩
-
Romero-Alvarado, D. et al. (including Hernández-Orallo) (2026). Capabilities Ain’t All You Need: Measuring Propensities in AI. arXiv:2602.18182. https://arxiv.org/abs/2602.18182 Measures propensity across contexts within an item-response-theory framework. Does not make an explicit ontological denial of a true value.
[requires primary verification](access limited to the abstract). ↩ -
Schaeffer, R., Miranda, B., Koyejo, S. (2023). Are Emergent Abilities of Large Language Models a Mirage? NeurIPS 2023. https://arxiv.org/abs/2304.15004 Argues emergence can be a product of the non-linearity of the measurement metric. Presupposes a continuous true capability underneath. ↩
-
von Oswald, J. et al. (2023). Transformers Learn In-Context by Gradient Descent. arXiv:2212.07677. https://arxiv.org/abs/2212.07677 / Garg, S. et al. (2022). What Can Transformers Learn In-Context? arXiv:2208.01066. https://arxiv.org/abs/2208.01066 / Akyürek, E. et al. (2023). What learning algorithm is in-context learning? ICLR 2023. https://arxiv.org/abs/2211.15661 Three lines of argument that in-context learning is mechanistically equivalent to gradient descent within the forward pass. ↩
-
Shen, H. et al. (2024). Position: Towards Bidirectional Human-AI Alignment. arXiv:2406.09264. https://arxiv.org/abs/2406.09264 Offers a bidirectional classification of aligning AI to humans and humans to AI. ↩
-
Li, Y. & Song, W. (2025). Co-Alignment: Rethinking Alignment as Bidirectional Human-AI Cognitive Adaptation. arXiv:2509.12179. https://arxiv.org/abs/2509.12179 Offers a framework of mutual adaptation and places optimal collaboration at the intersection of human and AI capability.
[requires primary verification](access limited to the abstract). ↩ -
Gibson, J. J. (1979). The Ecological Approach to Visual Perception. Houghton Mifflin. Formulated affordance as a relation between organism and environment. ↩