Shuichiro Ogawa
日本語

Notes · updated 2026-07-19

AI, Psychology, and Cognitive Science Research Trends: Machine Psychology and the Cognitive Modeling of LLMs (2024–2026)

An integrative summary of research trends that emerged at the intersection of large language models (LLMs) and psychology/cognitive science in 2024–2026, collected through a lightweight scoping review of 30 peer-reviewed papers and major preprints. Full bibliographic details for each publication appear in the “References” section below (with DOIs/URLs for traceability). The internal working ledger with provenance tracking, confidence ratings, and reachability-verification methods is at source/review/ai-psychology-cognitive-science-trends/papers.md (repository-internal, not published). For adjacent research trends, see ai-social-science-research-trends and ai-education-learning-sciences-trends; for the industry-vs-academy skill-gap debate, see ai-cognition-skill-gap-debate. This note keeps a separate scope: the research trends themselves in psychology and cognitive science. Collection protocol: .claude/rules/collection-protocol.md (peer-review priority, zero fabrication, provenance tracking).

The same model can be read, by one study, as showing six-year-old-level theory of mind, and by another as failing the moment a single word in the task is changed. Between 2023 and 2024, LLM research on theory of mind produced both images at once (Kosinski 2024, PNAS; Ullman 2023; Strachan et al. 2024, Nature Human Behaviour). The central question of this field became not which reading is correct, but why the same object yields such divergent pictures.

The question originates in a simple move: administering psychology’s tasks to machines rather than humans. Hagendorff and colleagues named this machine psychology in 2023 and proposed it as a methodological program for studying emergent LLM behavior across the subfields of psychology (Hagendorff et al. 2023/2024, arXiv). At almost the same time, Binz and Schulz showed in a 2023 PNAS paper that GPT-3, given classic decision-making and information-search tasks, reached human-level performance on standard vignettes but broke down when small, logically equivalent alterations were introduced (Binz & Schulz 2023, PNAS). This pattern of “passing the standard task, fragile to perturbation” prefigures the shape of the debates that followed.

This note tracks the research at this intersection across four clusters for 2024–2026: machine psychology and its methodological controversies; LLMs as models or theories of human cognition, and the critiques thereof; the psychology of human-AI interaction; and the turn in cognitive-science methodology itself. Every claim is tied to a source, and preprints are flagged in-text as [preprint].

Machine Psychology: Passing the Task, Fragile to Perturbation

The empirical program of machine psychology split most sharply over theory of mind. Kosinski reported in a 2024 PNAS paper that GPT-4-class models scored at roughly a six-year-old’s level on false-belief tasks (Kosinski 2024, PNAS). Against this, Ullman showed in 2023 that swapping the vignettes for logically equivalent variants with altered surface form was enough to make the models fail, and argued that the apparent success could be explained by reliance on surface heuristics (Ullman 2023, arXiv, [preprint]).

The closest thing to a resolution came from large-scale direct comparison with humans. Strachan and colleagues, in a 2024 Nature Human Behaviour study, compared GPT-family and LLaMA2-family models against 1,907 humans across false belief, indirect requests, irony, and faux pas (Strachan et al. 2024, Nature Human Behaviour). GPT-4 reached human level on false belief and indirect requests, while LLaMA2’s edge on some tasks was traced by follow-up manipulations to a bias toward over-attributing ignorance. The point is that “human-level” can mean genuinely solving the task, or it can mean a bias that happens to align with the scoring key.

Beyond theory of mind, the same ambivalence recurred. Suri and colleagues demonstrated in 2024 that GPT-3.5 exhibits anchoring, representativeness, availability, framing, and endowment effects in patterns resembling humans (Suri et al. 2024, JEP: General). That a machine reproduces human biases is evidence that psychological paradigms can be applied to LLMs. It is not, however, evidence that the machine shares the human cognitive processes that produce them.

On the methodological side, wariness about turning psychology’s instruments on LLMs became organized. Abdurahman and colleagues, in a 2024 PNAS Nexus paper, named the practice GPTology, mapped its opportunities and risks, and argued for distinguishing uses of LLMs as “models of the mind” from uses of LLMs as “tools” for text analysis (Abdurahman et al. 2024, PNAS Nexus). Gao and colleagues, in a 2025 PNAS paper, critically examined the use of LLMs as substitutes for human participants and showed that even in simple scenarios their responses were inconsistent with humans and idiosyncratic (Gao et al. 2025, PNAS). Shanahan sounded the alarm from the side of language, analyzing in a 2024 Communications of the ACM paper how ascribing philosophically loaded verbs like “knows” and “believes” to LLMs itself breeds anthropomorphic bias (Shanahan 2024, CACM). The controversy over the very word “understanding” had been framed earlier by Mitchell and Krakauer, who in 2023 proposed a comparative view of intelligence in which there is more than one mode of understanding (Mitchell & Krakauer 2023, PNAS).

LLMs as Cognitive Models: Accelerating Implementation, Contested Legitimacy

Where machine psychology asks whether an LLM behaves like a human, a second trend asks whether an LLM can be made into a model of human cognition itself. The landmark of this direction is Centaur, published in Nature in 2025 (Binz et al. 2025, Nature). Binz and colleagues fine-tuned an LLM on Psych-101, a dataset of roughly 10.6 million choices from over 60,000 participants across 160 experiments, and showed it predicted held-out participants’ behavior more accurately than existing cognitive models. Centaur generalized to task variations and entirely new experimental domains it had not seen in training, and after fine-tuning its internal representations aligned better with human brain activity. This extends the 2024 ICLR proof-of-concept from the same authors, “turning LLMs into cognitive models” (Binz & Schulz 2024, ICLR), into a Nature-scale, data-driven program.

On language, the implications LLMs bring to theory were confronted directly. Mahowald and colleagues, in a 2024 Trends in Cognitive Sciences paper, distinguished formal linguistic competence (knowledge of grammar and patterns) from functional linguistic competence (using and understanding language in the world), and organized within a neuroscience framework the asymmetry whereby LLMs are strong on the former and unstable on the latter (Mahowald et al. 2024, TiCS). On language acquisition, Contreras Kallens and colleagues argued in a 2023 Cognitive Science paper that LLMs can acquire grammatical language without an innate grammar, and that statistical learning can account for much of acquisition (Contreras Kallens et al. 2023, Cognitive Science). On meaning, Piantadosi and Hill made a positive case from conceptual-role semantics, in which meaning is defined by relations among internal states rather than external reference, that LLMs may capture meaning (Piantadosi & Hill 2022, arXiv, [preprint]).

This positive case comes paired with systematic rebuttals from cognitive science and linguistics. Katzir argued in 2023 that LLMs fail fundamentally as theories of human linguistic cognition, citing human learning biases, typological patterns, the competence/performance distinction, and the conflation of probability with grammaticality (Katzir 2023, Biolinguistics). At a more general level, van Rooij and colleagues argued in 2024 that engineering AI has been encroaching on theory-building in cognitive science, and proposed reclaiming AI as a theoretical tool to be scrutinized at the level of computational theory and complexity (van Rooij et al. 2024, Computational Brain & Behavior). Guest and Martin systematized the inferential leap from similarity of performance to identity of cognitive process as an abuse of “the logic of models” (Guest & Martin 2023, Computational Brain & Behavior).

Between rebuttal and positive case, some implementation research makes the model’s limits visible on its own terms. Oh and Schuler showed in 2023 that as language models grow larger and perplexity falls, their surprisal estimates fit human reading times monotonically worse, attributing the cause to memorization from massive training data (Oh & Schuler 2023, TACL). A “smarter model” is not necessarily a “better cognitive model”: a dissociation between scaling and cognitive plausibility. Taking a middle position, McGrath and colleagues argued in 2024 that the high-level representations revealed by interpretability research on deep networks can serve as “informative implementations” for confirming theory and generating hypotheses, bridging the critical and implementational camps (McGrath et al. 2024, Current Directions in Psychological Science).

The Psychology of Human-AI Interaction: Diminished Collaboration and Cognitive Debt

Psychology’s object moved not only inside the machine but into the space between human and machine. The broadest measurement of collaborative effect came from Vaccaro and colleagues, who reported in a 2024 Nature Human Behaviour paper a preregistered meta-analysis of 106 studies and 370 effect sizes (Vaccaro et al. 2024, Nature Human Behaviour). On average, human-AI combinations underperformed the better of human-alone and AI-alone (Hedges’ g = −0.23, 95% CI −0.39 to −0.07). The effect diverged by task, however: losses on decision-making tasks and gains on content-creation tasks. The naive expectation that “two together are stronger” is, on average, betrayed in some domains.

On trust and reliance, the process by which merely knowing something is AI distorts judgment was measured. Klingbeil and colleagues showed in 2024, in an incentivized behavioral experiment, that simply knowing advice came “from AI” produced overreliance strong enough to override contextual information and one’s own judgment, sometimes to the detriment of third parties (Klingbeil et al. 2024, Computers in Human Behavior). Lee and colleagues argued in a 2025 PNAS Nexus paper that the AI’s own metacognitive sensitivity, its ability to estimate its own confidence accurately, is key to human trust calibration and decision accuracy, and pointed to the problem that when AI expresses high confidence, humans raise their trust even when it is wrong (Lee et al. 2025, PNAS Nexus). On anthropomorphism itself, Marchegiani argued in 2025 that the false beliefs produced by anthropomorphizing conversational AI undermine users’ autonomy independently of any instrumental harm (Marchegiani 2025, Journal of Applied Philosophy).

The asymmetry of persuasion was also quantified in this period. Salvi and colleagues reported in a 2025 Nature Human Behaviour paper, using a preregistered debate RCT, that GPT-4 persuaded opponents 64.4% more often than human opponents did, and that personalizing to the opponent’s sociodemographic profile increased the odds by 81.2% (Salvi et al. 2025, Nature Human Behaviour).

The effect of reliance on human cognition was reported repeatedly in 2025. Gerlich showed in 2025, in a mixed-methods study of 666 participants, a significant negative correlation between frequency of AI use and measures of critical thinking, with cognitive offloading as the mediator and the tendency strongest among younger participants (Gerlich 2025, Societies). Lee and colleagues, in a 2025 CHI paper drawing on 936 instances from 319 knowledge workers, documented that high trust in GenAI is associated with reduced effort in critical thinking, and that critical thinking’s center of gravity is shifting from gathering to verifying information, from problem-solving to integrating responses, and from execution to oversight (Lee et al. 2025, CHI). At the neural level, Kosmyna and colleagues showed in 2025, via EEG during an essay-writing task, that the group using an LLM had the weakest brain-network connectivity and lower lexical and structural diversity, terming this accumulation cognitive debt (Kosmyna et al. 2025, arXiv, [preprint]). On longer-term psychosocial effects, Fang and colleagues reported in 2025, in a four-week randomized controlled trial of 981 participants, that self-directed high-frequency use was associated with greater loneliness, dependence, and problematic use (Fang et al. 2025, arXiv, [preprint]).

The Methodological Turn in Cognitive Science: Tool, Subject, Model, or Counterexample

The fourth trend questions not individual findings but the methodology of the science of mind itself. The anchoring framework is machine behaviour, proposed by Rahwan and colleagues in Nature in 2019, which established a field for studying machine behavior across behavioral science, social science, and computer science (Rahwan et al. 2019, Nature). In 2024, this framework developed into an epistemic risk argument in the wake of LLMs. Messeri and Crockett, in a 2024 Nature paper, classified the mechanisms by which AI produces illusions of understanding in science into four archetypes: AI as Oracle, as Surrogate, as Quant, and as Dreamer, and warned of a scientific monoculture in which reliance on AI thins the diversity of research (Messeri & Crockett 2024, Nature).

Whether LLMs can replace human participants became the center of this period’s methodological debate. The precursor is silicon sampling, shown by Argyle and colleagues in 2023: conditioning an LLM on a demographic backstory reproduces the distribution of actual survey responses, a match they named algorithmic fidelity (Argyle et al. 2023, Political Analysis). On the psychology side, Dillion and colleagues, in a 2023 Trends in Cognitive Sciences paper, offered a conditional model of when LLMs can substitute for human participants (Dillion et al. 2023, TiCS). For social science broadly, Bail, in a 2024 PNAS paper, weighed the potential of generative AI to improve surveys, online experiments, content analysis, and agent-based models against risks of bias and reproducibility (Bail 2024, PNAS).

The feasibility of substitution was pushed forward empirically by generative simulation of individuals. Park and colleagues reported in 2024 that LLM agents grounded in interviews with 1,052 people matched the individuals’ own responses two weeks later at 86%, and outperformed a demographic baseline in predicting personality and economic behavior (Park et al. 2024, arXiv, [preprint]). On large-scale replication, Cui and colleagues, in a 2025 Nature Computational Science paper, reproduced 156 psychology experiments with several LLMs, finding 73–81% agreement on main effects but inflated effect sizes and reduced agreement on socially sensitive topics (Cui et al. 2025, Nature Computational Science).

The brakes on this optimism were also systematized in the same period. Lin, in 2025, listed six epistemic fallacies inherent in substituting LLMs for human participants, including conflating token prediction with human intelligence and treating an LLM as “the average human” (Lin 2025, Advances in Methods and Practices in Psychological Science). A proposal to rethink the craft of evaluation itself came earlier from Frank, who in 2023 argued for applying developmental-psychology assessment methods to measuring LLM capacities and asking “what abstractions is the LLM using” (Frank 2023, Nature Reviews Psychology).

Cross-Cutting Themes

Several threads run through the four clusters.

First, the reproducibility problem of divergent pictures from the same observation in machine psychology. Success on standard tasks coexists with failure on equivalent perturbations (Binz & Schulz 2023; Ullman 2023), and only large-scale human comparison finally separated “genuinely solving” from “a bias aligned with the scoring key” (Strachan et al. 2024).

Second, the unsettled status of whether an LLM is a tool, a subject, a model, or a counterexample of cognition. Centaur demonstrated it can function as a model of behavior prediction (Binz et al. 2025), while van Rooij et al. (2024) and Guest and Martin (2023) reject the leap from matching performance to sharing process. The same object looks like a tool in implementation research and a counterexample in critical research.

Third, the diminished returns of human-AI collaboration and cognitive debt. Collaboration can, on average, underperform the best single agent in some domains (Vaccaro et al. 2024), and reliance can thin critical thinking and brain connectivity (Gerlich 2025; Lee et al. 2025, CHI; Kosmyna et al. 2025). The picture of individual efficiency coinciding with a loss of judgment and diversity echoes the “acceleration running alongside homogenization” reported in the design domain (see ai-cognition-skill-gap-debate).

Fourth, the validity of using LLMs as substitutes for human participants. High replication agreement but inflated effect sizes that break down on sensitive topics (Cui et al. 2025) accumulate at the same time as warnings about six epistemic fallacies (Lin 2025) and illusions of understanding (Messeri & Crockett 2024).

  • Turning points: the prediction accuracy and neural alignment assumptions of Centaur (Binz et al. 2025, Nature); the moderator analysis in Vaccaro et al. (2024).
  • Disentangling theory of mind: the follow-up manipulations in Strachan et al. (2024); the perturbation design in Ullman (2023).
  • Methodological adjudication: the four archetypes of Messeri & Crockett (2024); the six fallacies of Lin (2025); the effect-size inflation in Cui et al. (2025).
  • Cognitive debt: the EEG measures in Kosmyna et al. (2025, [preprint]); the mediation model in Gerlich (2025).
  • ai-research-gaps-abduction — A map of the gaps in AI research found through abduction-2, questioning the very frames these research currents take for granted

References

30 items total. [preprint] marks non-peer-reviewed preprints; [requires primary verification] marks partly unconfirmed bibliographic details (not affecting key findings). Links are DOI or arXiv. The internal working ledger is at source/review/ai-psychology-cognitive-science-trends/papers.md.

Machine Psychology and Its Methodological Controversies

  • Binz, M., & Schulz, E. (2023). Using cognitive psychology to understand GPT-3. PNAS 120(6), e2218523120. https://doi.org/10.1073/pnas.2218523120
  • Mitchell, M., & Krakauer, D. C. (2023). The debate over understanding in AI’s large language models. PNAS 120(13), e2215907120. https://doi.org/10.1073/pnas.2215907120
  • Hagendorff, T., Dasgupta, I., Binz, M., Chan, S. C. Y., Lampinen, A., Wang, J. X., Akata, Z., & Schulz, E. (2023/2024). Machine Psychology: Investigating Emergent Capabilities and Behavior in Large Language Models Using Psychological Methods. arXiv:2303.13988 [preprint]. https://arxiv.org/abs/2303.13988
  • Ullman, T. (2023). Large Language Models Fail on Trivial Alterations to Theory-of-Mind Tasks. arXiv:2302.08399 [preprint]. https://arxiv.org/abs/2302.08399
  • Kosinski, M. (2024). Evaluating large language models in theory of mind tasks. PNAS 121(45), e2405460121. https://doi.org/10.1073/pnas.2405460121
  • Strachan, J. W. A., et al. (2024). Testing theory of mind in large language models and humans. Nature Human Behaviour 8(7), 1285–1295. https://doi.org/10.1038/s41562-024-01882-z
  • Suri, G., Slater, L. R., Ziaee, A., & Nguyen, M. (2024). Do large language models show decision heuristics similar to humans? A case study using GPT-3.5. Journal of Experimental Psychology: General 153(4), 1066–1075. https://doi.org/10.1037/xge0001547
  • Abdurahman, S., et al. (2024). Perils and opportunities in using large language models in psychological research. PNAS Nexus 3(7), pgae245. https://doi.org/10.1093/pnasnexus/pgae245
  • Shanahan, M. (2024). Talking about Large Language Models. Communications of the ACM 67(2), 68–79. https://doi.org/10.1145/3624724
  • Gao, Y., Lee, D., Burtch, G., & Fazelpour, S. (2025). Take caution in using LLMs as human surrogates. PNAS 122(24), e2501660122. https://doi.org/10.1073/pnas.2501660122

LLMs as Cognitive Models and Their Critiques

  • Binz, M., Akata, E., Bethge, M., Brändle, F., et al. (2025). A foundation model to predict and capture human cognition. Nature 644(8078), 1002–1009. https://doi.org/10.1038/s41586-025-09215-4
  • Binz, M., & Schulz, E. (2024). Turning large language models into cognitive models. ICLR 2024. https://arxiv.org/abs/2306.03917
  • Mahowald, K., Ivanova, A. A., Blank, I., Kanwisher, N., Tenenbaum, J. B., & Fedorenko, E. (2024). Dissociating language and thought in large language models. Trends in Cognitive Sciences 28(6) [volume/issue requires primary verification]. https://doi.org/10.1016/j.tics.2024.01.011
  • Contreras Kallens, P., Kristensen-McLachlan, R. D., & Christiansen, M. H. (2023). Large Language Models Demonstrate the Potential of Statistical Learning in Language. Cognitive Science 47(3), e13256. https://doi.org/10.1111/cogs.13256
  • Piantadosi, S. T., & Hill, F. (2022). Meaning without reference in large language models. arXiv:2208.02957 [preprint]. https://arxiv.org/abs/2208.02957
  • Katzir, R. (2023). Why Large Language Models Are Poor Theories of Human Linguistic Cognition: A Reply to Piantadosi. Biolinguistics 17, e13153. https://doi.org/10.5964/bioling.13153
  • van Rooij, I., Guest, O., Adolfi, F., de Haan, R., Kolokolova, A., & Rich, P. (2024). Reclaiming AI as a theoretical tool for cognitive science. Computational Brain & Behavior 7(4), 616–636. https://doi.org/10.1007/s42113-024-00217-5
  • Guest, O., & Martin, A. E. (2023). On logical inference over brains, behaviour, and artificial neural networks. Computational Brain & Behavior 6, 213–227. https://doi.org/10.1007/s42113-022-00166-x
  • McGrath, S. W., Russin, J., Pavlick, E., & Feiman, R. (2024). How Can Deep Neural Networks Inform Theory in Psychological Science? Current Directions in Psychological Science 33(5), 325–333. https://doi.org/10.1177/09637214241268098
  • Oh, B.-D., & Schuler, W. (2023). Why Does Surprisal From Larger Transformer-Based Language Models Provide a Poorer Fit to Human Reading Times? Transactions of the ACL 11, 336–350. https://doi.org/10.1162/tacl_a_00548

The Psychology of Human-AI Interaction

  • Vaccaro, M., Almaatouq, A., & Malone, T. (2024). When combinations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour 8(12), 2293–2303. https://doi.org/10.1038/s41562-024-02024-1
  • Salvi, F., Horta Ribeiro, M., Gallotti, R., & West, R. (2025). On the conversational persuasiveness of GPT-4. Nature Human Behaviour 9(8), 1645–1653. https://doi.org/10.1038/s41562-025-02194-6
  • Klingbeil, A., Grützner, C., & Schreck, P. (2024). Trust and reliance on AI — An experimental study on the extent and costs of overreliance on AI. Computers in Human Behavior 160, 108352. https://doi.org/10.1016/j.chb.2024.108352
  • Lee, D., Pruitt, J., Zhou, T., Du, J., & Odegaard, B. (2025). Metacognitive sensitivity: The key to calibrating trust and optimal decision making with AI. PNAS Nexus 4(5), pgaf133. https://doi.org/10.1093/pnasnexus/pgaf133
  • Marchegiani, B. (2025). Anthropomorphism, False Beliefs, and Conversational AIs: How Chatbots Undermine Users’ Autonomy. Journal of Applied Philosophy 42(5), 1399–1419 [requires primary verification]. https://doi.org/10.1111/japp.70008
  • Gerlich, M. (2025). AI Tools in Society: Impacts on Cognitive Offloading and the Future of Critical Thinking. Societies 15(1), 6. https://doi.org/10.3390/soc15010006
  • Lee, H.-P., Sarkar, A., Tankelevitch, L., Drosos, I., Rintel, S., Banks, R., & Wilson, N. (2025). The Impact of Generative AI on Critical Thinking. CHI 2025. https://doi.org/10.1145/3706598.3713778
  • Kosmyna, N., et al. (2025). Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task. arXiv:2506.08872 [preprint]. https://arxiv.org/abs/2506.08872
  • Fang, C. M., et al. (2025). How AI and Human Behaviors Shape Psychosocial Effects of Extended Chatbot Use: A Longitudinal Randomized Controlled Study. arXiv:2503.17473 [preprint]. https://arxiv.org/abs/2503.17473

The Methodological Turn in Cognitive Science

Unverified Items

  • Hagendorff et al. (2023/2024, machine psychology), Ullman (2023, ToM perturbation), Piantadosi & Hill (2022), Park et al. (2024), Kosmyna et al. (2025), and Fang et al. (2025) are confirmed only as non-peer-reviewed preprints (flagged [preprint] in-text).
  • The exact volume/issue of Mahowald et al. (2024, TiCS) requires final verification (DOI and the paper itself are confirmed).
  • The volume/pages of Marchegiani (2025) were cross-checked via PhilPapers because the publisher page required authentication.
  • Some full texts on Nature / Cell / Sage / Elsevier / MDPI / Wiley were behind authentication or 403; bibliographic details were cross-checked via independent repositories (PubMed, IDEAS.repec, Cambridge Core, ACL Anthology, Crossref API), and DOI/volume/pages are confirmed.

Update Policy

This is a living page. As new peer-reviewed papers and major preprints appear, they will be added to the body and updated will be revised. [requires primary verification] items are resolved in the internal ledger and reflected in the references here.


← All Notes · Home