Shuichiro Ogawa
日本語

Notes · updated 2026-08-08

Scope and Method

The question of this note is how far the claim is academically supported that using generative AI early in learning lets the AI’s intervention inhibit the formation of the next thoughts that would otherwise have emerged from the learner. The claim is the reverse face of the lineage organized in Designing Failure Into Learning: an AI that preemptively supplies answers erases the situations in which expectation failure and desirable difficulties would have occurred. Collection was academic mode: experimental reports of harm (group A, 7 works), correlational/self-report studies (group B, 2), theoretical frameworks (group C, 3), and counter/conditioning evidence (group D, 3), plus the record of one retracted positive meta-analysis, for 16 works (source/review/ai-cognitive-offloading-learning/papers.md). Preprints carry flag: non-peer, and effect claims are reported as the papers’ claims.

Theoretical Frameworks: The Claim Is Not New

What looks specific to generative AI is a new instance of three established frameworks.

The assistance dilemma (Koedinger & Aleven 2007) was formulated in intelligent-tutoring research: too much assistance yields shallow learning and loss of autonomous motivation; too little yields wasted struggle. The current question of how much an AI should tell is the generative-AI version of this dilemma. Cognitive offloading (Risko & Gilbert 2016) organized the efficiency of delegating cognitive processing to external tools and its consequence for learning (delegated processing is not internalized). Its empirical predecessor is the Google effect (Sparrow et al. 2011): merely expecting external access to information reduces encoding of the content itself, and people remember where instead of what. Where search delegated the location of information, generative AI delegates the generation of thought itself, one level deeper as an offloading target. For the broader set of claimed technological harms to education (including the Google effect) and the robustness assessment of their peer-reviewed refutations, see Patterns of Technological Harm to Education.

Experimental Evidence: Where Harm Appears

The most direct test is the field RCT by Bastani et al. (PNAS 2025; Turkish high-school mathematics, about 1,000 students). Students with plain GPT-4 improved practice performance by 48%, but scored 17% below controls on an exam without AI access. The practice gains were an illusion of learning; the thinking the AI did on their behalf did not stay with the students. The same experiment’s other condition matters equally: with GPT Tutor, designed to withhold answers and instead prompt and scaffold, practice improved 127% with no significant exam decline.

The pattern of looking good with AI and retaining nothing without it replicates across independent experiments. Darvishi et al. (2024; 1,625 students, 10 courses) reported that students tend to rely on AI assistance rather than learn from it, with performance dropping once assistance is removed. Fan et al. (2025) showed a ChatGPT group outscoring others on the essay while showing no difference in knowledge gain or transfer, naming the suppression of metacognitive evaluation and orientation under AI reliance metacognitive laziness. Stadler et al. (2024) reported that a ChatGPT group experienced lower cognitive load than a web-search group yet produced significantly worse reasoning and justification: cognitive ease at the cost of depth.

In novice programming education, harm diverges by learner stratum. In Kazemitabaar et al. (CHI 2023; 69 novices), the Codex group scored 1.8× on code-writing but showed no difference on subsequent modification tasks, and outcomes split between passive users (copy-paste) and active users (verify and adapt). Prather et al. (ICER 2024) observed an illusion of competence in struggling novices (inflated self-assessment against objective evidence) and reported that AI-specific metacognitive difficulties correlate negatively with course performance (r = −0.503): high performers used AI for acceleration while weak performers depended on it and masked their lack of understanding — the “widening gap” of the title.

As preliminary neurophysiological evidence, Kosmyna et al. (MIT Media Lab, 2025, preprint, not yet peer reviewed) tracked essay writing over four months with EEG and conceptualized the LLM group’s weakest brain-network connectivity as cognitive debt. A critical commentary questions the sample size and analytic reproducibility, so pending peer review this stands as corroboration only.

Correlational and Self-Report Evidence

Cross-sectional studies point the same way without settling causality. Gerlich (2025; 666 participants) reported a negative correlation between AI-use frequency and critical-thinking scores, mediated by cognitive offloading (younger users more dependent). Lee et al. (CHI 2025; 319 knowledge workers, 936 cases) reported that higher confidence in generative AI predicts less critical-thinking effort while higher self-confidence predicts more, and that critical thinking shifts from generating one’s own thought to verifying and integrating AI output. Both carry the limits of self-report and cross-sectional design: they cannot separate “AI use erodes thinking” from “those who think less turn to AI.”

Counter and Conditioning Evidence: Reconciling the Positive Meta-Analyses

The strongest opposing evidence is the positive meta-analytic literature on ChatGPT and learning. Deng et al. (2025; 69 studies) reported a large positive effect on academic performance (g = 0.712) with gains in higher-order thinking, and Wu et al. (2026; 35 studies, 4,193 participants) a moderate positive effect (g = 0.670). Kestin et al.’s (2025) Harvard physics RCT found a course-specific AI tutor outperforming in-class active learning.

The apparent contradiction dissolves on inspection of what is measured. The positive meta-analyses aggregate mostly performance during or immediately after AI use, while the harms detected by Bastani et al. and Darvishi et al. appear in what remains after the AI is removed. The same intervention that raised practice scores 48% lowered exam scores 17% — a live demonstration that the two measurements can yield opposite conclusions. This is the same structure as Contention 3 in the failure-design debate note (the learning-performance dissociation of Soderstrom & Bjork): the desirable-difficulties warning that performance during training misleads is being replayed with generative AI. As a record of the field’s overheating, one positive meta-analysis (Wang & Fan 2025) was retracted for numerical inconsistencies and is retained in the corpus as a retraction record.

Synthesis: How Far the Claim Is Supported

On current evidence, the opening claim is supported with conditions.

That early generative-AI use inhibits the formation of thought is supported by multiple independent RCTs and quasi-experiments when (1) the AI acts as an answer provider (no guardrails, passive use), (2) the learners are novices weak in canonical knowledge or metacognition, and (3) the outcome is measured as retention and transfer without the AI. Conversely, designs that withhold answers in favor of prompting and scaffolding (GPT Tutor, course-specific tutors) show no detected harm and sometimes improved learning. What the evidence supports is thus neither “generative AI is harmful” nor “harmless” but the old lesson of the assistance dilemma: the harm of assistance is a function of its amount and design; an AI that does the thinking leaves no thinking behind, and an AI that demands thinking does. This reads as the same design problem as the missing post-failure consolidation that Designing Failure Into Learning found in industry services. The curriculum-side response in design education is covered by design-education-ai-adaptation.

Gaps

  • Long-term longitudinal evidence (beyond a semester) is nearly absent; Kosmyna et al. (four months) is among the longest but unreviewed.
  • No study directly operationalizes “early” (at which stage of learning AI access begins to harm); current evidence rests on the coarse novice/non-novice contrast.
  • Decomposition of mitigation design variables (what exactly a guardrail must restrict) does not go beyond Bastani et al.’s two-condition comparison.
  • No tests were found in domains without unique correct answers (design education, system design education); the evidence concentrates in mathematics, physics, programming, and essay writing.
  • No Japanese-language empirical work was confirmed in this search.

Unverified Items

  • Specific effect-size statistics for Stadler et al. (2024) and Gerlich (2025) (full texts unreached)
  • Final DOI confirmation for Kazemitabaar et al. (CHI 2023) (.3580919 versus .3580950)
  • Current peer-review status of Kosmyna et al. (2025) and the full author list of the critical commentary (arXiv:2601.00856)

References

All accessed 2026-08-08. For ledger details, see source/review/ai-cognitive-offloading-learning/papers.md.

Experimental reports of harm

  • Bastani, H., Bastani, O., Sungu, A., et al. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. PNAS, 122(26). https://doi.org/10.1073/pnas.2422633122
  • Fan, Y., Tang, L., et al. (2025). Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performance. British Journal of Educational Technology, 56(2), 489–530. https://doi.org/10.1111/bjet.13544
  • Stadler, M., Bannert, M., & Sailer, M. (2024). Cognitive ease at a cost: LLMs reduce mental effort but compromise depth in student scientific inquiry. Computers in Human Behavior, 160, 108386. https://doi.org/10.1016/j.chb.2024.108386
  • Darvishi, A., Khosravi, H., Sadiq, S., Gašević, D., & Siemens, G. (2024). Impact of AI assistance on student agency. Computers & Education, 210, 104967. https://doi.org/10.1016/j.compedu.2023.104967
  • Kazemitabaar, M., et al. (2023). Studying the effect of AI Code Generators on Supporting Novice Learners in Introductory Programming. CHI 2023. https://doi.org/10.1145/3544548.3580919
  • Prather, J., et al. (2024). The Widening Gap: The Benefits and Harms of Generative AI for Novice Programmers. ICER 2024, 469–486. https://doi.org/10.1145/3632620.3671116
  • Kosmyna, N., et al. (2025). Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task. arXiv:2506.08872. https://arxiv.org/abs/2506.08872 (preprint, non-peer)

Correlational and self-report

  • Gerlich, M. (2025). AI Tools in Society: Impacts on Cognitive Offloading and the Future of Critical Thinking. Societies, 15(1), 6. https://doi.org/10.3390/soc15010006
  • Lee, H.-P., et al. (2025). The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers. CHI 2025. https://doi.org/10.1145/3706598.3713778

Theoretical frameworks

Counter and conditioning evidence


← All Notes · Home