Shuichiro Ogawa
日本語

Notes · updated 2026-07-19

Research Currents in AI, Education, and the Learning Sciences: Evidence on Generative AI and Learning (2024–2026)

Students who used generative AI got better at solving practice problems. Yet on an exam where the AI was taken away, their scores actually dropped. A randomized experiment with roughly a thousand high school students in Turkey put a number on this gap (Bastani et al. 2025).

Read this one case as a miniature of the whole research field on generative AI and learning. The effect is real. But the sign of the effect can flip depending on how the AI is designed into the task.

Between 2024 and mid-2026, peer-reviewed research at the intersection of generative AI and the education and learning sciences moved past scattered case reports into a body that can support meta-analyses and reviews. This note is a synthesis of 39 peer-reviewed studies gathered by lightweight scoping over that window. The scope is general education and the learning sciences; work specific to design education is left to a separate note (design-education-ai-adaptation). Full bibliographic details are in “References” below; the internal ledger with provenance is at source/review/ai-education-learning-sciences-trends/papers.md (not public).

All effect sizes are reported as claims of the cited authors, not asserted as fact. Where the publisher site was paywalled, DOI existence was confirmed via Crossref, PubMed, ERIC, PMC, and similar metadata; figures that could not be verified in the full text are flagged [needs primary verification] in the internal ledger.

Is there an effect, and how large?

Start with the average. From 2024 through 2026 a series of meta-analyses aggregated the learning effects of generative AI. Their target studies and learner populations differ, yet the reported effect sizes cluster around a moderate positive value.

Ma & Zhong (2025) pooled 34 experimental and quasi-experimental studies and reported an overall learning-outcome effect of g=0.68 (p<0.001). Liu, Zuo & Lu (2025), focusing on ChatGPT, obtained g=0.577 across 37 studies; Sun & Zhou (2024) reported g=0.533 from 28 papers on college students. Even for higher-order thinking, Zhao et al. (2025) reported g=0.609 across 29 studies (59 effect sizes), broken down into problem-solving g=0.745, critical thinking g=0.691, and creativity g=0.444.

The convergence is reassuring. Before taking it at face value, though, two things are worth checking.

One is retraction. A meta-analysis published in 2025 and viewed more than 480,000 times (Wang & Fan 2025, reporting learning-performance g=0.867 among others) was retracted in April 2026 for data inconsistencies. Larger effect sizes get cited more readily, and once a cited value is retracted, the number keeps circulating on its own. This note does not use that paper as grounds for any claim.

The other is what the average hides. In the same data from Zhao et al. (2025), the effect was g=0.863 for learners with high self-regulated-learning ability and g=0.284 for those with low ability, nearly a threefold spread. A positive average g does not mean the effect is positive for everyone.

Did AI tutors surpass the classroom?

Now step away from averages toward the individual experiment with the strongest methodology. In 2025, the paper that became its emblem appeared.

Kestin et al. (2025) compared a research-based AI tutor against in-class active learning covering the same content, in a physics course at Harvard University. In a crossover randomized controlled trial, the AI-tutor condition produced significantly higher post-test scores, and in less learning time. The effect size was moderate to large (as reported; the exact range is behind a paywall, flagged [needs primary verification] in the ledger). The claim that shifted the field’s discussion was not “AI is as good as a human instructor” but “AI outperformed one.”

A well-designed AI tutor is powerful. That much seems solid. The catch lies in what “well-designed” contains.

Return to Bastani et al. (2025), and that content comes into view. Access to a GPT that gave direct answers improved practice performance by +48%, yet scores on a written exam with the AI removed fell by −17% relative to the control (as reported). In the same experiment, a GPT Tutor condition designed to release hints gradually mitigated this negative transfer. What separated the outcomes was not the presence of AI but whether the AI handed over answers or handed over scaffolding.

The effect of AI feedback is likewise not uniform. In a field experiment on statistics tasks, Bauer et al. (2025) found that for highly structured tasks, LLM-generated adaptive feedback showed no advantage over static expert solutions, and interest actually declined under the LLM condition. By contrast, Lu et al. (2026), in an “AI-mediated” design where the AI offered suggestions to a teaching assistant, improved the quality of student revisions by d=0.50 (introductory undergraduate economics, 354 students, 1,366 essays). The design that worked was one where AI assisted the instructor’s judgment rather than replacing it.

A systematic review of K-12 intelligent tutoring systems (ITS) (Létourneau et al. 2025, 28 studies) lands on the same caution. Effects are generally positive, but the margin narrows against human tutoring, and larger, longer, more diverse replications are needed.

Scores may rise while understanding does not

The effect sizes so far are mostly measured on the produced artifact. So when a score rises, does something rise in the learner’s head by the same amount? The study that reframed the question this way drew the most attention on the learning-sciences side over this period.

Fan et al. (2024/2025) randomly assigned university students to four conditions (ChatGPT, human expert, writing-analytics tool, control). The ChatGPT group improved essay scores relative to the other conditions. Yet knowledge acquisition and transfer showed no significant difference between conditions. The authors named this “metacognitive laziness,” arguing that reliance on AI makes learners more likely to relinquish the processes of planning, monitoring, and evaluating their own work.

This gap, “scores rise but understanding does not,” became a common starting point for later work. The cognitive-offloading line explains the gap in the language of cognition.

Gerlich (2025), in a mixed-methods study of 666 participants, found a significant negative correlation between frequency of AI use and critical thinking, with cognitive offloading as the mediator (pronounced among younger users). As evidence beyond the learning setting, Lee et al. (2025) surveyed 319 knowledge workers and showed that higher trust in generative AI was associated with less engagement in critical thinking, while higher confidence in one’s own judgment was associated with more. Both trace the same logic: when AI takes over part of the thinking, the capacity it takes over goes unused.

To avoid tipping too far toward the negative, one concession belongs here. There is evidence that scaffolding design can bend this logic the other way.

In an RCT by Xu et al. (2025), the group given metacognitive support showed no significant difference on achievement scores themselves, but improved on the self-regulated-learning facets of task strategy and self-evaluation. Melanou’s (2026) quasi-experiment reports that a structured prompt scaffold requiring learners to state goal, context, and constraints improved knowledge acquisition and reflection during AI use. A meta-analysis of a decade of AI support for self-regulated learning (Xu et al. 2025) placed the overall effect at g=0.507, with the effect larger in the enactment phase (g=0.574) than the preparation phase (g=0.401). When support is placed on the side of “supporting how to proceed” rather than “producing the answer,” there is room to run metacognitive laziness in reverse.

Even granting the concession, the gap itself does not vanish. That is why treating artifact scores alone as the measure of effect overlooks the case where AI is encouraging the outsourcing of learning.

When no one knows who wrote it, what do you assess?

If understanding and scores diverge, the assessment that produces those scores wobbles too. Generative AI made it ambiguous “who wrote” the text a student submits. Institutions first answered this wobble with detection.

Detection is less reliable than hoped. Liang et al. (2023) showed that seven GPT detectors misclassified an average of 61.3% of TOEFL essays by non-native English writers as AI-generated, while misflagging English by U.S. eighth-graders far less often. The detectors were responding not to AI-ness but to the simplicity of the English. When Hadra et al. (2026) evaluated commercial detectors across 192 texts, accuracy dropped for hybrid texts, EFL writers, and long texts, and they concluded detectors should remain an aid and not the sole basis for a misconduct charge.

If detection cannot be trusted, the next idea was to rebuild the tasks themselves into ones “hard to replace with AI.” A caveat attaches here too. Kofinas, Tsay & Pike (2025), in an exploratory experiment, showed that authentic assessment alone cannot guarantee academic integrity, and that work produced with generative AI is hard to distinguish during grading.

Student behavior, too, resists a simple “cheating or not.” In a survey of 2,555 students at the University of Liverpool (Johnston et al. 2024), only 7% were unaware of generative AI, 70.4% were negative about having ChatGPT write an entire essay, yet more than half were positive about grammar-support tools. In an Australian survey of 337 students (Gruenhagen et al. 2024), over a third used chatbots to assist with assessments, and many did not regard this as an integrity violation. On where the line should fall, student intuitions diverge.

In this situation, an all-or-nothing dichotomy of ban or permit spread in practice. Curtis (2025) rejects it directly. Full permission damages the learning process; a full ban ignores the educational value of constrained use. What is needed is a middle design that is neither. Assessment research is shifting its center of gravity from reliance on detection toward designing tasks and integrity around the assumption that AI is present (the scoping review by Xia et al. 2024 points the same direction).

Leveling up, or widening the gap?

Effects and assessment have been the focus so far, but talk of the average never asks “whose average.” Does generative AI lower the barrier to entry and lift learning broadly, or widen the gap between those who can use it and those who cannot?

Evidence for both sits side by side. On the inclusive-education side, Khlaif et al. (2025), following 21 Palestinian undergraduates with visual impairment, describe generative AI as a potential support tool for narrowing gaps. On the other hand, Capraro et al. (2024), examining generative AI’s effect on socioeconomic inequality, point to gender differences (lower ChatGPT-use rates among women) and the risk of a widening digital divide in education. Unless equal access is guaranteed, the leveling-up effect is distributed unevenly.

What tips the distribution is the teacher on the ground. In a survey of 89 K-12 teachers in Idaho, USA (Cheah et al. 2025), 61% reported being “hardly prepared,” citing absent district policy, ethical concerns, and infrastructure disparities as barriers. A scoping review of the changing teacher role across 43 studies (Tan 2025) finds that professional development using generative AI can connect to seven outcomes, including instructional design, self-efficacy, and AI literacy. On the institutional side, Jin et al. (2025), analyzing 98 university strategy documents worldwide, identify common patterns: emphasis on trialability, promotion of AI literacy, and a focus on human-centered competencies.

Institutional design also holds a finding that betrays a naive expectation. In two R1 universities in the U.S. (Jiang et al. 2025), students who were aware of the AI policy had, if anything, lower AI-use rates in writing and research. A policy does not communicate merely by being written down. Unless the design extends to how it is conveyed, an institution’s intent does not reach behavior on the ground.

How to read the currents

Return to the Turkish high school from the opening. Generative AI helped practice while lowering the exam not because the AI was superior or inferior. It was because a design that hands over answers took from learners the occasion to think for themselves.

The image running through this period’s research emerges as a tension across two levels. Individual artifact scores generally rise (meta-analytic g is moderately positive), while at the deeper layers of understanding, transfer, critical thinking, and metacognition, the effect moves up or down with design. And at the group level, leveling-up and gap-widening can arise from the same technology at once.

So the field’s center of gravity moved from “does AI work” to “under what design does it work, for which learners, on which capacities.” Separate designs that hand over scaffolding from those that hand over answers; separate artifact scores from understanding; separate the average from the distribution when measuring. The largest open question is whether the effects seen in controlled university experiments can be reproduced in K-12 classrooms with uneven resources and preparation, without widening the gap. That question stays open.

References

39 items, traceable by DOI or arXiv/official URL. All figures are as reported by the cited authors and are not asserted in the text. The retracted paper is listed separately at the end and is not used as grounds for any claim in this note. The internal working ledger is at source/review/ai-education-learning-sciences-trends/papers.md.

Tutoring, feedback, ITS, meta-analyses

  • Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. (2025). AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports 15. https://doi.org/10.1038/s41598-025-97652-6
  • Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakçı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. PNAS 122. https://doi.org/10.1073/pnas.2422633122 (correction: https://doi.org/10.1073/pnas.2518204122)
  • Ma, N., & Zhong, Z. (2025). A Meta-Analysis of the Impact of Generative Artificial Intelligence on Learning Outcomes. Journal of Computer Assisted Learning. https://doi.org/10.1111/jcal.70117
  • Liu, Z., Zuo, H., & Lu, Y. (2025). The Impact of ChatGPT on Students’ Academic Achievement: A Meta-Analysis. Journal of Computer Assisted Learning. https://doi.org/10.1111/jcal.70096
  • Sun, L., & Zhou, L. (2024). Does Generative Artificial Intelligence Improve the Academic Achievement of College Students? A Meta-Analysis. Journal of Educational Computing Research. https://doi.org/10.1177/07356331241277937
  • Liu, X., Guo, B., He, W., & Hu, X. (2025). Effects of Generative Artificial Intelligence on K-12 and Higher Education Students’ Learning Outcomes: A Meta-Analysis. Journal of Educational Computing Research 63(5). https://doi.org/10.1177/07356331251329185
  • Zhao, Y., Yue, Y., Sun, Z., Jiang, Q., & Li, G. (2025). Does Generative Artificial Intelligence Improve Students’ Higher-Order Thinking? A Meta-Analysis Based on 29 Experiments and Quasi-Experiments. Journal of Intelligence 13(12). https://doi.org/10.3390/jintelligence13120160
  • Mo, F., Huang, J., Yang, Y., Özen, Z., Maeda, Y., & Olenchak, F. R. (2025). Undergraduate students’ learning outcomes with ChatGPT: A meta-analytic study. Computers in Human Behavior Open. https://www.sciencedirect.com/science/article/pii/S2666920X25001766
  • Bauer, E., Greiff, S., Graesser, A. C., Scheiter, K., & Sailer, M. (2025). Looking beyond the Hype: Understanding the Effects of AI on Learning. Educational Psychology Review. https://doi.org/10.1007/s10648-025-10020-8
  • Bauer, E., et al. (2025). Effects of Artificial Intelligence on Educational Functioning: A Review and Meta-Analysis. Educational Psychology Review. https://doi.org/10.1007/s10648-025-10085-5
  • Létourneau, A., Deslandes Martineau, M., Charland, P., Karran, J. A., Boasen, J., & Léger, P. M. (2025). A systematic review of AI-driven intelligent tutoring systems (ITS) in K-12 education. npj Science of Learning. https://doi.org/10.1038/s41539-025-00320-7
  • Bauer, E., Greiff, S., Scheiter, K., & Sailer, M. (2025). Effects of AI-generated adaptive feedback on statistical skills and interest in statistics: A field experiment in higher education. British Journal of Educational Technology. https://doi.org/10.1111/bjet.13609
  • Lu, X., Ju, K. P., Dudley, M., Sano, L., & Wang, X. (2026). AI-Mediated Feedback Improves Student Revisions: A Randomized Trial with FeedbackWriter in a Large Undergraduate Course. CHI 2026. https://doi.org/10.1145/3772318.3791121
  • LearnLM Team (Google DeepMind). (2025). AI tutoring can safely and effectively support students: An exploratory RCT in UK classrooms. arXiv:2512.23633 [preprint]. https://arxiv.org/abs/2512.23633

Learning-science theory (self-regulated learning, cognitive offloading, scaffolding)

  • Fan, Y., Tang, L., Le, H., Shen, K., Tan, S., Zhao, Y., Shen, Y., Li, X., & Gašević, D. (2024/2025). Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performance. British Journal of Educational Technology 56. https://doi.org/10.1111/bjet.13544
  • Gerlich, M. (2025). AI Tools in Society: Impacts on Cognitive Offloading and the Future of Critical Thinking. Societies 15(1). https://doi.org/10.3390/soc15010006
  • Lee, H.-P., Sarkar, A., Tankelevitch, L., Drosos, I., Rintel, S., Banks, R., & Wilson, N. (2025). The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers. CHI 2025. https://doi.org/10.1145/3706598.3713778
  • Xu, X., Qiao, L., Cheng, N., Liu, H., & Zhao, W. (2025). Enhancing Self-Regulated Learning and Learning Experience in Generative AI Environments: The Critical Role of Metacognitive Support. British Journal of Educational Technology. https://doi.org/10.1111/bjet.13599
  • Xu, J., Luo, Y., et al. (2025). AI Support in Self-Regulated Learning: A Decade of Technological Evolution and Meta-Analysis. British Journal of Educational Technology. https://doi.org/10.1111/bjet.70058
  • Fütterer, T., Bardach, L., Kuhn, J., Keller, S. D., & Gerjets, P. (2026). Enhancing School Students’ Self-Regulated Learning through Generative AI Support: A Randomized Controlled Trial. Educational Psychology Review. https://doi.org/10.1007/s10648-026-10133-8
  • Ng, D.T.K., Tan, C.W., & Leung, J.K.L. (2024). Empowering Student Self-Regulated Learning and Science Education through ChatGPT: A Pioneering Pilot Study. British Journal of Educational Technology. https://doi.org/10.1111/bjet.13454
  • Melanou, C. (2026). Scaffolding Generative AI as a Tutor: A Quasi-Experimental Study of Learning Outcomes and Motivational, Cognitive and Metacognitive Processes. Education Sciences 16(4). https://doi.org/10.3390/educsci16040651
  • Lan, M., & Zhou, X. (2025). A Qualitative Systematic Review on AI Empowered Self-Regulated Learning in Higher Education. npj Science of Learning. https://doi.org/10.1038/s41539-025-00319-0

Assessment and academic integrity

  • Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns (Cell Press) 4. https://doi.org/10.1016/j.patter.2023.100779
  • Xia, Q., Weng, X., Ouyang, F., Lin, T. J., & Chiu, T. K. F. (2024). A scoping review on how generative artificial intelligence transforms assessment in higher education. International Journal of Educational Technology in Higher Education. https://doi.org/10.1186/s41239-024-00468-z
  • Kofinas, A. K., Tsay, C. H.-H., & Pike, D. (2025). The impact of generative AI on academic integrity of authentic assessments within a higher education context. British Journal of Educational Technology 56(6). https://doi.org/10.1111/bjet.13585
  • Johnston, H., Wells, R. F., Shanks, E. M., Boey, T., & Parsons, B. N. (2024). Student perspectives on the use of generative artificial intelligence technologies in higher education. International Journal for Educational Integrity. https://doi.org/10.1007/s40979-024-00149-4
  • Gruenhagen, J. H., Sinclair, P. M., Carroll, J.-A., Baker, P. R. A., Wilson, A., & Demant, D. (2024). The rapid rise of generative AI and its implications for academic integrity: Students’ perceptions and use of chatbots for assistance with assessments. Computers and Education: Artificial Intelligence. https://doi.org/10.1016/j.caeai.2024.100273
  • Curtis, G. J. (2025). The two-lane road to hell is paved with good intentions: why an all-or-none approach to generative AI, integrity, and assessment is insupportable. Higher Education Research & Development 44(8). https://doi.org/10.1080/07294360.2025.2476516
  • Hadra, M. G., Cambridge, K., & Mesbah, M. (2026). Evaluating the accuracy and reliability of AI content detectors in academic contexts. International Journal for Educational Integrity. https://doi.org/10.1007/s40979-026-00213-1

Equity, the changing role of teachers, and institutional responses

  • Yusuf, A., Pervin, N., Román-González, M., & Noor, N.M. (2024). Generative AI in education and research: A systematic mapping review. Review of Education 12, e3489. https://doi.org/10.1002/rev3.3489
  • Lin, X., & Tan, H. (2025). A Systematic Review of Generative AI in K–12: Mapping Goals, Activities, Roles, and Outcomes via the 3P Model. Systems 13(10), 840. https://doi.org/10.3390/systems13100840
  • Jin, Y., Yan, L., Echeverria, V., Gašević, D., & Martinez-Maldonado, R. (2025). Generative AI in higher education: A global perspective of institutional adoption policies and guidelines. Computers and Education: Artificial Intelligence 8, 100348. https://doi.org/10.1016/j.caeai.2024.100348
  • Tan, Q. (2025). Reimagining teacher development in the era of generative AI: A scoping review. Teaching and Teacher Education 168, 105236. https://doi.org/10.1016/j.tate.2025.105236
  • Cheah, Y.H., Lu, J., & Kim, J. (2025). Integrating generative artificial intelligence in K-12 education: Examining teachers’ preparedness, practices, and barriers. Computers and Education: Artificial Intelligence 8, 100363. https://doi.org/10.1016/j.caeai.2025.100363
  • Capraro, V., Lentsch, A., Acemoglu, D., et al. (2024). The impact of generative artificial intelligence on socioeconomic inequalities and policy making. PNAS Nexus 3(6), pgae191. https://doi.org/10.1093/pnasnexus/pgae191
  • Khlaif, Z.N., Alshakhshir, R., Hamamra, B., & Joma, A. (2025). Reimagining inclusive education: The assistive power of generative AI in promoting accessibility and equity. British Journal of Visual Impairment. https://doi.org/10.1177/02646196251382469
  • Jiang, Y., Xie, L., & Cao, X. (2025). Exploring the Effectiveness of Institutional Policies and Regulations for Generative AI Usage in Higher Education. Higher Education Quarterly 79(4). https://doi.org/10.1111/hequ.70054
  • UNESCO (Miao, F., & Holmes, W.). (2023). Guidance for Generative AI in Education and Research. UNESCO. https://www.unesco.org/en/articles/guidance-generative-ai-education-and-research

Retracted (not used as grounds in this note; flagged for caution)

ai-psychology-cognitive-science-trends / ai-humanities-digital-humanities-trends / design-education-ai-adaptation


← All Notes · Home