Shuichiro Ogawa
日本語

Notes · updated 2026-07-19

AI and Social Science Research Trends: Computational Social Science and LLM Agents (2024–2026)

As material for a review article, this is an integrated summary of 29 peer-reviewed papers, major perspective pieces, and highly cited preprints on the intersection of AI and the social sciences, collected via lightweight scoping. Full bibliographic details for each work are in the “References” section at the end (with DOI/URL, reachability confirmed). The internal working ledger with provenance tracking, confidence, and verification routes is source/review/ai-social-science-research-trends/papers.md (repository-internal, not published). Collection protocol: .claude/rules/collection-protocol.md (zero fabrication, primary-source priority, provenance tracking). Related notes: design-social-science-nexus, ai-economics-research-trends, ai-humanities-digital-humanities-trends.

Survey Metadata

  • Collected: 2026-07-19 / Count: 29 (lightweight scoping, prioritizing avoidance of blind spots over exhaustiveness)
  • Venue center of gravity: Science, PNAS, Nature, Nature Human Behaviour, Political Analysis, Sociological Methods & Research, Computational Linguistics, arXiv (cs.CL / cs.CY)
  • Confidence notes: Several works are at the preprint stage (unconfirmed peer-reviewed publication is collected at the end as [要一次検証], “needs primary verification”). Publisher home sites (science.org / pnas.org / nature.com) can be login-walled; in those cases bibliography was cross-checked via PubMed / PMC / ACL Anthology / Cambridge Core / RePEc / institutional pages.

Introduction

The path by which the social sciences came to reckon with AI lies where two currents crossed in 2023.

One current is the moment LLMs became usable as tools for processing human language data at scale. Empirical social science has spent human labor classifying and coding language data: survey responses, political texts, social media posts. If an LLM can take over that process, the cost and speed of measurement change abruptly. When Gilardi and colleagues reported in 2023 in PNAS that “ChatGPT outperforms crowd workers for text annotation,” this was not merely a tool evaluation; it became an occasion to reexamine social science’s measurement practice itself.

The other current is the possibility that LLMs might be usable as devices that simulate humans themselves. Argyle and colleagues showed that conditioning GPT-3 on demographic information could reproduce the opinion distributions of particular population segments, and called this a “silicon sample.” That same year, Park and colleagues released a system in which 25 generative agents, endowed with memory, reflection, and planning, behaved emergently inside a village. Human responses and behaviors might be obtained without assembling humans. This prospect shook the premises of survey research.

These two currents pose distinct questions to social scientists. Used as a tool, the question is whether its measurement is valid and reproducible. Used as a substitute for humans, the question is whether the simulated responses represent real people. The research currents from 2024 through 2026 can be read as an oscillation between this optimism and this caution. Below they are organized into four clusters: the LLM turn in computational social science, LLMs as research tools and their validity critiques, social agents and silicon sampling, and AI as a substantive object of study.

The LLM Turn in Computational Social Science

The agenda for the field as a whole was set by perspective pieces spanning 2023 to 2024.

Grossmann and colleagues argued in Science in 2023 that LLMs could change theory testing, simulated data generation, scaling, and decision support in the social sciences. This was one of the earliest influential short pieces to raise the field’s transformation as an agenda, and at the same time it named bias management and data fidelity as challenges. That the presentation of the agenda and the presentation of caution coexist in the same paper captures well the character of this current’s starting point.

Bail, in PNAS in 2024, systematically organized the applicability of generative AI across four domains: survey research, online experiments, content analysis, and agent-based models. At the same time he arrayed concerns about training-data bias, ethics, reproducibility, environmental cost, and the proliferation of low-quality research, proposing as a remedy the construction of open-source infrastructure led by social scientists. By not stopping at an enumeration of opportunities but discussing the side effects that opportunities bring, this paper structured the field’s debate.

Ziems and colleagues, in 2024, carved out the limits of capability empirically. They evaluated 13 LLMs zero-shot across 25 computational social science benchmarks, showing that while models fall short of fine-tuned ones on classification tasks, they can surpass crowd workers on open-ended coding. The contribution of this paper is that it decomposed the question “can LLMs transform social science” into answers that differ by task type.

LLMs as Research Tools and Their Validity Critiques

This cluster is where the empirics of optimism and the empirics of critique come closest.

On the optimistic side, Gilardi and colleagues showed in 2023 in PNAS that ChatGPT’s zero-shot annotation accuracy exceeded crowd workers by an average of about 25 percentage points, at roughly one-thirtieth of the MTurk cost. Törnberg, that same year, reported that in classifying politicians’ tweets, ChatGPT-4 surpassed both experts and crowd workers, and did so with smaller bias. The move to replace the labor of measurement, from humans to LLMs, obtained concrete numbers here.

The critical side arose by testing the premise of replacement. Pangakis and colleagues, in 2023, reproduced 27 tasks across 11 datasets with GPT-4 and showed that performance varies greatly by task, arguing that LLM annotation cannot be used without validation against human labels. Halterman and Keith, in 2025, asked whether LLMs faithfully follow actual codebook definitions, showing in Political Analysis that zero-shot F1 ranged from a low 0.21 to 0.65 and only approached usable levels after instruction tuning. Reports of high accuracy and reports of non-adherence to definitions coexist about the same tool.

From 2025 into 2026, papers followed that shaped this tension into methodological norms. Lin and Zhang, in 2025, proposed a framework for using LLMs as primary annotators or as assistants, systematizing four risks: validity, reliability, reproducibility, and transparency. Abdurahman and colleagues, the same year, published in Advances in Methods and Practices in Psychological Science a set of evaluation guidelines for researchers using LLMs for data analysis and synthetic data generation. Desai and colleagues, in a 2026 preprint, surveyed LLM measurement studies published across eight flagship social science journals and demonstrated empirically that validation practice is inadequate. The current is shifting from “can LLMs be used” to “what must be validated when using LLMs.”

Social Agents and Silicon Sampling

The idea of obtaining human responses or behaviors without assembling humans appears most sharply in this cluster.

There are two methodological origins. Argyle and colleagues showed in 2023 in Political Analysis that conditioning GPT-3 on a first-person demographic backstory could reproduce human survey response distributions, and formalized this goodness of fit under the concept of algorithmic fidelity. The same year, Park and colleagues released at UIST a system in which 25 generative agents endowed with a memory stream, reflection, and planning displayed emergent social behavior inside a sandbox. The former established “silicon sampling,” the latter the LLM-ization of agent-based social simulation, each as a methodology.

Optimism extended to the individual level from 2024 through 2026. Park and colleagues, in a 2024 preprint, grounded agents in semi-structured interviews with 1,052 real people and reported that the reproduction accuracy of the General Social Survey (GSS), which those people had themselves re-answered two weeks later, reached 83 to 86 percent of the people’s own test-retest consistency (the original title “Generative Agent Simulations of 1,000 People” has since been revised). Ashokkumar, Hewitt, and colleagues, in Nature in 2026, showed that across 70 pre-registered, nationally representative survey experiments, GPT-4’s simulated treatment-effect predictions achieved accuracy on par with the wisdom of the crowd of human forecasters (correlation r=0.85, and r=0.90 even for unpublished experiments not in the model’s training data). The claim is that both individual responses and experimental results can be predicted within a certain range.

Caution met this optimism with counter-evidence on the same ground. Bisbee and colleagues, in 2024, in the same Political Analysis, showed that synthetic public opinion generated by ChatGPT was close to real surveys in the mean but had markedly lower variance, and that roughly 48 percent of the regression coefficients differed significantly from the actual values. They further recorded that results diverged after an OpenAI update in June 2023, issuing a strong warning against using synthetic data for statistical inference. Lin, in 2025, systematically enumerated six fallacies latent in the argument that LLMs substitute for human participants, contending that “functional equivalence” and “mechanistic equivalence” must be distinguished. Qu and Wang, in 2024, showed that evaluating with the World Values Survey, LLM performance skews toward the West, English, and developed countries, faring much worse in non-Western contexts. The range over which simulation holds is not uniform, whether in mean and variance or across cultural spheres.

The current of this cluster can be read as a chain: Argyle (optimism), Bisbee (critique), Park and Ashokkumar (extension), Lin (conceptual critique). The compromise Dillion and colleagues offered in 2023 in Trends in Cognitive Sciences, positioning LLMs not as full substitutes but as tools for hypothesis generation and piloting, continues to be referenced as a practical landing point placed amid this oscillation.

AI as a Substantive Object of Study

The current that treats LLMs not as tools or substitutes but as the object of study itself also gained depth after 2024.

Effects on labor were measured through the frame of LLMs as a general-purpose technology. Eloundou and colleagues, in Science in 2024, estimated against the O*NET task taxonomy that about 80 percent of occupations are exposed to GPT-4 in at least 10 percent of their tasks, and reported the inversion that higher-wage occupations have higher exposure. Brynjolfsson and colleagues, in Quarterly Journal of Economics in 2025, demonstrated from a quasi-experiment covering 5,179 customer-support staff that a generative AI assistant produced an average 14 percent productivity gain (34 percent among novices) through a mechanism that diffuses experts’ tacit knowledge to novices. The finding that the distribution of effects is not uniform is the starting point of this domain.

Persuasion and misinformation were studied as measurements of AI’s power to act on public opinion. Costello and colleagues, in Science in 2024, showed that individual dialogues with GPT-4 Turbo reduced conspiracy beliefs by an average of about 20 percent, and that the effect persisted two months later (this paper carries an Editorial Expression of Concern dated June 2026; in response to the flagged data-processing errors, the authors report that the conclusions are unchanged. Cite with attention to that concern). Salvi and colleagues, in Nature Human Behaviour in 2025, showed that a GPT-4 personalized with demographic data raised the post-hoc odds of agreement about 81 percent higher than a human debater, demonstrating the danger of AI microtargeted persuasion. On the other hand, Argyle and colleagues reported in PNAS in 2025 that among generative-AI persuasion strategies, personalization and interactive elaboration showed no additive advantage over generic messages, and Bai and colleagues, the same year, identified a mechanistic difference whereby LLMs persuade through perceived “facts, evidence, and logic” while humans persuade through perceived “uniqueness and originality.” The 2025 meta-analysis by Hölbling and colleagues concluded that the overall difference between LLM and human persuasiveness was not significant (Hedges’ g=0.02), while contextual factors such as conversation design and domain explained most of the variance. By 2026 a picture is coming to be shared in which AI’s persuasive power does not uniformly surpass humans but varies greatly with conditions.

Cross-Cutting Issues (Candidate Axes for a Review)

  1. The oscillation of optimism and caution: The structure in which Argyle’s optimism and Bisbee’s critique converse in the same venue (Political Analysis) forms the type of debate in this field. Optimism extended to the individual level (Park 2024, Ashokkumar 2026), and critique deepened into concepts and norms (Lin 2025, Desai 2026).
  2. Non-uniformity of representativeness: The fit of silicon samples is close in the mean but low in variance (Bisbee 2024) and skewed in performance by cultural sphere (Qu & Wang 2024). The range over which simulation holds must be stated conditionally.
  3. From measurement to norms: The current of LLMs as research tools moved from “can they be used” to “what must be validated” (Lin & Zhang 2025, Abdurahman 2025, Desai 2026). Concern over LLM measurement without validation is becoming normalized.
  4. Condition-dependence of persuasiveness: AI’s persuasive power is not uniform but depends on post-training, prompt design, and domain (Hölbling 2025, Argyle 2025). Measurement that avoids both over- and under-estimation is the challenge.
  5. Reproducibility as a shared weakness: The problem of results changing with model updates (Bisbee 2024) and the data-processing fragility signaled by the EEC (Costello 2024) span both research that uses AI and research that studies AI.
  • The core of the silicon-sampling validity debate: Argyle et al. 2023 (origin), Bisbee et al. 2024 (critique).
  • The frontier of individual-level simulation: Park et al. 2024 (grounding in 1,052 people), Ashokkumar/Hewitt et al. 2026 (prediction of experimental results).
  • The normalization of measurement: Ziems et al. 2024 (decomposition of capability), Desai et al. 2026 (survey of validation practice).
  • AI as object: Eloundou et al. 2024 (labor exposure), Costello et al. 2024 (persuasion, with EEC note).
  • ai-research-gaps-abduction — A map of the gaps in AI research found through abduction-2, questioning the very frames these research currents take for granted

References

29 items in total. [要一次検証] (“needs primary verification”) indicates that part of the bibliographic record (body-text reachability, peer-reviewed publication, volume/issue, or author order) is unconfirmed. Links are DOIs or primary/quasi-primary pages whose reachability was confirmed. The internal working ledger is source/review/ai-social-science-research-trends/papers.md.

The LLM Turn in Computational Social Science

  • S01 Grossmann, I., Feinberg, M., Parker, D.G., Christakis, N.A., Tetlock, P.E., & Cunningham, W.A. (2023). AI and the transformation of social science research. Science 380(6650), 1108–1109. https://doi.org/10.1126/science.adi1778
  • S02 Bail, C.A. (2024). Can Generative AI improve social science? PNAS 121(21), e2314021121. https://doi.org/10.1073/pnas.2314021121
  • S03 Ziems, C., Held, W., Shaikh, O., Chen, J., Zhang, Z., & Yang, D. (2024). Can Large Language Models Transform Computational Social Science? Computational Linguistics 50(1), 237–291. https://doi.org/10.1162/coli_a_00502

LLMs as Research Tools and Their Validity Critiques

  • S04 Gilardi, F., Alizadeh, M., & Kubli, M. (2023). ChatGPT outperforms crowd workers for text-annotation tasks. PNAS 120(30), e2305016120. https://doi.org/10.1073/pnas.2305016120
  • S05 Törnberg, P. (2023). ChatGPT-4 Outperforms Experts and Crowd Workers in Annotating Political Twitter Messages with Zero-Shot Learning. arXiv:2304.06588 [peer-reviewed publication [要一次検証]]. https://arxiv.org/abs/2304.06588
  • S06 Pangakis, N., Wolken, S., & Fasching, N. (2023). Automated Annotation with Generative AI Requires Validation. arXiv:2306.00176 [peer-reviewed publication [要一次検証]]. https://arxiv.org/abs/2306.00176
  • S07 Lin, H., & Zhang, Y. (2025). Navigating the Risks of Using Large Language Models for Text Annotation in Social Science Research. Sociological Methods & Research. https://doi.org/10.1177/08944393251366243
  • S08 Halterman, A., & Keith, K.A. (2025). Codebook LLMs: Evaluating LLMs as Measurement Tools for Political Science Concepts. Political Analysis 34(2), 188–204. https://doi.org/10.1017/pan.2025.10017
  • S09 Abdurahman, S., Ziabari, A.S., Moore, A.K., Bartels, D.M., & Dehghani, M. (2025). A Primer for Evaluating Large Language Models in Social-Science Research. Advances in Methods and Practices in Psychological Science 8(2). https://doi.org/10.1177/25152459251325174
  • S10 Desai, M., Card, D., & Jacobs, A.Z. (2026). Validating LLMs in social science: Epistemic threats and emerging norms. arXiv:2607.07915 [peer-reviewed publication [要一次検証]]. https://arxiv.org/abs/2607.07915
  • S11 Heseltine, M., & Clemm von Hohenberg, B. (2024). Large language models as a substitute for human experts in annotating political text. Research & Politics 11(1), 1–10. https://doi.org/10.1177/20531680241236239

Social Agents and Silicon Sampling

  • S12 Park, J.S., O’Brien, J.C., Cai, C.J., Morris, M.R., Liang, P., & Bernstein, M.S. (2023). Generative Agents: Interactive Simulacra of Human Behavior. ACM UIST 2023. https://doi.org/10.1145/3586183.3606763 (full text on ACM DL; bibliography confirmed via arXiv:2304.03442)
  • S13 Argyle, L.P., Busby, E.C., Fulda, N., Gubler, J.R., Rytting, C., & Wingate, D. (2023). Out of One, Many: Using Language Models to Simulate Human Samples. Political Analysis 31(3), 337–351. https://doi.org/10.1017/pan.2023.2
  • S14 Bisbee, J., Clinton, J.D., Dorff, C., Kenkel, B., & Larson, J.M. (2024). Synthetic Replacements for Human Survey Data? The Perils of Large Language Models. Political Analysis 32(4), 401–416. https://doi.org/10.1017/pan.2024.5
  • S15 Dillion, D., Tandon, N., Gu, Y., & Gray, K. (2023). Can AI language models replace human participants? Trends in Cognitive Sciences 27(7), 597–600. https://doi.org/10.1016/j.tics.2023.04.008 (body-text reachability [要一次検証]; DOI matches across multiple bibliographic sources)
  • S16 Lin, Z. (2025). Six Fallacies in Substituting Large Language Models for Human Participants. Advances in Methods and Practices in Psychological Science 8(3), 1–19. https://doi.org/10.1177/25152459251357566
  • S17 Park, J.S., Zou, C.Q., Kamphorst, J., Egan, N., Shaw, A., Hill, B.M., Cai, C., Morris, M.R., Liang, P., Willer, R., & Bernstein, M.S. (2024). LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals (original title “Generative Agent Simulations of 1,000 People”). arXiv:2411.10109 [peer-reviewed status [要一次検証]]. https://arxiv.org/abs/2411.10109
  • S18 Sun, S., Lee, E., Nan, D., Zhao, X., Lee, W., Jansen, B.J., & Kim, J.H. (2024). Random Silicon Sampling: Simulating Human Sub-Population Opinion Using a Large Language Model Based on Group-Level Demographic Information. arXiv:2402.18144 [peer-reviewed publication [要一次検証]]. https://arxiv.org/abs/2402.18144
  • S19 Qu, Y., & Wang, J. (2024). Performance and biases of Large Language Models in public opinion simulation. Humanities and Social Sciences Communications 11, 1095. https://doi.org/10.1057/s41599-024-03609-x (author order/volume [要一次検証])
  • S20 Ashokkumar, A., Hewitt, L., Ghezae, I., & Willer, R. (2026). Large language models can predict the results of social science experiments. Nature. https://doi.org/10.1038/s41586-026-10742-x (volume/author order [要一次検証]; bibliography confirmed via the Stanford AI4PB page)

AI as a Substantive Object of Study (Persuasion, Misinformation, Public Opinion, Labor)

  • S21 Costello, T.H., Pennycook, G., & Rand, D.G. (2024). Durably reducing conspiracy beliefs through dialogues with AI. Science 385(6714), eadq1814. https://doi.org/10.1126/science.adq1814 (Editorial Expression of Concern dated June 2026; EEC original text reachability [要一次検証])
  • S22 Eloundou, T., Manning, S., Mishkin, P., & Rock, D. (2024). GPTs are GPTs: Labor market impact potential of LLMs. Science 384(6702), 1306–1308. https://doi.org/10.1126/science.adj0998
  • S23 Brynjolfsson, E., Li, D., & Raymond, L. (2025). Generative AI at Work. Quarterly Journal of Economics 140(2), 889–942. https://doi.org/10.1093/qje/qjae044
  • S24 Salvi, F., Horta Ribeiro, M., Gallotti, R., & West, R. (2025). On the conversational persuasiveness of GPT-4. Nature Human Behaviour 9(8), 1645–1653. https://doi.org/10.1038/s41562-025-02194-6
  • S25 Argyle, L.P., Busby, E.C., Gubler, J.R., Lyman, A., Olcott, J., Pond, J., & Wingate, D. (2025). Testing theories of political persuasion using AI. PNAS 122(18), e2412815122. https://doi.org/10.1073/pnas.2412815122
  • S26 Bai, H., Voelkel, J.G., Muldowney, S., Eichstaedt, J.C., & Willer, R. (2025). LLM-generated messages can persuade humans on policy issues. Nature Communications 16(1). https://doi.org/10.1038/s41467-025-61345-5
  • S27 Hölbling, L., Maier, S., & Feuerriegel, S. (2025). A meta-analysis of the persuasive power of large language models. Scientific Reports. https://doi.org/10.1038/s41598-025-30783-y
  • S28 Hackenburg, K., et al. (2025). The levers of political persuasion with conversational AI. Science 390(6777). https://doi.org/10.1126/science.aea3884 (Science home-site reachability [要一次検証]; bibliography confirmed via LSE ResearchOnline + arXiv:2507.13919)
  • S29 Lin, H., et al. (2025). Persuading voters using human–artificial intelligence dialogues. Nature 648(8093), 394–401. https://doi.org/10.1038/s41586-025-09771-9

Unverified Items

  • [要一次検証] S05 Törnberg 2023 / S06 Pangakis et al. 2023 / S10 Desai et al. 2026 / S18 Sun et al. 2024: all on arXiv only, with peer-reviewed publication unconfirmed.
  • [要一次検証] S12 Park et al. 2023 (UIST): ACM DL body text not reached (bibliography confirmed via arXiv).
  • [要一次検証] S15 Dillion et al. 2023: publisher body text not reached (DOI matches across multiple bibliographic sources).
  • [要一次検証] S17 Park et al. 2024: peer-reviewed status unconfirmed. Title revision from the original recorded.
  • [要一次検証] S19 Qu & Wang 2024 / S20 Ashokkumar & Hewitt et al. 2026: publisher home sites login-walled, so author order and volume are unconfirmed (S20 bibliography confirmed via the Stanford AI4PB page. Figures differ across routes, 476 effect sizes / 105,165 participants vs. 469 effect sizes / 119,330 participants; this note adopts the former from the directly reached Stanford page).
  • [要一次検証] S21 Costello et al. 2024: Editorial Expression of Concern original (science.org) not reached, confirmed via EurekAlert.
  • [要一次検証] S28 Hackenburg et al. 2025: Science home site not reached (bibliography confirmed via LSE ResearchOnline + arXiv).

Update Policy

This note is a living page. When new empirical studies, meta-analyses, or perspective pieces are obtained, they will be added to the body text and updated will be revised. The [要一次検証] items will be resolved on the internal ledger side (source/review/ai-social-science-research-trends/papers.md) and reflected in this page’s references.


← All Notes · Home