Shuichiro Ogawa
日本語

Notes · updated 2026-07-19

Research Currents in AI and Economics: Productivity, Labor Markets, and Methods (2024–2026)

A synthesis of 16 peer-reviewed articles and major working papers on the intersection of generative AI and economics, collected via lightweight scoping. Full bibliographic entries are in the “References” section below (DOI/URL, traceable). The internal ledger with provenance and confidence ratings lives at source/review/ai-economics-research-trends/papers.md (repository-internal, unpublished). For the design-specific intersection see ai-economics-and-design; for adjacent fields see ai-law-governance-research-trends and ai-social-science-research-trends.

The same tool was handed to everyone, and the results did not converge. That is the single point the generative-AI research of the past few years keeps making. In call centers, in writing tasks, among micro-entrepreneurs in Kenya, average productivity rises. But open up the distribution and who gains and who loses is decided on the side of the user, not the tool.

What economics turned toward over these years is not the question “is AI useful?” Usefulness is taken as given; the question moved to how that effect is distributed, how much it lifts macro growth, and where in the labor market it cuts. This note follows that shift across five currents.

Productivity field experiments and the unevenness of the effect

The first turning point was a series of randomized experiments measuring the productivity effect of generative AI. Brynjolfsson, Li, and Raymond (E01, QJE 2025) analyzed the staggered rollout of an AI assistant to 5,172 customer-support agents as a quasi-experiment. Issues resolved per hour rose by about 15% on average. The breakdown matters: less-experienced, lower-skilled workers improved in both speed and quality, while the most skilled saw small speed gains and a slight decline in quality. What AI lifted was not the average but the lower part of the distribution.

Noy and Zhang (E02, Science 2023) had shown the same picture earlier in a near-lab setting. They assigned occupation-specific writing tasks to 453 college-educated professionals and gave ChatGPT to half. Time taken fell 40%, output quality rose 18%, and inequality between workers narrowed. The finding that weaker-skilled participants benefited most was read as evidence that AI might compress productivity inequality.

Coding produced a similar number. In Peng et al. (E03, arXiv 2023), a group using GitHub Copilot completed an HTTP-server task 55.8% faster (95% CI [21%, 89%], p=.0017). Completion rates did not differ; only speed compressed.

So far this reads as “AI lifts everyone from the bottom.” Two studies that pushed into harder work broke that story. Dell’Acqua et al. (E04, HBS/SSRN 2023) ran a field experiment with 758 BCG consultants. Inside the frontier of AI capability, the AI group completed 12.2% more tasks, 25.1% faster, at much higher quality. Outside that frontier, the AI group scored 19 percentage points worse than the non-AI group. The authors called this boundary the “jagged technological frontier,” where apparent difficulty and actual AI-fitness diverge.

The unevenness was starkest in a developing-country setting. Otis et al. (E05, SSRN 2024) randomly gave GPT-4-based business advice to 640 Kenyan entrepreneurs over five months. The average treatment effect was zero. Yet high-performing entrepreneurs improved by about 20% while low performers did about 10% worse. The difference lay not in the advice the AI returned but in which pieces they chose to implement. The already-strong used AI as a lever; the struggling asked for help on harder problems and missed.

Labor-market exposure and the first cut to employment

If productivity rises, how does employment move? Research built up in two layers.

One layer measures how exposed occupations are to AI. Eloundou et al. (E06, Science 2024) scored O*NET tasks with a GPT-4 rubric and estimated that about 1.8% of jobs have over half their tasks affected by simple LLMs, rising to about 46% once complementary software is included. This is exposure potential, not an employment forecast; what actually happens is a separate matter.

The other layer began measuring the “actual.” Hui, Reshef, and Zhou (E07, Organization Science 2024) tracked online labor markets after ChatGPT and found that freelancers in highly affected occupations, such as writing and translation, lost both employment and earnings. Past success offered no protection; top performers were hit disproportionately.

Yet national administrative data reverse the picture. Humlum and Vestergaard (E08, NBER 2025) linked Danish administrative records with adoption surveys and reported near-zero average effects of chatbots on earnings and hours (confidence intervals ruling out effects beyond roughly ±1%). Time savings were about 3%. Employers absorbed AI as added tasks (content generation, oversight, integration) rather than as changes large enough to move average pay or hours.

These two results are not a contradiction but a difference in where you look. Online self-employment is the most thinly protected layer, and the cut appeared there first. Where employment protection and room for reallocation are large, the change dispersed into task reorganization and stayed hidden in the average. The answer changes not by which is true but by which layer and which metric you observe.

Brynjolfsson, Chandar, and Chen (E09, Stanford 2025) sliced out that most fragile layer by age. Using high-frequency U.S. payroll data, they showed a roughly 13% relative decline in employment for early-career workers (ages 22–25) in highly exposed occupations, after controlling for firm-level shocks. Where AI complemented labor, employment stayed stable or grew. The authors likened young workers to canaries in a coal mine.

Macro and growth: the modest estimate

If individual productivity moves this much, how much does the whole economy grow? The most-cited estimate here is Acemoglu (E10, NBER 2024). Using a task-based model and Hulten’s theorem, he estimated total-factor-productivity (TFP) gains of no more than 0.53–0.66% over ten years, roughly 0.064% annually. The basis is two narrowings. Only about 20% of U.S. labor tasks are exposed to AI, and among these only about 23% can be profitably automated. The large effect sizes from the lab and the modest macro number are reconciled here.

Beneath this estimate lies the task-based theory of Acemoglu and Restrepo (E11, JEP 2019). Automation carries a displacement effect, where capital replaces labor in existing tasks, while the creation of new tasks in which labor holds a comparative advantage brings a reinstatement effect that restores labor demand. This “displacement or reinstatement” decomposition is applied directly to generative AI. Which dominates depends on the character of the technology, and from this comes the caveat that labor demand can fall even as productivity rises.

LLMs as a research tool: both subject and experimenter

Generative AI became not only an object of analysis but a change to the method of economics itself. Korinek (E12, JEL 2023) classified LLM uses across six areas (ideation, writing, background research, data analysis, coding, mathematical derivations), rating capabilities from experimental to highly useful. It is a practical account of how economists can gain substantial productivity by automating micro-tasks.

More ambitious is the attempt to use LLMs as experimental subjects. Horton, Filippas, and Manning (E13, NBER 2023) treated LLMs as “Homo silicus,” implicit computational models of humans, and reproduced classic economic experiments. Given endowments, information, and preferences, the LLMs produced results qualitatively close to the originals, and where they diverged, the divergence became fuel for further research. Manning, Zhu, and Horton (E14, NBER 2024) automated this further, combining structural causal models with LLMs to generate and test social-science hypotheses (negotiation, bail hearing, job interview, auction). This arrangement, with the LLM as both subject and experimenter, raises the validity question anew. How far the regularities shown by silicon subjects can be trusted as proxies for real humans remains an open question.

Markets and competition: algorithms that learn to collude

The last current is what price-setting algorithms do to competition. Calvano et al. (E15, AER 2020) showed by simulation that Q-learning pricing algorithms converge to supracompetitive prices without communicating with one another. High prices were sustained by strategies with a finite punishment phase and gradual return to cooperation. That collusion can emerge without explicit communication posed a new problem for competition policy.

Real-market data showed similar signs. Assad et al. (E16, JPE 2024) analyzed the German retail gasoline market around 2017, when algorithmic pricing software became widely available. Margin increases from adoption did not occur in monopoly markets, only in non-monopoly ones. In duopoly markets, margins rose about 28% when both stations adopted. The theoretical implication that algorithms can tacitly learn collusive strategies was supported in real data.

Cross-cutting questions

The questions running through the five currents reduce to a few.

First, acceleration and unevenness advance together. Average productivity rises, but its distribution depends strongly on the user’s initial ability, so both a lift from the bottom (E01/E02) and a widening of the gap (E04/E05) are observed. Which one appears depends on whether the task lies inside or outside AI’s fitness (E04).

Second, the employment effect changes with the observation point. In thinly protected online markets employment falls (E07/E09); in countries with room for reallocation the average does not move (E08). This is not a contradiction but a matter of which layer and which metric are in view.

Third, the reconciliation of lab effect sizes with macro modesty. Large per-task improvements (E01–E03), passed through the two narrowings of exposed-task share and profitable automation, converge to TFP of no more than 0.66% over ten years (E10).

Fourth, the validity of LLMs as a research method. How far silicon subjects (E13/E14) hold up as proxies for real humans is an unresolved question that reopens the premises of empirical economics.

If only one is left open, it should be the last. Research that measures generative AI as an object of analysis has thickened over these years. But the validity of using generative AI as a research tool is not yet established. Whether the measuring instrument can be trusted anticipates the reliability of what is measured.

References

16 entries in total. Unpublished working papers (arXiv, SSRN, NBER, Stanford Digital Economy Lab) carry reservations on confidence; the external validity of effect sizes is stated in the body. Links are DOIs or canonical landing pages. The internal ledger is source/review/ai-economics-research-trends/papers.md.

A. Productivity field / lab experiments

B. Labor economics (exposure, employment, wages)

C. Macro, growth, automation theory

D. LLMs as a research tool

  • E12 Korinek, A. (2023). Generative AI for Economic Research: Use Cases and Implications for Economists. Journal of Economic Literature 61(4), 1281–1317. https://doi.org/10.1257/jel.20231736
  • E13 Horton, J.J., Filippas, A., & Manning, B.S. (2023). Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus? NBER Working Paper 31122 [WP]. https://www.nber.org/papers/w31122
  • E14 Manning, B.S., Zhu, K., & Horton, J.J. (2024). Automated Social Science: Language Models as Scientist and Subjects. NBER Working Paper 32381 [WP]. https://www.nber.org/papers/w32381

E. Markets, competition, algorithmic pricing

  • E15 Calvano, E., Calzolari, G., Denicolò, V., & Pastorello, S. (2020). Artificial Intelligence, Algorithmic Pricing, and Collusion. American Economic Review 110(10), 3267–3297. https://doi.org/10.1257/aer.20190623
  • E16 Assad, S., Clark, R., Ershov, D., & Xu, L. (2024). Algorithmic Pricing and Competition: Empirical Evidence from the German Retail Gasoline Market. Journal of Political Economy 132(3), 723–771. https://doi.org/10.1086/726906

← All Notes · Home