Shuichiro Ogawa
日本語

Notes · updated 2026-07-19

AI and the Humanities: Digital Humanities and Large Language Models (2024–2026)

The humanities have placed reading text at the center of the discipline. When a machine that generates, classifies, and translates text at scale enters that center, what does the field lose, and what can it newly ask?

Reading the peer-reviewed articles and major preprints from 2024 through mid-2026, two flows run at once. One uses the large language model (LLM) as a tool, reaching into sources and scales that were previously out of reach. The other lets the machine’s very presence unsettle the foundational concepts the tool presupposes, namely “meaning,” “author,” and “understanding.” This note follows both, sorted into five currents.

The materials span peer-reviewed articles and major preprints across digital humanities (DH), computational literary studies, history and linguistics, the philosophy of AI, and critical AI studies (19 items; see “References”). Each DOI/arXiv identifier was verified for existence via Crossref or the arXiv abstract page (collection protocol .claude/rules/collection-protocol.md, zero fabrication). The internal working ledger is source/review/ai-humanities-digital-humanities-trends/papers.md (not published). Sister notes: ai-social-science-research-trends, ai-arts-culture-studies-trends, ai-education-learning-sciences-trends. The industrial discourse on art production is left to ai-arts-culture-studies-trends.

Three turning points that characterize the field

The 2024 Poetics Today special issue is one of the earliest collective records of literary theory responding directly to LLMs (H01). The editors asked literary scholars what LLMs mean for authorship, academic writing, pedagogy, and the future of the profession, and central figures such as Hayles and Kirschenbaum contributed. What makes the issue a watershed is that its subject is the discipline’s self-understanding, not the tool’s performance.

In the same issue, Bajohr distinguished artificial texts (text generated by a machine) from post-artificial texts (the stance readers take once machine generation has become normal) (H02). By moving the question from the writer’s technique to the reader’s expectations, this distinction is becoming shared vocabulary for later debate.

On the tool side, a quiet reversal took place. Humphries and colleagues showed that for transcribing 18th- and 19th-century handwritten documents, multimodal LLMs outperform dedicated handwritten text recognition (HTR) software (H12). With a self-correction step, the character error rate (CER) fell to 1.8%, approaching human levels, they report. The claim that a general-purpose model has caught up with and passed specialized models in transcription, long the province of the latter, is prompting DH to reconsider its toolkit.

The LLM turn in digital humanities: scaling transcription and analysis

The bottleneck DH has long faced is the step of turning sources into text. Transcribing handwritten sources demands skill and is hard to scale. The LLM is first breaking this bottleneck.

The Humphries report (H12) is not an isolated case. Greif and colleagues report that for historical German city directories, multimodal LLMs surpassed conventional OCR, and with post-correction the CER dropped below 1% (2025, H14). That this level was reached without image preprocessing or fine-tuning carries practical weight.

Yet the picture that a general model always wins is not accurate. The benchmark by Crosilla and colleagues (2025, H13) shows that while commercial models such as Claude 3.5 Sonnet beat open-source ones in a zero-shot setting, LLMs are limited in their ability to autonomously fix errors in their own transcriptions. Transcription accuracy and error-correction ability are distinct, and the latter still requires human verification.

Beyond transcription, LLMs are entering the step of analyzing text. Stewart and Sinha combined OCR, custom segmentation, and LLM prompting to extract structured metadata from over 45,000 pages of unstructured USDA plant inventory records (2025, H15). Zeng built HistoLens, a framework for multi-layered analysis of historical texts, taking the Western Han “Yantie Lun” as a case (2024, H16). Dobranić and colleagues layered topic modeling with LLM-based sentiment analysis to read how identity and political ideology appeared in early-20th-century Slovene newspapers (2026, H18). Distant reading, with the LLM in hand, extends its reach toward macroscopic patterns invisible to individual interpretation.

Still, this scaling is not unconditional. Stewart and Sinha themselves caution that validating the extracted results and human oversight remain essential (H15). The faster the tool moves through a step, the more the question of where the human verifies comes to the front.

Machine translation and low-resource languages: when safety filters halt the text

The languages the humanities handle include many with few speakers, many that are historically extinct, and many that are non-standardized. The translation capability of LLMs would seem to help precisely with such low-resource and classical languages.

But Tekgurler’s report (2025, H19) breaks the naive expectation with fact. When an 18th-century Ottoman Turkish manuscript is translated with Gemini, the safety mechanism blocks 14% to 23% of the body text, preventing complete and accurate translation of the historical text. A mechanism built to censor contemporary harmful expression halts the work of reading a document from the past. When a general-purpose LLM is brought into the humanities, the design of the model’s “safety” becomes a constraint on access to the source itself.

This single case illuminates a problem invisible from translation-quality metrics alone. In low-resource translation, before asking how fluently a model can translate, one must ask whether it will translate at all.

The philosophy of AI around meaning and understanding

The more fluent the text an LLM produces, the more an old question returns with new weight. Is the model handling the meaning of words, or only the statistics of their form?

The point of departure is the position that, because a model is trained only on form (text), it cannot in principle learn meaning. Grindrod reframes this dispute through the philosophical notion of intentionality, examining whether LLM outputs can be said to possess meaning or intentionality (2024, H03). The symbol grounding problem, the missing circuit that would ground symbols in their referents in the world, re-ignites here.

Shanahan, one of the participants in the dispute, corrected the misreading of his own position as reductionism (2024, H05). He repositions his argument as a Wittgensteinian look at how language is used, not a metaphysical assertion about the nature of LLMs. The debate is shifting from the binary “does the LLM understand or not” toward “in which context is it appropriate to use the word understanding.”

Millière and Buckner systematized these individual disputes into a philosophy of language models (2026, Philosophy Compass, H04). By laying out the topics of meaning, understanding, representation, and grounding, this review gives a map to a scattered debate. The existence of this review shows that the philosophy of AI is beginning to take shape as a subfield with LLMs as its subject.

Authorship, interpretation, and disclosure: humanistic disputes

While the philosophy of meaning is contested in the abstract, the disputes over authorship and rights proceed at the level of institutions. The question of who counts as having written a work forces concrete decisions in both law and scholarly norms.

Ramos-Zaga proposes a normative framework for the minimum threshold of human creativity required for a generative-AI-assisted work to merit copyright protection (2025, H10). He reorganizes, as a question of thresholds, the legal current holding that prompts alone provide insufficient human control.

On the interpretive side, how far LLMs “understand” style was tested empirically. Hicke and Mimno probed how models identify authors and genres, showing that authorial voice is easier to detect than genre, and that mechanisms relying on memorization coexist with those that learn style (2025, H09). O’Sullivan used the classical methods of stylometry to compare human and AI-generated creative writing, measuring where the stylistic fingerprints match and where they diverge (2025, H08). Stylometry brings the discipline’s tradition of treating authorship as measurable features, not sentiment, into the age of generative AI.

There is one more dispute in which the field questions its own footing. How far should the use of generative AI in research be disclosed? The survey by Ma and colleagues (2025, H11) reveals that DH scholars recognize the importance of disclosure while actual disclosure rates remain low, and views on which activities warrant disclosure diverge. A gap has opened between the speed at which the field adopts the new tool and the speed at which norms for making that use visible are set.

Critical AI studies: questioning the outside of the tool

The currents so far all treat the LLM as an object of analysis or as a tool. Critical AI studies steps back a level, questioning the power and materiality of the technical regime of AI itself.

Hogan argued, as humanistic critique, about the environmental burden of generative AI and the extractivist power structure embodied by cloud companies (2024, Critical AI, H06). The focus of critique is placed not on the model’s output but on the infrastructure that runs it.

Seddone argued that as algorithmic knowledge spreads, the humanities take on the role of preserving the human component of knowledge (2025, AI & Society, H07). This is a defense of the humanities, and at the same time a proposal for what the humanities should shoulder in the age of AI. The existence of critical AI studies shows that the field seeks to be not only a user of AI but also a subject that questions it.

Cross-cutting issues (candidate axes for a review)

  1. The LLM as tool and as object of critique: the current that uses LLMs for transcription and analysis (H12/H14/H15/H16/H18) runs alongside the current that questions their meaning, power, and materiality (H03/H04/H06/H07). Both are different faces of the same field.
  2. The accuracy reversal, and its conditions: the claim that general LLMs beat dedicated HTR (H12/H14) is strong, but comes with conditions: weak self-correction (H13) and the need for human verification (H15). It should not be read as an unconditional performance claim.
  3. The tension between safety filters and source access: safety design keyed to contemporary standards obstructs the work of reading past texts (H19). A friction specific to bringing general-purpose models into the humanities.
  4. The re-ignition and systematization of the philosophy of meaning: the old disputes of symbol grounding and intentionality (H03/H05) return with LLMs and are mapped as a philosophy of language models (H04).
  5. The institutional lag in authorship and disclosure: the creativity threshold (H10), style detection (H08/H09), and disclosure of generative AI use (H11) proceed unsettled on both the legal and scholarly-norm fronts.
  • Grasping the turning points: H01 (Poetics Today special-issue introduction), H02 (artificial/post-artificial texts).
  • The conditions of the tool’s limits: H13 (weak self-correction in HTR), H19 (safety-filter blocking in translation).
  • The map of the philosophy: H04 (philosophy of language models, review), H03 (intentionality).
  • Institutional disputes: H11 (the reality of disclosure), H10 (the copyright threshold).

References

19 items in total. Peer-reviewed articles by DOI, preprints by arXiv. [preprint] marks a pre-peer-review preprint; [to be verified] marks partly unconfirmed bibliographic details (not affecting the main facts). The internal ledger is source/review/ai-humanities-digital-humanities-trends/papers.md.

Turning points, literary theory, authorship

  • H01 Evron, N., & Tartakovsky, R. (2024). The AI Revolution: Speculations on Authorship, Pedagogy, and the Future of the Profession. Poetics Today 45(2), 189–195. https://doi.org/10.1215/03335372-11092765
  • H02 Bajohr, H. (2024). On Artificial and Post-artificial Texts: Machine Learning and the Reader’s Expectations of Literary and Non-literary Writing. Poetics Today 45(2), 331–361. https://doi.org/10.1215/03335372-11092990
  • H08 O’Sullivan, J. (2025). Stylometric comparisons of human versus AI-generated creative writing. Humanities and Social Sciences Communications 12. https://doi.org/10.1057/s41599-025-05986-3
  • H09 Hicke, R.M.M., & Mimno, D. (2025). Looking for the Inner Music: Probing LLMs’ Understanding of Literary Style. arXiv:2502.03647 [preprint]. https://arxiv.org/abs/2502.03647
  • H10 Ramos-Zaga, F.A. (2025). Reconceptualizing Human Authorship in the Age of Generative AI: A Normative Framework for Copyright Thresholds. Laws 14. https://doi.org/10.3390/laws14060084

Philosophy of AI (meaning, understanding, intentionality)

Critical AI studies and the role of the humanities

Digital humanities (transcription, OCR, analysis)

  • H12 Humphries, M., Leddy, L.C., Downton, Q., Legace, M., McConnell, J., Murray, I., & Spence, E. (2024). Unlocking the Archives: Using Large Language Models to Transcribe Handwritten Historical Documents. arXiv:2411.03340 [preprint]. https://arxiv.org/abs/2411.03340
  • H13 Crosilla, G., Klic, L., & Colavizza, G. (2025). Benchmarking Large Language Models for Handwritten Text Recognition. arXiv:2503.15195 [preprint]. https://arxiv.org/abs/2503.15195
  • H14 Greif, G., Griesshaber, N., & Greif, R. (2025). Multimodal LLMs for OCR, OCR Post-Correction, and Named Entity Recognition in Historical Documents. arXiv:2504.00414 [preprint]. https://arxiv.org/abs/2504.00414
  • H15 Stewart, S.D., & Sinha, S. (2025). Retrieving information from unstructured historical sources using large language models. Computational Humanities Research 1. https://doi.org/10.1017/chr.2025.10019
  • H16 Zeng, Y. (2024). HistoLens: An LLM-Powered Framework for Multi-Layered Analysis of Historical Texts — A Case Application of Yantie Lun. arXiv:2411.09978 [preprint]. https://arxiv.org/abs/2411.09978
  • H17 Yang, Z., et al. (2024). Analyzing Nobel Prize Literature with Large Language Models. arXiv:2410.18142 [preprint]. https://arxiv.org/abs/2410.18142
  • H18 Dobranić, F., et al. (2026). Approaches to Analysing Historical Newspapers Using LLMs. arXiv:2603.25051 [preprint]. https://arxiv.org/abs/2603.25051

Machine translation (historical and low-resource languages)

  • H19 Tekgurler, M. (2025). LLMs for Translation: Historical, Low-Resourced Languages and Contemporary AI Models. arXiv:2503.11898 [preprint]. https://arxiv.org/abs/2503.11898

Items to be verified

  • H04 (Millière & Buckner, Philosophy Compass) shows volume 21 / 2026 in Crossref. A discrepancy between early-view and the finalized volume remains, hence [to be verified] (publication itself confirmed).
  • H12 (Humphries et al.): the figures “CER 1.8% / WER 3.5%” and “50x faster, ~1/50 the cost” are the authors’ experimental claims; independent replication is unconfirmed [to be verified]. Treated here as the authors’ report.
  • H10 (Ramos-Zaga, Laws/MDPI) is peer-reviewed, but given debate over the rigor of that publisher’s review, confidence is medium. It is a normative framework proposal, with limited factual claims.
  • H18 (arXiv:2603.25051) and some other preprints carry a 2026 submission year. The arXiv ID, title, and authors are confirmed via the abstract page; the peer-reviewed venue is undetermined.

Update policy

This note is a living page. As new peer-reviewed articles and major preprints appear, they will be added to the body and the updated date revised. On the internal ledger side, [to be verified] markers will be resolved and reflected in the references here.


← All Notes · Home