Notes · updated 2026-07-29
The Genesis of LLM-as-a-Judge: A Confluence of Three Lineages and Four Demands
Scope and Method
This note is a genesis-history literature note that organizes from which lineages the concept of LLM-as-a-Judge (the technique of using an LLM as an evaluator of other generated outputs) emerged, and by which demands it was pushed forward until it took hold. The research currents after the concept’s establishment (the three generations of judges, bias research, creativity evaluation) are covered by llm-as-a-judge-literature; this note covers the stage that precedes it.
The corpus extends the existing source/review/llm-as-a-judge/.
This extension comprises 38 academic publications (C1–C38; 38 explored, 0 duplicates, 0 excluded) and 10 primary sources from vendors and projects (P1–P10); the core publications of the formative period (MT-Bench, G-Eval, Chatbot Arena, and others) are referenced by the IDs of the existing corpus (the A series).
The covered period has its prehistory anchor in 2002, with the main focus on 2016–2024.
Rather than restricting to the creative domain, the scope covers the whole of NLG evaluation, including machine translation, summarization, dialogue, and alignment.
Prehistory: The Double Crisis of NLG Evaluation
The prehistory of LLM-as-a-Judge can be read as a process in which both automatic evaluation and human evaluation lost trust.
Distrust of Surface Metrics
BLEU for machine translation (Papineni et al. 2002) and ROUGE for summarization (Lin 2004) established automatic evaluation that measures quality by n-gram overlap with reference texts, and remained the standard of generation-system research for two decades. Distrust, however, had accumulated from early on. Callison-Burch et al. (2006) showed that improvements in BLEU do not imply improvements in actual translation quality, and Liu et al. (2016) demonstrated at scale that surface metrics are almost uncorrelated with human judgment in dialogue response generation. Novikova et al. (2017) confirmed weak correlations across metrics in end-to-end NLG and stated explicitly the need for new metrics, and Reiter (2018), in a meta-analysis of 284 correlations, limited the validity of BLEU to “diagnostics of MT systems.” Mathur et al. (2020) went as far as pointing out that the meta-evaluation protocol for evaluating metrics is itself flawed (correlations distorted by outliers). Distrust of surface metrics had deepened in stages, from individual counterexamples to criticism of the methodology as a whole.
Human Evaluation Offered No Refuge Either
Simply returning to human evaluation turned out not to be an option. Howcroft et al. (2020), in a meta-analysis of 165 papers spanning 20 years, showed that more than 200 quality terms proliferate in human evaluation and that methods are not standardized. The ReproGen shared task (Belz et al. 2021) made the reproducibility of human evaluation itself the object of measurement and quantified the difficulty of reproduction. Problems also surfaced in crowdsourcing. Karpinska et al. (2021) demonstrated problems in the quality and reproducibility of Mechanical Turk evaluation, and Clark et al. (2021) showed that non-experts can barely distinguish GPT-3-generated text from human writing. As the fluency of generative models began to exceed evaluators’ discriminative capacity, the very premise that “human evaluation is the gold standard” was shaken.
Learned Metrics as an Intermediate Stage
On the automatic-evaluation side, an intermediate stage intervened: learning the metric itself. BERTScore (Zhang et al. 2020) replaced surface overlap with the similarity of contextual embeddings; BLEURT (Sellam et al. 2020), through fine-tuning on human evaluation data, and COMET (Rei et al. 2020), through multilingual quality prediction, each brought into practical use “metrics learned to approximate human judgment.” BARTScore (Yuan et al. 2021), which uses generation probability itself as the score, and UniEval (Zhong et al. 2022), with its Boolean QA format, moved still closer to the generative model as evaluator. GPTScore (Fu et al. 2023) is the terminus of this line and became the direct precursor of the prompted judge, specifying evaluation criteria through natural-language instructions alone. Learned metrics established that “evaluation can be learned,” but dependence on reference texts and evaluation training data remained. The field had come to one step short of discarding that dependence.
Adjacent Precursors: Judges Inside the Training Loop
A separate stream, apart from evaluation research, lies in preference learning and alignment research.
RLHF and Reward Models
The reward model (a judge learned from human preference comparisons that returns the goodness of a model output as a scalar) was formulated by Christiano et al. (2017) for training RL agents. Ziegler et al. (2019) and Stiennon et al. (2020) applied it to the generation quality of language models, and InstructGPT (Ouyang et al. 2022) brought it into practical use as the three-stage pipeline of SFT, reward model, and PPO. At this point, the structure in which “another model grades a model’s outputs” already existed inside the training loop. LLM-as-a-Judge can be positioned as taking this judge out of the training loop and repurposing it for evaluation in general. The continuity between the two also shows in the fact that the meta-evaluation of reward models (RewardBench) and the meta-evaluation of judges later came to be discussed within the same framework.
AI Feedback
The step of replacing the reward model’s training data from humans to AI was taken by Constitutional AI (Bai et al. 2022). The model critiques and revises its own outputs based on a set of principles (a constitution), and a preference model is trained from AI preference labels. Anthropic also published the set of principles itself (2023-05), indicating a direction of making AI judgment criteria externally verifiable. RLAIF by Lee et al. (2023) demonstrated through scale comparison that AI feedback can reach performance on par with human feedback, giving empirical grounding to the claim that “the evaluator may be an AI.”
Scalable Oversight
The lineage on the normative side is scalable oversight (the problem setting of whether oversight can be maintained even when AI capability exceeds human evaluative capacity). Amodei et al. (2016) formulated “scalable oversight” as a concrete problem of AI safety; Irving et al. (2018) proposed a structure in which humans referee debates between AIs, and Leike et al. (2018) proposed recursive reward modeling. Empirical work had also begun. Saunders et al. (2022) showed that model self-critique helps human evaluators find flaws, and Bowman et al. (2022) experimented on whether non-experts can evaluate specialized tasks with LLM assistance. OpenAI, in its research direction of August 2022, stated explicitly that “RLHF has a fundamental limitation in its premise that humans can accurately evaluate tasks,” placing AI-assisted evaluation on its official research agenda. In other words, “AI evaluating AI” was, prior to the 2023 boom, a research direction normatively demanded within alignment research.
The Establishment of the Concept (2023)
Practice Preceded the Term
In the first half of 2023, the evaluation-research lineage and the preference-learning lineage converged in the post-ChatGPT situation. On the academic side, GPTScore (arXiv 2302.04166) showed evaluation design via instructions, GEMBA (Kocmi & Federmann 2023) surpassed learned metrics in machine translation evaluation, Wang et al. (2023) systematically measured ChatGPT’s aptitude as an NLG evaluator across 5 datasets, and Chiang & Lee (2023) asked in their very title whether LLMs can be an alternative to human evaluations. G-Eval (Liu et al. 2023) raised human correlation with CoT and form-filling.
In parallel, practice ran ahead of academia. The Vicuna blog of March 30, 2023, which claimed model quality by having GPT-4 grade answers to 80 questions, is the first widely known example of the practice of using GPT-4 as an evaluator. Although the authors themselves stated the limitation that the method was “not yet rigorous or mature,” the figure “90% of ChatGPT” took on a life of its own (separated as a marketing-claim in the corpus). In the same month, OpenAI released the evaluation framework Evals alongside GPT-4, incorporating model-graded evals, in which a model grades outputs, from the very beginning. Around May, AlpacaFarm (Dubois et al. 2023) implemented evaluation 50 times cheaper than crowdworkers through preference simulation with GPT-4-family APIs, and AlpacaEval began operating that evaluator as a leaderboard. Chatbot Arena of the same period (announced 2023-05-03) conversely placed human voting at its core, building the infrastructure for collecting the human preference data against which automatic judges would be checked.
The Term Takes Hold
It was Zheng et al. (2023) who put the term “LLM-as-a-judge” in a title and established it as a concept. Published in June 2023, the paper used MT-Bench and Chatbot Arena to show that a GPT-4 judge agrees with human preference at over 80%, while enumerating from the outset the limitations of position bias, verbosity bias, and self-enhancement bias. Within the range of this corpus, earlier publications referred to the same practice with terms such as “evaluator” and “automatic evaluation”; the strict first occurrence of the term cannot be identified, but it can be confirmed that subsequent surveys (Gu et al. 2024 and others) attribute the establishment of the paradigm to this paper. The order in which practice (Vicuna in March) preceded the term (the MT-Bench paper in June) indicates that this concept was not a product of theory but a retrospective naming of field practice pressed by the need for evaluation.
The Four Demands That Pushed the Concept Forward
Viewed from the side of demands rather than as a sequence of individual publications, the establishment of the concept can be organized into four.
- The absence of correct answers in open-ended generation: Reference-based automatic metrics presuppose the existence of reference texts, but the outputs of instruction-following chat models (InstructGPT onward) admit no single correct answer. The demonstration that surface metrics were nearly uncorrelated in dialogue (Liu et al. 2016) is an early manifestation of this structural problem. An evaluator that can judge quality without references became necessary, and the only candidates were humans, or models approximating humans.
- Benchmark saturation and contamination: Dynabench (Kiela et al. 2021) problematized the rapid saturation of static benchmarks and their divergence from real-world performance. As the scale of LLM training data expanded, contamination, in which evaluation data leaks into training, was added on top. Sainz et al. (2023) called for standardizing contamination measurement, Zhou et al. (2023) quantified score distortion through deliberate contamination experiments, and Golchin & Surdeanu (2023) built detection tools. The more evaluation against fixed correct answers lost trust, the higher the relative value rose of approaches that judge dynamically generated outputs on the spot.
- The supply-demand gap in evaluation: Human evaluation was not only costly but neither standardized nor reproducible (Howcroft et al. 2020; Belz et al. 2021). The scale of HELM (Liang et al. 2022), 30 models by 42 scenarios, shows that the labor required for systematic evaluation was exceeding what human hands could cover. AlpacaFarm’s “50 times cheaper” was a price quote to fill this supply-demand gap.
- The need for alignment evaluation: The practical deployment of RLHF constantly requires evaluation along a dimension with no correct answer, namely whether outputs are helpful and harmless. The scalable-oversight lineage formulated this as the problem of the limits of human evaluative capacity, and Constitutional AI and RLAIF had already implemented the solution of AI feedback. Seen from the side of alignment research, LLM-as-a-Judge was not a tool borrowed from outside but an apparatus demanded by its own problem setting.
Institutionalization
After the concept took hold, institutionalization advanced from late 2023 through 2026.
On the survey side, Chang et al. (2024, ACM TIST), which systematized LLM evaluation as a whole, positioned LLM-as-a-Judge as one type of evaluation method, and three judge-dedicated surveys (Gu et al. 2024; Li et al. 2024, From Generation to Judgment; Li et al. 2024, LLMs-as-Judges) put taxonomies in place.
On the meta-evaluation side, the apparatus for verifying judges themselves was put in place. JUDGE-BENCH (Bavaresco et al. 2024) verified agreement with human judgment at scale across 20 datasets; RewardBench and JudgeBench provided benchmarks for judges, and Length-Controlled AlpacaEval (Dubois et al. 2024) provided statistical control of bias.
On the platform side, Chatbot Arena became academized as the reference infrastructure for human preference (Chiang et al. 2024), and AlpacaEval and OpenAI Evals were operated as evaluation infrastructure. The static-benchmark Open LLM Leaderboard continued to hold down the contrasting position of “reproducible but vulnerable to saturation and contamination.”
On the standardization side, public institutions took in the concept within less than three years of its emergence. NIST’s AI 800-2 initial draft (2026-01) defines LLM-as-a-judge as one method for automated benchmark evaluation, and UK AISI implemented model-graded scorers in its evaluation framework Inspect. The current state on the institutional side is detailed in the industry sections of llm-as-a-judge-literature.
Why 2023?
Read through the genesis history, the establishment of the concept can be explained as the confluence of three streams.
- The internal circumstances of evaluation research: Surface metrics lost validity between 2016 and 2020, human evaluation could be neither standardized nor reproduced, and learned metrics had established that “evaluation can be learned.” The evaluator’s seat stood vacant.
- The technical foundation: Preference learning had normalized “models that grade model outputs” inside the training loop, and a general-purpose model capable of approximating expert evaluation (GPT-4) appeared in March 2023.
- The normative demand: The scalable-oversight lineage had already legitimized “AI evaluating AI” as a research direction.
Even with these three in place, the trigger was the explosion of open-ended outputs after ChatGPT. When the objects of evaluation (a mass of outputs with no correct answers) were supplied, the evaluator (GPT-4) was supplied, and the justification (scalable oversight) was supplied, all that remained was the naming. The three months from the Vicuna practice (March) to the naming in the MT-Bench paper (June) lie at the center of the concept’s emergence.
This confluence, however, did not solve the problems of the prehistory; it relocated them. The validity problem of surface metrics persists in altered form as the judge bias problem, and the reproducibility problem of human evaluation as the judge meta-evaluation problem. That post-establishment research concentrated on the quantification of bias and meta-evaluation benchmarks (Cluster 2 of llm-as-a-judge-literature) is a consequence of this continuity.
Unverified Items
- The acceptance venues of C17 (Ziegler 2019), C23 (Irving 2018), C24 (Leike 2018), C26 (Bowman 2022), C34 (Golchin & Surdeanu 2023), and C35 (Zhou et al. 2023) (currently recorded as preprints) [requires primary verification]
- The initial release date of the AlpacaEval repository (estimated at around 2023-05 from the bibtex and the Internet Archive) [requires primary verification]
- The release date of OpenAI Evals, 2023-03-14 (per secondary reporting; no explicit date in the repository) [requires primary verification]
- The publication dates and body texts of the three OpenAI blog posts (InstructGPT, Learning to Summarize, Our Approach to Alignment Research) (openai.com returned 403; body texts not retrieved) [requires primary verification]
- The exact initial release date of Open LLM Leaderboard v1 [requires primary verification]
- Whether the term “LLM-as-a-judge” was used before arXiv 2306.05685 (unidentified within the range of this corpus)
References
All accessed 2026-07-29.
Prehistory (The Crisis of NLG Evaluation)
- Papineni, K., Roukos, S., Ward, T., Zhu, W.-J. (2002). BLEU: a Method for Automatic Evaluation of Machine Translation. ACL 2002. https://doi.org/10.3115/1073083.1073135
- Lin, C.-Y. (2004). ROUGE: A Package for Automatic Evaluation of Summaries. Text Summarization Branches Out (ACL Workshop). https://aclanthology.org/W04-1013/
- Callison-Burch, C., Osborne, M., Koehn, P. (2006). Re-evaluating the Role of Bleu in Machine Translation Research. EACL 2006. https://aclanthology.org/E06-1032/
- Liu, C.-W. et al. (2016). How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation. EMNLP 2016. https://doi.org/10.18653/v1/D16-1230
- Novikova, J., Dušek, O., Curry, A. C., Rieser, V. (2017). Why We Need New Evaluation Metrics for NLG. EMNLP 2017. https://doi.org/10.18653/v1/D17-1238
- Reiter, E. (2018). A Structured Review of the Validity of BLEU. Computational Linguistics, 44(3). https://doi.org/10.1162/coli_a_00322
- Mathur, N., Baldwin, T., Cohn, T. (2020). Tangled up in BLEU: Reevaluating the Evaluation of Automatic Machine Translation Evaluation Metrics. ACL 2020. https://doi.org/10.18653/v1/2020.acl-main.448
- Howcroft, D. M. et al. (2020). Twenty Years of Confusion in Human Evaluation: NLG Needs Evaluation Sheets and Standardised Definitions. INLG 2020. https://doi.org/10.18653/v1/2020.inlg-1.23
- Karpinska, M., Akoury, N., Iyyer, M. (2021). The Perils of Using Mechanical Turk to Evaluate Open-Ended Text Generation. EMNLP 2021. https://doi.org/10.18653/v1/2021.emnlp-main.97
- Clark, E. et al. (2021). All That’s ‘Human’ Is Not Gold: Evaluating Human Evaluation of Generated Text. ACL 2021. https://doi.org/10.18653/v1/2021.acl-long.565
- Belz, A., Shimorina, A., Agarwal, S., Reiter, E. (2021). The ReproGen Shared Task on Reproducibility of Human Evaluations in NLG. INLG 2021. https://doi.org/10.18653/v1/2021.inlg-1.24
- Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., Artzi, Y. (2020). BERTScore: Evaluating Text Generation with BERT. ICLR 2020. https://arxiv.org/abs/1904.09675
- Sellam, T., Das, D., Parikh, A. P. (2020). BLEURT: Learning Robust Metrics for Text Generation. ACL 2020. https://doi.org/10.18653/v1/2020.acl-main.704
- Rei, R., Stewart, C., Farinha, A. C., Lavie, A. (2020). COMET: A Neural Framework for MT Evaluation. EMNLP 2020. https://doi.org/10.18653/v1/2020.emnlp-main.213
- Fu, J., Ng, S.-K., Jiang, Z., Liu, P. (2023). GPTScore: Evaluate as You Desire. arXiv:2302.04166. https://arxiv.org/abs/2302.04166
Adjacent Precursors (Preference Learning and Scalable Oversight)
- Christiano, P. et al. (2017). Deep Reinforcement Learning from Human Preferences. NeurIPS 2017. https://proceedings.neurips.cc/paper_files/paper/2017/hash/d5e2c0adad503c91f91df240d0cd4e49-Abstract.html
- Ziegler, D. M. et al. (2019). Fine-Tuning Language Models from Human Preferences. arXiv:1909.08593. https://arxiv.org/abs/1909.08593
- Stiennon, N. et al. (2020). Learning to Summarize from Human Feedback. NeurIPS 2020. https://arxiv.org/abs/2009.01325
- Ouyang, L. et al. (2022). Training Language Models to Follow Instructions with Human Feedback. NeurIPS 2022. https://arxiv.org/abs/2203.02155
- Bai, Y. et al. (2022). Constitutional AI: Harmlessness from AI Feedback. arXiv:2212.08073. https://arxiv.org/abs/2212.08073
- Lee, H. et al. (2023). RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback. ICML 2024. https://arxiv.org/abs/2309.00267
- Amodei, D. et al. (2016). Concrete Problems in AI Safety. arXiv:1606.06565. https://arxiv.org/abs/1606.06565
- Irving, G., Christiano, P., Amodei, D. (2018). AI Safety via Debate. arXiv:1805.00899. https://arxiv.org/abs/1805.00899
- Leike, J. et al. (2018). Scalable Agent Alignment via Reward Modeling: A Research Direction. arXiv:1811.07871. https://arxiv.org/abs/1811.07871
- Saunders, W. et al. (2022). Self-Critiquing Models for Assisting Human Evaluators. arXiv:2206.05802. https://arxiv.org/abs/2206.05802
- Bowman, S. R. et al. (2022). Measuring Progress on Scalable Oversight for Large Language Models. arXiv:2211.03540. https://arxiv.org/abs/2211.03540
The Establishment of the Concept (Around 2023)
- Chiang, C.-H., Lee, H.-y. (2023). Can Large Language Models Be an Alternative to Human Evaluations? ACL 2023. https://doi.org/10.18653/v1/2023.acl-long.870
- Wang, J. et al. (2023). Is ChatGPT a Good NLG Evaluator? A Preliminary Study. NewSum Workshop @ EMNLP 2023. https://doi.org/10.18653/v1/2023.newsum-1.1
- Kocmi, T., Federmann, C. (2023). Large Language Models Are State-of-the-Art Evaluators of Translation Quality (GEMBA). EAMT 2023. https://aclanthology.org/2023.eamt-1.19/
- Dubois, Y. et al. (2023). AlpacaFarm: A Simulation Framework for Methods that Learn from Human Feedback. NeurIPS 2023. https://arxiv.org/abs/2305.14387
- tatsu-lab (2023). AlpacaEval: An Automatic Evaluator of Instruction-Following Language Models. GitHub. https://github.com/tatsu-lab/alpaca_eval
- Zheng, L. et al. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023 Datasets and Benchmarks. https://arxiv.org/abs/2306.05685
- Liu, Y. et al. (2023). G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment. EMNLP 2023. https://aclanthology.org/2023.emnlp-main.153/
The Demands That Pushed the Concept Forward
- Kiela, D. et al. (2021). Dynabench: Rethinking Benchmarking in NLP. NAACL 2021. https://doi.org/10.18653/v1/2021.naacl-main.324
- Sainz, O. et al. (2023). NLP Evaluation in Trouble: On the Need to Measure LLM Data Contamination for Each Benchmark. Findings of EMNLP 2023. https://doi.org/10.18653/v1/2023.findings-emnlp.722
- Golchin, S., Surdeanu, M. (2023). Data Contamination Quiz: A Tool to Detect and Estimate Contamination in Large Language Models. arXiv:2311.06233. https://arxiv.org/abs/2311.06233
- Zhou, K. et al. (2023). Don’t Make Your LLM an Evaluation Benchmark Cheater. arXiv:2311.01964. https://arxiv.org/abs/2311.01964
- Liang, P. et al. (2023). Holistic Evaluation of Language Models (HELM). TMLR. https://arxiv.org/abs/2211.09110
Institutionalization
- Chang, Y. et al. (2024). A Survey on Evaluation of Large Language Models. ACM Transactions on Intelligent Systems and Technology, 15(3). https://doi.org/10.1145/3641289
- Bavaresco, A. et al. (2024). LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks. ACL 2025. https://arxiv.org/abs/2406.18403
Primary Sources (Vendors and Projects)
- LMSYS Org (2023-03-30). Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90%* ChatGPT Quality. https://lmsys.org/blog/2023-03-30-vicuna/
- LMSYS Org (2023-05-03). Chatbot Arena: Benchmarking LLMs in the Wild with Elo Ratings. https://lmsys.org/blog/2023-05-03-arena/
- OpenAI (2023). OpenAI Evals (GitHub). https://github.com/openai/evals
- OpenAI (2022). Aligning Language Models to Follow Instructions (InstructGPT). https://openai.com/index/instruction-following/
- OpenAI (2020). Learning to Summarize with Human Feedback. https://openai.com/index/learning-to-summarize-with-human-feedback/
- Anthropic (2022-12-15). Constitutional AI: Harmlessness from AI Feedback. https://www.anthropic.com/research/constitutional-ai-harmlessness-from-ai-feedback
- Anthropic (2023-05-09). Claude’s Constitution. https://www.anthropic.com/news/claudes-constitution
- OpenAI (2022). Our Approach to Alignment Research. https://openai.com/index/our-approach-to-alignment-research/
- Hugging Face (2023–2024). Open LLM Leaderboard v1 (archive). https://huggingface.co/docs/leaderboards/en/open_llm_leaderboard/archive
References from the Existing Corpus (details in llm-as-a-judge-literature)
- Yuan, W., Neubig, G., Liu, P. (2021). BARTScore: Evaluating Generated Text as Text Generation. NeurIPS 2021. https://proceedings.neurips.cc/paper/2021/hash/e4d2b6e6fdeca3e60e0f1a62fee3d9dd-Abstract.html
- Zhong, M. et al. (2022). Towards a Unified Multi-Dimensional Evaluator for Text Generation (UniEval). EMNLP 2022. https://aclanthology.org/2022.emnlp-main.131/
- Chiang, W.-L. et al. (2024). Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. ICML 2024. https://arxiv.org/abs/2403.04132
- Dubois, Y. et al. (2024). Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators. COLM 2024. https://arxiv.org/abs/2404.04475
- Lambert, N. et al. (2024). RewardBench: Evaluating Reward Models for Language Modeling. Findings of NAACL 2025. https://arxiv.org/abs/2403.13787
- Tan, S. et al. (2024). JudgeBench: A Benchmark for Evaluating LLM-based Judges. ICLR 2025. https://arxiv.org/abs/2410.12784
- Gu, J. et al. (2024). A Survey on LLM-as-a-Judge. arXiv:2411.15594. https://arxiv.org/abs/2411.15594
- Li, D. et al. (2024). From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge. EMNLP 2025. https://arxiv.org/abs/2411.16594
- Li, H. et al. (2024). LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods. arXiv:2412.05579. https://arxiv.org/abs/2412.05579
- NIST (2026). AI 800-2 (Initial Public Draft): Practices for Automated Benchmark Evaluations of Language Models. https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.800-2.ipd.pdf
- UK AI Security Institute. Inspect. https://inspect.aisi.org.uk/