Large language models (LLMs) are increasingly used to classify, label, summarize, and interpret large text collections, creating new possibilities for corpus linguistics. Their capacity for zero-shot and few-shot instruction following could reduce the cost of linguistic annotation and extend analysis beyond the categories handled by conventional part-of-speech taggers, parsers, and dictionary-based tools. At the same time, LLM outputs are probabilistic, prompt-sensitive, model-dependent, and potentially biased, raising fundamental questions about measurement validity, annotation reliability, and reproducibility. This article critically synthesizes foundational corpus-annotation principles with recent evidence on LLM-based text annotation and develops a validated human–LLM workflow for corpus research. The framework distinguishes token-, span-, sentence-, document-, and discourse-level annotation; requires a human-coded gold sample before large-scale deployment; treats prompt design as part of the annotation manual; and evaluates accuracy, precision, recall, F1, inter-annotator agreement, stability across repeated runs, subgroup performance, and error types. Recent studies show that LLMs can approach or exceed crowd-worker performance on some well-specified classification tasks, but that performance varies substantially across datasets, languages, models, prompts, text lengths, and annotation complexity. Span-level annotation and context-dependent semantic or pragmatic coding remain particularly challenging. The article therefore argues against unvalidated full automation and proposes selective automation, disagreement-based human adjudication, model/version documentation, and preservation of raw outputs. For corpus linguistics, the strongest near-term use of LLMs is as flexible annotators within a transparent, theory-driven, and auditable pipeline rather than as replacements for linguistic expertise. The resulting framework supports scalable corpus annotation while preserving the empirical principles on which corpus-based linguistic inference depends.
Using Large Language Models for Automated Corpus Annotation and Linguistic Analysis: A Critical Methodological Framework
DOI:
Keywords:
Abstract
References
Bhat, S., & Varma, V. (2023). Large language models as annotators: A preliminary evaluation for annotating low-resource language content. In Proceedings of the 4th Workshop on Evaluation and Comparison of NLP Systems (pp. 100–107). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.eval4nlp-1.8
Biber, D. (1988). Variation across speech and writing. Cambridge University Press.
Biber, D., Conrad, S., & Reppen, R. (1998). Corpus linguistics: Investigating language structure and use. Cambridge University Press.
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., et al. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877–1901.
Calderon, N., Reichart, R., & Dror, R. (2025). The Alternative Annotator Test for LLM-as-a-Judge: How to statistically justify replacing human annotators with LLMs. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 16051–16081). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.acl-long.782
Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37–46. https://doi.org/10.1177/001316446002000104
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of NAACL-HLT 2019 (pp. 4171–4186). Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1423
Gilardi, F., Alizadeh, M., & Kubli, M. (2023). ChatGPT outperforms crowd-workers for text-annotation tasks. Proceedings of the National Academy of Sciences, 120(30), e2305016120. https://doi.org/10.1073/pnas.2305016120
Gries, S. T. (2009). Quantitative corpus linguistics with R: A practical introduction. Routledge.
Heseltine, M., & von Hohenberg, B. C. (2024). Large language models as a substitute for human experts in annotating political text. Research & Politics, 11(1), 1–10. https://doi.org/10.1177/20531680241236239
Kasner, Z., Zouhar, V., Schmidtová, P., Kartáč, I., Onderková, K., Platek, O., Gkatzia, D., Mahamood, S., Dusek, O., & Balloccu, S. (2026). LLMs as span annotators: A comparative study of LLMs and humans. In Proceedings of the First Workshop on Multilingual Multicultural Evaluation (pp. 1–22). Association for Computational Linguistics. https://doi.org/10.18653/v1/2026.mme-main.1
Kristensen-McLachlan, R. D., Canavan, M., Kárdos, M., Jacobsen, M., & Aarøe, L. (2025). Are chatbots reliable text annotators? Sometimes. PNAS Nexus, 4(4), pgaf069. https://doi.org/10.1093/pnasnexus/pgaf069
Krippendorff, K. (2018). Content analysis: An introduction to its methodology (4th ed.). SAGE.
Leech, G. (2005). Adding linguistic annotation. In M. Wynne (Ed.), Developing linguistic corpora: A guide to good practice (pp. 17–29). Oxbow Books.
McEnery, T., & Hardie, A. (2012). Corpus linguistics: Method, theory and practice. Cambridge University Press. https://doi.org/10.1017/CBO9780511981395
Nasuto, A., Iacus, S., Rowe, F., et al. (2026). Large language models identify immigration attitudes in online discourse regardless of language. Scientific Reports, 16, 20118.
Nivre, J., de Marneffe, M.-C., Ginter, F., Hajič, J., Manning, C. D., Pyysalo, S., Schuster, S., Tyers, F., & Zeman, D. (2020). Universal Dependencies v2: An evergrowing multilingual treebank collection. In Proceedings of the Twelfth Language Resources and Evaluation Conference (pp. 4034–4043). European Language Resources Association.
Pangakis, N., Wolken, S., & Fasching, N. (2023). Automated annotation with generative AI requires validation. arXiv. https://doi.org/10.48550/arXiv.2306.00176
Reiss, M. V. (2023). Testing the reliability of ChatGPT for text annotation and classification: A cautionary remark. arXiv. https://doi.org/10.48550/arXiv.2304.11085
Sinclair, J. (2005). Corpus and text: Basic principles. In M. Wynne (Ed.), Developing linguistic corpora: A guide to good practice (pp. 1–16). Oxbow Books.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30.
Volkanovska, E. (2025). Large language models as annotators of named entities in climate change and biodiversity: A preliminary study. In Proceedings of the 1st Workshop on Ecology, Environment, and Natural Language Processing (pp. 24–33). University of Tartu Library.
Wang, A., Morgenstern, J., & Dickerson, J. P. (2025). Large language models that replace human participants can harmfully misportray and flatten identity groups. Nature Machine Intelligence, 7, 400–411.
Wynne, M. (Ed.). (2005). Developing linguistic corpora: A guide to good practice. Oxbow Books.
Zhang, R., Li, Y., Ma, Y., Zhou, M., & Zou, L. (2023). LLMaAA: Making large language models as active annotators. In Findings of the Association for Computational Linguistics: EMNLP 2023 (pp. 13088–13103). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.findings-emnlp.872
Ziems, C., Held, W., Shaikh, O., Chen, J., Zhang, Z., & Yang, D. (2024). Can large language models transform computational social science? Computational Linguistics, 50(1), 237–291. https://doi.org/10.1162/coli_a_00502