This article analyzes the philological, methodological, and algorithmic limitations of digital resources and text corpora in the Uzbek language. The study reveals how the exact character-based matching principles used in national information retrieval systems, the lack of “fuzzy matching” models, and the absence of robust morphological analyzers lead to a significant “silent loss” of scientific data. Furthermore, the paper provides a scientific justification for the necessity of implementing linguistic annotation based on international TEI P5 and LAF standards, as well as the urgent need to develop a customized national TEI schema for the Uzbek language.
Methodological Problems and Structural Solutions in Preparing Digitalized Texts for Scientific Analysis
DOI:
Keywords:
Abstract
References
Baeza-Yates R., Ribeiro-Neto B. Modern Information Retrieval: The Concepts and Technology behind Search. 2nd ed. – Harlow: Addison-Wesley, 2011. – 913 p.
Gavrilova T.A., Horoshevskij V.F. Bazy znanij intellektual'nyh sistem. – SPb.: Piter, 2000. – 384 s.
Ide N., Romary L. International Standard for a Linguistic Annotation Framework // Natural Language Engineering. – 2004. – Vol. 10, № 3–4. – P. 211–225.
Kettunen K. Keep, Change or Delete Setting up a Low Resource OCR Post-correction Framework for a Digitized Old Finnish Newspaper Collection // Proceedings of the 11th Italian Research Conference on Digital Libraries (IRCDL). – Bolzano, 2015. – P. 95–103.
Manning C., Schütze H. Foundations of Statistical Natural Language Processing. – Cambridge, MA: MIT Press, 1999. – 680 p.
Sperberg-McQueen C.M., Burnard L. TEI P5: Guidelines for Electronic Text Encoding and Interchange. – Charlottesville: TEI Consortium, 2007. – 1230 p.
Zaharov V.P., Bogdanova S.Ju. Korpusnaja lingvistika: uchebnik. 2-e izd. – SPb.: Izd-vo S.-Peterb. un-ta, 2013. – 148 s.