Large Language Models for Low-Resource Languages: A Conceptual Framework for an Electronic Explanatory Dictionary of the Tajik Language
Large Language Models for Low-Resource Languages: A Conceptual Framework for an Electronic Explanatory Dictionary of the Tajik Language
面向低资源语言的大语言模型:塔吉克语电子释义词典的概念框架
Abstract: This paper presents a conceptual framework for developing an electronic explanatory dictionary of the Tajik language using large language models (LLMs). The relevance of the work stems from the absence of a comprehensive digital lexicographic resource for Tajik that is comparable in functionality to dictionaries for high-resource languages, and from the limited adaptation of modern natural language processing technologies to low-resource language systems.
摘要: 本文提出了一个利用大语言模型(LLM)开发塔吉克语电子释义词典的概念框架。该研究的现实意义在于,目前塔吉克语缺乏功能上可与高资源语言词典相媲美的综合性数字词典资源,且现代自然语言处理技术在低资源语言系统中的应用也十分有限。
Based on a systematic survey of existing linguistic, statistical, and corpus resources, we propose a dictionary architecture that integrates modules for morphological analysis, lemmatization, semantic clustering, and dictionary entry generation using LLMs. The choice of subword tokenization is justified by the agglutinative nature of Tajik morphology and its high morphological variability, along with a parameter-efficient fine-tuning (PEFT) strategy suitable for limited annotated data.
基于对现有语言学、统计学和语料库资源的系统调查,我们提出了一种词典架构,该架构集成了形态分析、词形还原、语义聚类以及利用大语言模型生成词条的模块。选择子词分词(subword tokenization)是基于塔吉克语黏着语的形态特征及其高度的形态变化,同时采用了适用于有限标注数据的参数高效微调(PEFT)策略。
The novelty of the work lies in proposing the first holistic conceptual architecture of an explanatory dictionary for Tajik that unifies classical lexicographic methods, language statistics, and generative capabilities of LLMs into a single system. The practical significance of the study is the formation of a methodological foundation for developing a full-featured electronic dictionary that can serve both as a lexicographic tool and as a core resource for machine translation, automatic summarization, sentiment analysis, and other applied NLP tasks.
本研究的创新之处在于提出了首个塔吉克语释义词典的整体概念架构,将传统的词典编纂方法、语言统计学和大语言模型的生成能力统一到一个系统中。该研究的实际意义在于为开发功能齐全的电子词典奠定了方法论基础,使其既能作为词典编纂工具,也能作为机器翻译、自动摘要、情感分析及其他应用型自然语言处理任务的核心资源。
The paper is intended for specialists in computational linguistics, lexicography, and developers of natural language processing systems working with low-resource languages.
本文旨在为计算语言学、词典编纂领域的专家,以及从事低资源语言自然语言处理系统开发的工程师提供参考。