On Measuring Semantic Preservation in Legal Ontology Learning

On Measuring Semantic Preservation in Legal Ontology Learning

关于衡量法律本体学习中语义保持的研究

Abstract: Ontology learning transforms unstructured text into structured representations for automated reasoning. Yet structuring information risks losing it, and current evaluation methodologies cannot detect such loss, focusing on structural correctness while failing to measure whether meaning survives transformation.

摘要: 本体学习(Ontology learning)将非结构化文本转换为用于自动推理的结构化表示。然而,对信息进行结构化处理存在丢失信息的风险,而当前的评估方法无法检测到这种损失,因为它们侧重于结构正确性,却未能衡量意义在转换过程中是否得以保留。

We propose an evaluation methodology that addresses this: comparing LLM task performance on source documents against performance on transformed representations, with the difference quantifying semantic loss. We demonstrate this approach on legal merger agreement analysis, a domain chosen for its complex language and precise semantic requirements, comparing direct LLM application against three ontology learning methods across six language models.

我们提出了一种解决该问题的评估方法:通过比较大语言模型(LLM)在原始文档上的任务表现与在转换后表示上的表现,利用两者之间的差异来量化语义损失。我们在法律并购协议分析领域演示了该方法,选择该领域是因为其语言复杂且对语义要求精确;我们对比了直接应用 LLM 与三种本体学习方法在六种语言模型上的表现。

The results reveal systematic semantic loss with significant variation based on reasoning complexity and model-method interactions. Our contributions are: (1) an evaluation framework for measuring semantic preservation in ontology learning, and (2) empirical evidence that semantic loss varies dramatically with model-method pairing, providing guidance for selecting optimal configurations in legal knowledge systems.

结果显示,系统性的语义损失确实存在,且根据推理复杂度和模型-方法交互作用的不同,损失程度存在显著差异。我们的贡献包括:(1)一个用于衡量本体学习中语义保持的评估框架;(2)实证证据表明,语义损失随模型与方法的配对而剧烈变化,这为法律知识系统中选择最优配置提供了指导。