Classifying Interpretive Canons at the Sentence Level: A Benchmark from the German Federal Constitutional Court

Classifying Interpretive Canons at the Sentence Level: A Benchmark from the German Federal Constitutional Court

句子级解释准则分类:来自德国联邦宪法法院的基准测试

Abstract: Judicial reasoning remains challenging for large language models (LLMs) to analyze. This paper contributes a sentence-level benchmark for evaluating the ability of LLMs to classify interpretive canons as articulated by Larenz in the tradition of Savigny. 摘要: 司法推理对于大语言模型(LLM)的分析而言仍然具有挑战性。本文贡献了一个句子级的基准测试,用于评估大语言模型对萨维尼(Savigny)传统下由拉伦茨(Larenz)所阐述的解释准则进行分类的能力。

Our contributions are threefold. First, we operationalize this conception of interpretation as classification criteria. Second, we provide a dataset of decisions of the German Federal Constitutional Court annotated at the sentence level. Third, we report baseline evaluations of four LLMs from three model families under expert hand-written prompts, compared against prompts optimized with Genetic-Pareto (GEPA). 我们的贡献主要有三点。首先,我们将这种解释概念操作化为分类标准。其次,我们提供了一个在句子层面标注的德国联邦宪法法院判决数据集。第三,我们报告了来自三个模型家族的四种大语言模型在专家手写提示词下的基准评估结果,并将其与通过遗传帕累托(GEPA)优化的提示词进行了对比。

Mean F1 over the seven binary subtasks clusters between 70.4 and 79.2 across models, with grammatical interpretation usually the easiest canon to identify and systematic interpretation usually the hardest; under the tested configuration, GEPA-optimized prompts do not systematically outperform the hand-written ones, suggesting that the expert prompts provide a meaningful baseline. 在七项二元子任务中,各模型的平均 F1 分数集中在 70.4 到 79.2 之间。其中,语法解释通常是最容易识别的准则,而体系解释通常是最难识别的;在所测试的配置下,GEPA 优化的提示词并未系统性地优于专家手写提示词,这表明专家提示词提供了一个有意义的基准。