CARAT: Do Materials LLMs Reason or Recite?
CARAT: Do Materials LLMs Reason or Recite?
CARAT:材料大语言模型是在推理还是在背诵?
Abstract: When a materials LLM answers a question about crystal structure, does it reason from the structure or copy an answer already printed in its input? Accuracy cannot tell: a structural description often prints the very field it is scored against.
摘要: 当材料大语言模型(LLM)回答有关晶体结构的问题时,它是根据结构进行推理,还是仅仅复制了输入中已有的答案?准确率无法说明这一点:因为结构描述中往往直接包含了评分所依据的字段。
CARAT holds question and gold answer fixed across eight matched views, names each structural relation separately in GraphSpace, and adds matched fine-tuning, answer masking, evidence injection, paired inference, and a rule that can withhold claims.
CARAT 在八种匹配视图中保持问题和标准答案不变,在 GraphSpace 中分别命名每个结构关系,并增加了匹配微调、答案掩码、证据注入、配对推理以及一项可以拒绝回答的规则。
First, on the benchmark’s hardest families the grounded view is worth 17.3 points over formula inputs. Second, we turn that scrutiny on ourselves. GraphSpace beats a plain periodic graph by 19.3 points, but that margin is two effects at once: where the plain rendering carries everything the question needs it is 1.96 points, and where it omits those fields entirely, 46.7 points. The headline mostly measures what the baseline lacked, not how evidence is presented.
首先,在基准测试中最困难的类别中,基于事实(grounded)的视图比公式输入高出 17.3 分。其次,我们对自己进行了审视。GraphSpace 比普通的周期图高出 19.3 分,但这一差距同时包含了两种效应:在普通渲染包含问题所需全部信息的情况下,差距仅为 1.96 分;而在完全省略这些字段的情况下,差距则高达 46.7 分。这一总分主要衡量的是基准模型缺失了什么,而不是证据是如何呈现的。
Third, we attack our own benchmark. A rule that skips the link and reads the list directly answers four of seven hardened families, so we rebuilt it until eleven such shortcuts sat near chance. The frozen model quotes that link yet answers the same when we redirect it, on 95.6% of paired cases: it repeats the relation without using it. After matched supervision it reaches 99.8%, and deleting the link drops it to 23.4%, below the 27.0% the best shortcut reaches: both steps are learnable.
第三,我们对自己的基准测试进行了“攻击”。我们发现,一条跳过链接直接读取列表的规则可以回答七个强化类别中的四个,因此我们对其进行了重建,直到十一个此类“捷径”的准确率接近随机水平。在 95.6% 的配对案例中,冻结的模型在重定向后仍引用该链接并给出相同答案:它只是在重复该关系,而没有真正使用它。经过匹配监督后,准确率达到 99.8%,而删除链接后准确率降至 23.4%,低于最佳捷径所能达到的 27.0%:这证明这两个步骤都是可学习的。