A generative-informed neuro-symbolic framework for syntactic ambiguity resolution: Evidence from Arabic DPs

A generative-informed neuro-symbolic framework for syntactic ambiguity resolution: Evidence from Arabic DPs

一种基于生成式神经符号框架的句法歧义消解方法:来自阿拉伯语名词短语(DPs)的证据

Abstract: Syntactic ambiguity poses a persistent challenge for Arabic NLP, particularly in morphologically rich nominal constructions where multiple structural interpretations may be compatible with the same surface sequence.

摘要: 句法歧义是阿拉伯语自然语言处理(NLP)中一个长期存在的挑战,特别是在形态丰富的名词结构中,同一个表层序列往往对应多种可能的结构解释。

This study proposes a generatively informed neuro-symbolic framework for resolving structural ambiguity in Modern Standard Arabic (MSA) DPs. The framework integrates generative syntactic notions with AraBERT by representing ambiguity as a candidate-based decision task in which linguistically motivated alternatives are explicitly constructed and evaluated through candidate-conditioned input representations.

本研究提出了一种基于生成式神经符号框架的方法,用于消解现代标准阿拉伯语(MSA)名词短语(DPs)中的结构歧义。该框架将生成句法概念与 AraBERT 相结合,将歧义表示为一种基于候选对象的决策任务,通过候选条件输入表示,显式地构建并评估具有语言学依据的替代方案。

Findings indicate that the model achieved 96.88% accuracy, 95.92% macro-F1, 96.83% weighted F1, and 93.94% binary F1 on the unseen evaluation set. Class-level analysis revealed asymmetric performance, with recall of 99.71% for High/VP Attachment (N1) and 89.26% for Low/NP/Embedded Attachment (N2), indicating greater difficulty in recovering the embedded interpretation.

研究结果表明,该模型在未见过的评估集上达到了 96.88% 的准确率、95.92% 的宏观 F1 值、96.83% 的加权 F1 值以及 93.94% 的二元 F1 值。类级分析显示出性能的不对称性:高位/动词短语依附(N1)的召回率为 99.71%,而低位/名词短语/嵌入式依附(N2)的召回率为 89.26%,这表明恢复嵌入式解释的难度更大。

The study concludes that formal syntactic representations can be operationalized within Transformer-based NLP as an explicit interface between linguistic structure and contextual neural modeling, providing a controlled and interpretable approach to Arabic syntactic ambiguity resolution and beyond.

本研究得出结论,形式句法表示可以在基于 Transformer 的 NLP 模型中被操作化,作为语言结构与上下文神经建模之间的显式接口,从而为阿拉伯语句法歧义消解及相关领域提供了一种可控且可解释的方法。