Explaining GAND: A Resource on Gender-Ambiguous Natural Data & Contrastive Attribution
Explaining GAND: A Resource on Gender-Ambiguous Natural Data & Contrastive Attribution
解读 GAND:关于性别模糊自然数据与对比归因的资源
Abstract: Machine translation (MT) systems continue to produce gender-biased translations. In a time where self-expression is paramount, mistranslations based on default behaviour and stereotyping can lead to harm for users of these systems. To better understand how these systems translate gender in the absence of clear gender cues, we need benchmarking resources that reflect gender-ambiguous scenarios in a natural way.
摘要: 机器翻译(MT)系统持续产生带有性别偏见的翻译结果。在自我表达至关重要的时代,基于默认行为和刻板印象的错误翻译可能会对系统用户造成伤害。为了更好地理解这些系统在缺乏明确性别线索时如何处理性别翻译,我们需要能够以自然方式反映性别模糊场景的基准测试资源。
To this end, we present GAND, a gender-ambiguous natural data benchmarking resource for MT consisting of English source sentences, specifically designed to analyse the influence of contextual cues on gender in translation. We leverage GAND to conduct an interpretability analysis: we translate a subset of GAND into two grammatical gender languages and extend these with manually crafted contrastive translations.
为此,我们提出了 GAND,这是一个用于机器翻译的性别模糊自然数据基准测试资源,由英语源句子组成,专门用于分析上下文线索对翻译中性别的影响。我们利用 GAND 进行可解释性分析:我们将 GAND 的一个子集翻译成两种具有语法性别的语言,并辅以人工编写的对比翻译进行扩展。
A following feature attribution analysis reveals source words in context that inform the gender translation of an ambiguous referent entity in the target translation.
随后的特征归因分析揭示了语境中的哪些源词汇影响了目标翻译中模糊指代实体的性别翻译。
Paper Details:
- Authors: Janiça Hackenbuchner, Jasper Degraeuwe, Arda Tezcan, Joke Daems
- Subject: Computation and Language (cs.CL)
- arXiv ID: 2607.22546
- Submission Date: 6 May 2026
论文详情:
- 作者: Janiça Hackenbuchner, Jasper Degraeuwe, Arda Tezcan, Joke Daems
- 学科: 计算与语言 (cs.CL)
- arXiv ID: 2607.22546
- 提交日期: 2026年5月6日