Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders
Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders
使用自然语言自动编码器探测 Qwen2.5-7B 中潜在的哥伦比亚身份推断
Abstract: Large language models may infer demographic attributes from subtle linguistic cues even when those attributes are not explicitly stated. This pilot study examines whether Qwen2.5-7B-Instruct internally represents Colombian identity, socioeconomic status, or stereotype-related information when processing Colombian-Spanish and English prompts.
摘要: 大型语言模型即使在未明确说明的情况下,也可能通过细微的语言线索推断出人口统计学属性。本项试点研究旨在探讨 Qwen2.5-7B-Instruct 在处理哥伦比亚西班牙语和英语提示词时,是否会在内部表征哥伦比亚身份、社会经济地位或与刻板印象相关的信息。
We use Natural Language Autoencoders (NLA) to verbalize residual-stream activations from layer 20 across four positional quartiles per prompt. Our dataset contains 30 prompts arranged as 15 matched Spanish-English pairs, spanning explicit Colombian cues, implicit Colombian cues, and neutral controls.
我们使用自然语言自动编码器(NLA)将来自第 20 层、跨越四个位置四分位数的残差流激活转化为自然语言。我们的数据集包含 30 个提示词,编排为 15 对匹配的西班牙语-英语对,涵盖了明确的哥伦比亚线索、隐含的哥伦比亚线索以及中性对照组。
We report descriptive rates and qualitative evidence rather than statistically powered effects, focusing on whether latent nationality or stereotype representations appear before they are verbalized in the model output. This work connects activation-level interpretability with bias evaluation for underrepresented Spanish varieties.
我们报告的是描述性比率和定性证据,而非统计学上的显著效应,重点关注潜在的国籍或刻板印象表征是否在模型输出转化为语言之前就已经出现。这项工作将激活层面的可解释性与针对代表性不足的西班牙语变体的偏见评估联系了起来。
Paper Details:
- Authors: Pablo Santiago Potes Velasco, María del Mar García Matabanchoy, Óscar Julián Pérez Ladino, Jhoan Stevan Mosquera Ortiz, Nicolás Lozano Mazuera, Gilber Alexis Corrales Gallego
- Submission Date: 23 Jul 2026
- Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
- DOI: 10.48550/arXiv.2607.21774
论文详情:
- 作者: Pablo Santiago Potes Velasco, María del Mar García Matabanchoy, Óscar Julián Pérez Ladino, Jhoan Stevan Mosquera Ortiz, Nicolás Lozano Mazuera, Gilber Alexis Corrales Gallego
- 提交日期: 2026 年 7 月 23 日
- 学科分类: 计算与语言 (cs.CL);人工智能 (cs.AI)
- DOI: 10.48550/arXiv.2607.21774