The Multilingual Quantization Tax: Structural Collapse and Typological Fragility in Edge SLMs
The Multilingual Quantization Tax: Structural Collapse and Typological Fragility in Edge SLMs
多语言量化税:边缘侧小语言模型(SLM)的结构性崩溃与类型学脆弱性
Abstract: While 4-bit weight quantization is critical for deploying Small Language Models (SLMs) on edge devices, evaluations of the resulting performance degradation—the quantization tax—remain overwhelmingly English-centric.
摘要: 虽然 4-bit 权重量化对于在边缘设备上部署小语言模型(SLM)至关重要,但针对由此产生的性能下降(即“量化税”)的评估目前仍主要集中在英语领域。
We present a zero-shot multilingual evaluation of 4-bit quantization across the Gemma 4 and Qwen 3.5 architectures. Evaluating on eight typologically diverse languages using MMLU ProX Lite and GlobalPIQA, we show parameter truncation exposes deep pre-training inequalities.
我们针对 Gemma 4 和 Qwen 3.5 架构,对 4-bit 量化进行了零样本(zero-shot)多语言评估。通过使用 MMLU ProX Lite 和 GlobalPIQA 对八种类型学差异显著的语言进行测试,我们发现参数截断暴露了预训练过程中深层次的不平等性。
We identify four phenomena: (1) Typological Fragility: low-resource and specific non-Latin scripts suffer representational collapse via architecture-specific double dissociations, failing to generate valid task logits; (2) Home Language Fragility Paradox: foundational pre-training pathways provide limited precision loss protection; (3) Domain-Specific Forgetting: multi-step cross-lingual routing degrades while associative soft-science recall remains robust; and (4) Quantization Resistance: highly saturated, typologically aligned domains resist deterministic degradation, with post-quantization performance gains bounded by statistical noise.
我们识别出四种现象:(1) 类型学脆弱性:低资源语言和特定的非拉丁语系脚本通过架构特有的“双重解离”(double dissociations)遭受表征崩溃,导致无法生成有效的任务 Logits;(2) 母语脆弱性悖论:基础预训练路径对精度损失的保护作用有限;(3) 特定领域遗忘:多步跨语言路由能力下降,而关联性社会科学知识的召回能力依然稳健;(4) 量化抗性:高度饱和且类型学对齐的领域能够抵御确定性退化,量化后的性能增益受限于统计噪声。