Is Human-Readable Text Necessary for Effective LLM Fine-Tuning?
Is Human-Readable Text Necessary for Effective LLM Fine-Tuning?
人类可读文本对于大语言模型的高效微调是必要的吗?
Abstract: Is human readability necessary for effective fine-tuning of large language models? We investigate whether model-conditioned training representations can preserve or improve adaptation utility without requiring a human-readable textual form.
摘要: 人类可读性对于大语言模型(LLM)的高效微调是必要的吗?我们研究了模型条件下的训练表征是否可以在无需人类可读文本形式的情况下,保持或提升模型的适应效能。
We propose Desired-Update-Aligned Synthetic Data (DASA), which uses activation-gradient feedback from a frozen reference model to guide the optimization of continuous synthetic input embeddings. Inspired by the role of activation gradients in local risk reduction, DASA targets useful adaptation updates rather than source-text reconstruction or linguistic fluency.
我们提出了“期望更新对齐合成数据”(Desired-Update-Aligned Synthetic Data, DASA),该方法利用冻结参考模型的激活梯度反馈,来指导连续合成输入嵌入(embeddings)的优化。受激活梯度在局部风险降低中作用的启发,DASA 的目标是实现有效的适应性更新,而非源文本重构或语言流畅性。
The resulting embeddings are used directly for downstream fine-tuning; discrete token projections are employed only for qualitative inspection. Experiments on six models from the Llama and Qwen families, ranging from 1B to 32B parameters, cover six benchmarks spanning knowledge, mathematical reasoning, code generation, and commonsense reasoning.
生成的嵌入直接用于下游微调;离散标记投影仅用于定性检查。我们在 Llama 和 Qwen 系列的六个模型上进行了实验,参数规模从 1B 到 32B 不等,涵盖了知识、数学推理、代码生成和常识推理等六个基准测试。
Under matched LoRA adaptation settings, DASA achieves performance comparable to the source natural-language data and surpasses it in multiple configurations, while outperforming GRADMM in most comparisons. Further experiments cover general-domain and task-specialized source data. Under the evaluated synthesis settings, DASA provides a $3.6$—$4.9\times$ speedup over GRADMM with comparable peak GPU memory.
在匹配的 LoRA 适应设置下,DASA 实现了与源自然语言数据相当的性能,并在多种配置下超越了后者,同时在大多数对比中优于 GRADMM。进一步的实验涵盖了通用领域和任务专用源数据。在评估的合成设置下,DASA 相比 GRADMM 实现了 3.6 到 4.9 倍的加速,且峰值 GPU 内存占用相当。