A survey detection channel overrides the pixels in an astronomical foundation model, and biases tomographic mean redshifts
A survey detection channel overrides the pixels in an astronomical foundation model, and biases tomographic mean redshifts
巡天检测通道覆盖了天文基础模型中的像素,并导致层析平均红移偏差
Abstract: Foundation models for astronomy are trained on survey pixels together with the catalogue products derived from those pixels. Those catalogues are incomplete at a measurable rate, and a model trained on both inherits that incompleteness as a systematic.
摘要: 天文基础模型通常基于巡天像素以及从这些像素中导出的目录产品进行训练。这些目录存在可测量的不完整性,而同时使用这两者进行训练的模型会继承这种不完整性,并将其转化为系统性误差。
We audit AION-1, a 39-modality transformer trained on more than 200 million objects, using causal interventions on its inputs. Holding the image tokens byte-identical and editing only the survey segmentation map changes every quantity the model reports — flux, size, ellipticity, redshift — by 110-4400 times a matched placebo.
我们对 AION-1(一个在超过 2 亿个天体上训练的 39 模态 Transformer 模型)进行了审计,并对其输入进行了因果干预。在保持图像 Token 完全一致的情况下,仅编辑巡天分割图,就会导致模型报告的每一个量(如通量、大小、椭圆率、红移)发生 110 到 4400 倍于对照组的变化。
The mechanism is detection gating, presence at the field centre (r = 0.47), not the light the mask encloses (r = 0.30); across 322 real blends the model ignores how the pipeline partitioned the light (R = -0.006). Nor is the preference specific to that channel: contradicted catalogue photometry leaves the model nine times worse than supplying no metadata at all.
其机制在于检测门控(detection gating),即天体在视场中心的存在性(r = 0.47),而非掩模所包含的光度(r = 0.30);在 322 个真实混合天体样本中,模型完全忽略了流水线是如何划分光度的(R = -0.006)。这种偏好并非仅限于该通道:当目录测光数据与图像矛盾时,模型表现比不提供任何元数据时还要差 9 倍。
The Legacy Survey pipeline leaves 3.68% of targets with no segment covering their position. Propagating that rate, with a miss represented by the fields the pipeline actually returns, shifts tomographic mean redshifts by a median 0.71 times the LSST DESC requirement over 40 assignments and exceeds it in 12; observed positional errors take the worst bin to 8.3 times.
Legacy Survey 流水线导致 3.68% 的目标天体在其位置上没有对应的分割区域。将这一比率进行传播,并以流水线实际返回的视场作为缺失样本,结果显示层析平均红移在中位数上偏移了 LSST DESC 要求的 0.71 倍(共 40 次分配),其中 12 次超过了该要求;在观测位置误差的影响下,最差的区间偏移达到了要求的 8.3 倍。
Drawing the misses by their measured magnitude dependence rather than uniformly does not change it. Spectroscopy removes the effect, withholding the detection channel removes it at no measurable cost, and the effect grows with model scale.
根据测量的星等依赖性而非均匀分布来抽取缺失样本,并不会改变这一结果。光谱数据可以消除这种影响,剔除检测通道也能在不产生可测量代价的情况下消除该影响,且这种效应会随着模型规模的扩大而增强。
Two further limits lie in the tokeniser: its image codec resolves 28 effective states on source patches against 934 for the spectrum codec, and the redshift readout is quantisation-limited. Sparse dictionaries are unreliable causal handles: across 15, recovery spans 26-75% and moves up to 18 points on the seed alone.
该模型还存在两个局限性:其图像编解码器在源补丁上仅能解析出 28 个有效状态,而光谱编解码器可解析 934 个;此外,红移读出受到量化限制。稀疏字典作为因果分析手段是不可靠的:在 15 次实验中,恢复率在 26% 到 75% 之间波动,且仅随机种子一项就能导致高达 18 个点的变动。