Hallucination as a Feature, not a Defect: Evaluating a multi-agent architecture to transform speculative language-model outputs into testable scientific hypotheses
Hallucination as a Feature, not a Defect: Evaluating a multi-agent architecture to transform speculative language-model outputs into testable scientific hypotheses
幻觉是特性而非缺陷:评估一种将大模型推测性输出转化为可验证科学假设的多智能体架构
Abstract: 摘要:
Contemporary Large Language Models (LLMs) are increasingly aligned to suppress hallucinations, prioritizing factual retrieval over combinatorial creativity. While crucial for mitigating misinformation, this alignment may also restrict speculative Research and Development (R&D) by encouraging what this work operationally treats as semantic overfitting and diversity collapse. 当代大语言模型(LLM)正日益通过对齐(Alignment)来抑制幻觉,优先考虑事实检索而非组合式创造力。虽然这对减少错误信息至关重要,但这种对齐方式也可能通过诱导本文所定义的“语义过拟合”和“多样性崩溃”,从而限制推测性的研发(R&D)工作。
In this paper, we propose a Rust-based multi-agent orchestration that uses the contrast between narrative daydreaming and executive control as a functional analogy, not as a neurocognitive claim. The system instigates an Epistemological Friction loop between a high-entropy generating agent and a web-grounded evaluating agent, mediated by a low-entropy semantic bottleneck intended to reduce noise and repetition. 在本文中,我们提出了一种基于 Rust 的多智能体编排系统,该系统利用“叙事白日梦”与“执行控制”之间的对比作为功能类比,而非神经认知层面的主张。该系统在“高熵生成智能体”与“基于网络检索的评估智能体”之间建立了一个认识论摩擦循环,并通过一个旨在减少噪声和重复的“低熵语义瓶颈”进行中介调节。
Initial experiments generated diverse, viability-rated hypotheses across physical and social-science domains. We additionally report an exploratory paired baseline and ablation study comparing the full system against direct prompting, self-reflection, removal of the semantic filter, removal of search grounding, and removal of lateral lenses. 初步实验在物理学和社会科学领域生成了多样化且具备可行性评估的假设。此外,我们还报告了一项探索性的配对基准测试和消融研究,将完整系统与直接提示(Direct Prompting)、自我反思(Self-reflection)、移除语义过滤器、移除搜索基础(Search Grounding)以及移除横向视角(Lateral Lenses)进行了对比。
The results place direct prompting among the weakest conditions across most observed metrics, but they do not show a general superiority of the full system over simple self-reflection. Instead, they suggest that each architecture shifts the balance between originality, feasibility, diversity, and empirical grounding in different ways, and that the full system provides its main advantages when hypotheses must survive strong physical, empirical, or institutional constraints. 结果显示,在大多数观测指标中,直接提示法表现最弱,但完整系统并未表现出对简单自我反思的全面优越性。相反,研究表明每种架构都在原创性、可行性、多样性和实证基础之间以不同方式调整了平衡;完整系统的主要优势在于当假设必须经受严格的物理、实证或制度约束时,其表现更为稳健。
These findings do not show that hallucination is useful in isolation; they suggest that speculative generation gains value only when constrained by architecture, empirical grounding, and explicit evaluation. 这些发现并不意味着幻觉本身是有用的;它们表明,推测性生成只有在受到架构、实证基础和明确评估的约束时,才具有实际价值。