LLM Agents Perform Controlled Experiments Using Simulation Models
LLM Agents Perform Controlled Experiments Using Simulation Models
大语言模型智能体利用仿真模型进行受控实验
Abstract: Large language models (LLMs) have shown strong capabilities in reasoning, planning, and tool use, but many scientific and engineering tasks require more than plausible text and code generation. They require understanding how a system responds to intervention, which in practice depends on controlled experimentation.
摘要: 大语言模型(LLMs)在推理、规划和工具使用方面展现出了强大的能力,但许多科学和工程任务需要的不仅仅是生成看似合理的文本和代码。它们需要理解系统如何对干预做出反应,而这在实践中依赖于受控实验。
In this work, we propose a multi-agent framework that enables LLM agents to conduct controlled experiments with scientific simulation models for pharmaceutical process design. Given a user query and a baseline configuration, the system constructs a structured task representation, designs experiments, executes comparative simulation, interprets the resulting outcomes, and synthesizes evidence-based recommendations for process parameter optimization.
在这项工作中,我们提出了一个多智能体框架,使大语言模型智能体能够利用科学仿真模型进行药物工艺设计的受控实验。给定用户查询和基准配置,该系统能够构建结构化的任务表示、设计实验、执行对比仿真、解读结果,并综合得出基于证据的工艺参数优化建议。
By coupling language models with high-fidelity simulation models in an interactive agent framework, the proposed system supports reasoning through intervention, comparison, and observation. As a result, it produces more specific and actionable outputs than language-only reasoning.
通过在交互式智能体框架中将语言模型与高保真仿真模型相结合,该系统支持通过干预、比较和观察进行推理。因此,它比单纯的语言推理能产生更具体、更具可操作性的输出。
In an industrial application setting, this advantage is reflected in higher output specificity as well as improved user-rated correctness and helpfulness. Ablation studies and visualized case analyses further demonstrate the effectiveness and practical utility of simulation-integrated experimental reasoning.
在工业应用场景中,这一优势体现为更高的输出特异性,以及用户评分中更高的准确性和实用性。消融研究和可视化案例分析进一步证明了仿真集成实验推理的有效性和实际应用价值。
Paper Details:
- Authors: Yuchen Xia, Michael Weyrich, Nasser Jazdi, Johannes Stümpfle, Johannes Sigel, Akshay Narla, Gavin K. Reynolds, Anna Jawor-Baczynska, Pol Llopart
- arXiv ID: 2608.23622
- Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Multiagent Systems (cs.MA); Software Engineering (cs.SE)
论文详情:
- 作者: Yuchen Xia, Michael Weyrich, Nasser Jazdi, Johannes Stümpfle, Johannes Sigel, Akshay Narla, Gavin K. Reynolds, Anna Jawor-Baczynska, Pol Llopart
- arXiv ID: 2608.23622
- 学科分类: 人工智能 (cs.AI);计算与语言 (cs.CL);多智能体系统 (cs.MA);软件工程 (cs.SE)