Open-sourcing AstaBrief, the fast report-generation model in Asta

Open-sourcing AstaBrief, the fast report-generation model in Asta

开源 AstaBrief:Asta 中快速生成报告的模型

Language models can already help researchers search the literature, synthesize evidence, and work through complex questions. But scientific work places particular demands on these models—answers need to stay grounded in evidence, the models need to preserve what the evidence actually supports rather than quietly broadening a study’s conclusions, and researchers need to be able to verify the final outputs. 语言模型已经能够帮助研究人员搜索文献、综合证据并处理复杂问题。但科学工作对这些模型提出了特殊要求——答案必须基于证据,模型需要准确呈现证据所支持的内容,而不是悄悄扩大研究结论,且研究人员必须能够验证最终输出。

We see that in how scientists use Asta, our agentic platform for scientific work. Instead of simple keyword searches, users often bring substantial context and many constraints—for example, asking Asta to compare approaches across a body of literature while accounting for a particular method, population, or setting. Many also return to generated reports later, treating them as working research artifacts rather than one-off answers. 我们在科学家使用 Asta(我们用于科学工作的智能体平台)的过程中观察到了这一点。用户往往不只是进行简单的关键词搜索,而是提供大量的背景信息和约束条件——例如,要求 Asta 在考虑特定方法、人群或环境的情况下,对文献中的各种方法进行比较。许多用户还会回头查看生成的报告,将其视为研究工作产物,而非一次性的答案。

We wanted to help scientists generate cited reports faster, with a model they could download and run themselves. To do that, we tested whether a small, open model trained specifically for scientific report generation could match the report quality of the proprietary models we were using, while reducing generation time and serving costs. 我们希望帮助科学家更快地生成带有引用的报告,并提供一个他们可以下载并自行运行的模型。为此,我们测试了一个专门针对科学报告生成而训练的小型开源模型,看它是否能在降低生成时间和服务成本的同时,达到我们所使用的专有模型的报告质量。

We built AstaBrief 8B, a model that turns a research question and retrieved literature excerpts into a cited report. AstaBrief is available in Asta’s Generate a report feature today as Fast mode alongside Claude-powered Thinking mode, and we’re also open-sourcing it and the training data so others can study, reproduce, and build on our approach. 我们构建了 AstaBrief 8B,这是一个能将研究问题和检索到的文献片段转化为带引用报告的模型。AstaBrief 现已在 Asta 的“生成报告”功能中作为“快速模式”(Fast mode)提供,与 Claude 驱动的“思考模式”(Thinking mode)并列。我们同时将其模型权重和训练数据开源,以便他人研究、复现并在此基础上进行开发。

Developing AstaBrief required tens of thousands of real research queries, citation-focused filtering, preference data, and a redesigned report-generation pipeline that writes the full report in one pass rather than section by section. The result is nearly an order-of-magnitude reduction in report generation time compared to the proprietary models we tracked—across the full Asta pipeline, Fast mode averages 51.1 seconds per report compared with 178.5 seconds for Thinking mode, about 3.5× faster. 开发 AstaBrief 需要数以万计的真实研究查询、以引用为重点的过滤、偏好数据,以及重新设计的报告生成流程——该流程能一次性写出完整报告,而不是分章节撰写。结果显示,与我们追踪的专有模型相比,报告生成时间缩短了近一个数量级——在整个 Asta 流程中,快速模式平均每份报告耗时 51.1 秒,而思考模式为 178.5 秒,速度提升了约 3.5 倍。

Together, those efficiency gains made AstaBrief a useful test case for a broader goal: building open language models that can be adapted to the specific demands of scientific work. Open weights will also let institutions run AstaBrief on their own infrastructure, which is necessary when research questions reveal sensitive or unpublished work. Alongside the model weights, we’re releasing an example workflow that researchers can adapt to create reports from their own PDFs, providing a starting point for local report generation. 这些效率提升使 AstaBrief 成为实现更宏大目标的有用测试案例:构建能够适应科学工作特定需求的开源语言模型。开放权重还将允许机构在自己的基础设施上运行 AstaBrief,这在研究问题涉及敏感或未发表的工作时尤为必要。除了模型权重,我们还发布了一个示例工作流,研究人员可以根据自己的 PDF 文件进行调整以创建报告,为本地报告生成提供了一个起点。

This post covers how we trained AstaBrief, what we learned about grounding it in scientific evidence, and which parts of our approach we think can carry forward to future models for science. Most of the training and evaluation described was completed in 2025, so the proprietary models used to generate training data and as comparison points reflect the frontier at the time. We haven’t rerun the full evaluation against today’s frontier models; the results below are best read as evidence about the particular training and system design choices we tested. 本文介绍了我们如何训练 AstaBrief,我们在将其基于科学证据进行锚定方面学到了什么,以及我们认为哪些方法可以延续到未来的科学模型中。文中描述的大部分训练和评估工作于 2025 年完成,因此用于生成训练数据和作为对比点的专有模型反映了当时的前沿水平。我们尚未针对当今的前沿模型重新进行全面评估;以下结果最好被解读为我们所测试的特定训练和系统设计选择的证据。

Training the model: Our goal with AstaBrief was to build an open-weights model with all the qualities that matter most for long-form scientific synthesis: answer quality, relevance, structure, and citation grounding. We started from Qwen3-8B and focused most of our effort on the post-training data, evaluation, and surrounding report-generation scaffolding. 训练模型:我们开发 AstaBrief 的目标是构建一个开源权重模型,使其具备长篇科学综述最重要的品质:答案质量、相关性、结构和引用锚定。我们从 Qwen3-8B 开始,将大部分精力集中在训练后数据、评估以及配套的报告生成框架上。

Adapting general-purpose models for scientific work – and training new scientific models from scratch – is something we’re exploring broadly across Ai2. Through NSF OMAI, a U.S. national initiative led by Ai2 to build fully open AI infrastructure and models for scientific discovery, our researchers are working directly with scientific communities to understand what they need from future open models and where today’s general-purpose models fall short. That includes studying how needs differ across scientific fields and workflows, with more findings from that research to share in the future. 将通用模型应用于科学工作——以及从零开始训练新的科学模型——是我们正在 Ai2 全面探索的方向。通过由 Ai2 领导的美国国家倡议 NSF OMAI(旨在构建完全开放的科学发现 AI 基础设施和模型),我们的研究人员正直接与科学界合作,了解他们对未来开源模型的需求,以及当今通用模型的不足之处。这包括研究不同科学领域和工作流之间的需求差异,未来我们将分享更多相关研究成果。

Recent work, including our DR Tulu, has shown that reinforcement-learning-based (RL) methods can improve long-form report generation for open-weights models, especially when judge models are involved in the training loop. We considered that path for AstaBrief, but ultimately focused on a simpler recipe built around supervised fine-tuning (SFT) and direct preference optimization (DPO). RL-based training can be unstable and expensive. We wanted to see how far we could push report generation quality with a cheaper, more operationally manageable setup—one that’s also easier to debug and iterate on. 近期的工作(包括我们的 DR Tulu)表明,基于强化学习(RL)的方法可以改善开源权重模型的长篇报告生成效果,特别是在训练循环中引入判别模型时。我们曾考虑过 AstaBrief 采用这种路径,但最终选择了围绕监督微调(SFT)和直接偏好优化(DPO)构建的更简单的方案。基于 RL 的训练可能不稳定且昂贵。我们希望通过一种更廉价、更易于操作的设置,看看能将报告生成质量提升到什么程度——这种设置也更容易调试和迭代。

That made the quality of the training data especially important. Rather than relying on a more complex optimization method to compensate for noisy examples, we spent much of the project figuring out how to generate, select, and filter examples that actually demonstrated the report-writing behavior we wanted. We also wanted AstaBrief to be faster so that users could get preliminary reports quickly that they could then iterate over in subsequent turns. For speed improvements, we decided to train AstaBrief to directly generate the final report in one pass given a user query and relevant retrieved snippets, bypassing the expensive snippet summarization and clustering stages our Claude-based Thinking mode uses and not writing out the answer section-by-section. Interestingly, we found it was possible to do so without sacrificing performance. 这使得训练数据的质量变得尤为重要。我们没有依赖更复杂的优化方法来弥补噪声样本,而是花费了大量时间研究如何生成、选择和过滤那些真正体现我们所需报告撰写行为的样本。我们还希望 AstaBrief 更快,以便用户能迅速获得初步报告,并在后续轮次中进行迭代。为了提高速度,我们决定训练 AstaBrief 在给定用户查询和相关检索片段的情况下,直接一次性生成最终报告,跳过我们基于 Claude 的“思考模式”所使用的昂贵的片段摘要和聚类阶段,也不分章节撰写答案。有趣的是,我们发现这样做并不会牺牲性能。

Collecting SFT training data: The training pipeline began with real user queries submitted through the system described in our paper “Synthesizing scientific literature with retrieval-augmented LMs” and ScholarQA, the framework that now underpins Asta’s Generate a report feature. 收集 SFT 训练数据:训练流程始于通过我们论文《利用检索增强语言模型综合科学文献》中所述系统提交的真实用户查询,以及目前支撑 Asta“生成报告”功能的框架 ScholarQA。