Gemini 4 Argon: our next era of frontier intelligence

Gemini 4 Argon: our next era of frontier intelligence

Gemini 4 Argon:我们前沿智能的新纪元

Gemini 4 Argon delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense. Gemini 4 Argon 在现实世界的软件工程、法律和金融等企业知识工作以及网络安全防御等复杂工作流程中,提供了前沿的性能表现。

Google’s new Gemini 4 Argon model brings advanced reasoning to complex, long-horizon professional tasks. The model features an industry-leading 1 million token limit for deep, multi-step problem solving. It excels at coding, financial research, legal drafting, and autonomous cybersecurity vulnerability patching. Argon is currently rolling out to trusted cyber defenders through the Fairwind Program. Google is prioritizing safety and rigorous testing before a wider release to the public. 谷歌全新的 Gemini 4 Argon 模型为复杂、长周期的专业任务带来了先进的推理能力。该模型具备行业领先的 100 万 token 上限,支持深度、多步骤的问题解决。它在编程、金融研究、法律起草和自主网络安全漏洞修复方面表现出色。Argon 目前正通过“Fairwind 计划”向受信任的网络防御者开放。在向公众广泛发布之前,谷歌正优先考虑安全性和严格的测试。

Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program. Built to sustain deep reasoning across complex, long-horizon workflows, Argon is fundamentally changing the way we work and build at Google. It delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense. 今天,我们宣布推出全新的前沿模型 Gemini 4 Argon,该模型正通过我们的“Fairwind 计划”向一批受信任的网络防御者开放。Argon 旨在支持跨复杂、长周期工作流程的深度推理,它正在从根本上改变我们在谷歌的工作和构建方式。它在现实世界的软件工程、法律和金融等企业知识工作以及网络安全防御的复杂工作流程中,提供了前沿的性能。

Safely releasing frontier capabilities at this level requires a phased approach. We are actively engaged in the U.S. government’s voluntary process for pre-release model access while we gradually expand access. We’ll continue to gather feedback from early testers as we iterate on guardrails before making Argon available to developers, enterprises, and consumers as soon as possible. 安全地发布这一级别的前沿能力需要分阶段进行。我们正积极参与美国政府关于模型预发布访问的自愿流程,同时逐步扩大访问范围。我们将继续收集早期测试者的反馈,并在尽快向开发者、企业和消费者提供 Argon 之前,不断迭代安全护栏。

Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off input token price. Argon 的首发定价为每百万输入 token 2 美元,每百万输出 token 10 美元;缓存的输入 token 可享受输入 token 价格 95% 的折扣。

Changing how we work and build at Google

改变我们在谷歌的工作与构建方式

Gemini 4 Argon is already powering our internal workflows, with thousands of Googlers highlighting the model’s strengths in specialized coding tasks, conducting deeper research, and writing quality. It’s helping teams build faster and push the boundaries of engineering productivity and accelerating breakthroughs. Gemini 4 Argon 已经在驱动我们的内部工作流程,数以千计的谷歌员工强调了该模型在专业编程任务、进行深度研究以及写作质量方面的优势。它正在帮助团队更快地构建产品,突破工程生产力的极限,并加速实现技术突破。

Quantum algorithmic optimization: Argon is helping our quantum computing researchers optimize the spacetime resources (qubits × gates) of subroutines that bottleneck important applications. In one example, it beat the published baseline by 40% in a matter of minutes. 量子算法优化: Argon 正在帮助我们的量子计算研究人员优化那些阻碍重要应用发展的子程序的时空资源(量子比特 × 门)。在一个案例中,它在几分钟内就将已发布的基准性能提升了 40%。

Memory efficiency: A team of Argon agents analyzed fleet-wide profiling telemetry to autonomously identify and apply memory optimizations across Google’s data centers, freeing up over 300 TiB of memory once rolled out, with an estimated 500 TiB to 1 PiB in total savings. 内存效率: 一组 Argon 智能体通过分析全网范围的性能遥测数据,自主识别并应用了谷歌数据中心的内存优化方案。部署后释放了超过 300 TiB 的内存,预计总节省量可达 500 TiB 至 1 PiB。

Large Scale Codebase Migrations and Optimizations: Argon agents are working on migrating C/C++ codebases to Rust across Google — scaling from tens of thousands of lines in core libraries like re2, libgav1 up to 800K+ lines for the Fuchsia Zircon kernel. Given the criticality of many of these systems, such large-scale rewrites are undergoing rigorous automated and manual auditing, emulation testing, and review before rolling out to production. 大规模代码库迁移与优化: Argon 智能体正在谷歌内部进行将 C/C++ 代码库迁移至 Rust 的工作——规模从 re2、libgav1 等核心库的数万行代码,扩展到 Fuchsia Zircon 内核超过 80 万行的代码。鉴于其中许多系统的关键性,这些大规模重写在投入生产之前,都要经过严格的自动化和人工审计、仿真测试及评审。

Working harder on your most complex problems

更深入地解决您最复杂的问题

To support Gemini 4 Argon’s capabilities across longer, more complex use cases, we are significantly expanding the model’s output token limit to an industry-leading 1M tokens, up from the previous 64K tokens. When the model has the headroom to think deeply and generate hundreds of thousands of tokens in a single trajectory, it adds a new level of depth in reasoning to solve tough problems in one go. 为了支持 Gemini 4 Argon 在更长、更复杂的用例中的能力,我们将模型的输出 token 上限从之前的 64K 大幅扩展至行业领先的 100 万 token。当模型拥有足够的空间进行深度思考并在单次轨迹中生成数十万个 token 时,它在推理深度上达到了一个新的水平,能够一次性解决棘手的问题。

Enabling coding and enterprise workflows across domains

赋能跨领域的编程与企业工作流程

Gemini 4 Argon’s capabilities across coding, reasoning, and multimodality and its ability to sustain long, multi-step tasks enable it to excel across a range of enterprise workflows. Gemini 4 Argon 在编程、推理和多模态方面的能力,以及其处理长周期、多步骤任务的能力,使其能够在各种企业工作流程中表现出色。

Google engineers have been using Argon for their daily tasks, from everyday debugging to large-scale codebase migrations and algorithm designs. It sets a new state of the art on DeepSWE v1.1 (77.9%), which measures a model’s performance in real-world long-horizon software engineering tasks. 谷歌工程师一直在将 Argon 用于日常任务,从日常调试到大规模代码库迁移和算法设计。它在 DeepSWE v1.1(77.9%)上创下了新的行业最佳水平,该基准测试衡量了模型在现实世界长周期软件工程任务中的表现。

Beyond coding, Argon is the leading model on the Vals Index, which measures economic impact across finance, coding, legal, and tax work, with every sector weighted by its contribution to U.S. GDP. We see similarly leading performance across other domain specific evaluations, like Vals Finance Agent v2 (multi-step financial research) and Harvey’s Legal Agent Benchmark (legal research and drafting). On AutomationBench, Zapier’s benchmark measuring end-to-end execution across core business functions, Argon ranks #1 with a score of 51.3%. 除了编程之外,Argon 还是 Vals Index 上的领先模型。该指数衡量金融、编程、法律和税务工作中的经济影响,并根据各行业对美国 GDP 的贡献进行加权。我们在其他特定领域的评估中也看到了类似的领先表现,例如 Vals Finance Agent v2(多步骤金融研究)和 Harvey’s Legal Agent Benchmark(法律研究与起草)。在 Zapier 用于衡量核心业务功能端到端执行情况的 AutomationBench 基准测试中,Argon 以 51.3% 的得分排名第一。

Argon is also uniquely strong when knowledge work requires visual understanding. It’s able to drive professional chart analysis, identify details from long videos, and take action based on a series of documents. For example, on LVBench, which measures long video understanding, Argon is state of the art with a score of 91.7%. 当知识工作需要视觉理解时,Argon 也表现出独特的优势。它能够进行专业的图表分析,从长视频中识别细节,并根据一系列文档采取行动。例如,在衡量长视频理解能力的 LVBench 上,Argon 以 91.7% 的得分处于行业领先地位。

Leading in defensive cybersecurity

引领防御性网络安全

To better equip cyber defenders for the new era of cyberattacks, we trained Gemini 4 Argon to be highly… 为了让网络防御者更好地应对新时代的网络攻击,我们训练了 Gemini 4 Argon,使其具备高度的……