Introducing Gemini 3.8 Flash and 3.8 Flash Cyber
Introducing Gemini 3.8 Flash and 3.8 Flash Cyber
隆重推出 Gemini 3.8 Flash 与 3.8 Flash Cyber
Our newest Gemini models deliver next-generation intelligence for agentic workflows and cybersecurity. 我们最新的 Gemini 模型为智能体工作流和网络安全领域带来了新一代的智能。
Building on the momentum of 3.7 Flash from three weeks ago and marking our third Flash release in only six weeks, today we’re introducing Gemini 3.8, our best reasoning & coding model yet, at the same speed and low cost of 3.7. 继三周前发布 3.7 Flash 之后,我们仅在六周内就迎来了第三次 Flash 版本更新。今天,我们正式推出 Gemini 3.8——这是我们迄今为止最强大的推理与编程模型,且保持了与 3.7 相同的速度和低廉成本。
Gemini 3.8 introduces 2 variants: Gemini 3.8 包含两个版本:
Gemini 3.8 Flash: our most intelligent workhorse model, delivering significant improvements from 3.7 Flash across software engineering, agentic tasks, and critical, multi-step reasoning in specialized domains. It is available at the same introductory price as 3.7 Flash at $0.75 per million input tokens and $3.75 per million output tokens. Gemini 3.8 Flash: 我们最智能的主力模型。相较于 3.7 Flash,它在软件工程、智能体任务以及专业领域的关键多步推理方面均有显著提升。其定价与 3.7 Flash 的首发价格相同,即每百万输入 Token 0.75 美元,每百万输出 Token 3.75 美元。
Gemini 3.8 Flash Cyber: our most capable cybersecurity model with frontier-level performance in vulnerability detection and automated patching, available to trusted defenders through our new Fairwind Program. Gemini 3.8 Flash Cyber: 我们能力最强的网络安全模型,在漏洞检测和自动修复方面具备前沿水平的性能。该版本通过我们全新的“Fairwind 计划”向受信任的防御者开放。
While tailored for different deployment environments, both of today’s releases are powered by the same foundational intelligence, and further accelerated by long-running agentic loops designed to recursively evaluate and refine the underlying models. The significant coding and reasoning gains across this shared core were driven by a number of innovations, including rigorous training in the highly demanding domain of cybersecurity. 尽管针对不同的部署环境进行了优化,但今天发布的两个版本均由相同的底层智能驱动,并由旨在递归评估和优化底层模型的长效智能体循环进一步加速。这一共享核心在编程和推理能力上的显著提升,得益于多项创新,其中包括在要求极高的网络安全领域进行的严苛训练。
Gemini 3.8 Flash: built for long-horizon coding and autonomous agents
Gemini 3.8 Flash:专为长周期编程与自主智能体打造
Gemini 3.8 Flash delivers substantial gains from 3.7 Flash, often approaching the performance of higher-cost frontier models. On DeepSWE v1.1 (Long-Horizon Software Engineering), 3.8 Flash outperforms most larger frontier models in autonomously solving complex engineering problems end to end, only at a fraction of the cost. Gemini 3.8 Flash 较 3.7 Flash 有了实质性的提升,性能往往接近成本更高的前沿模型。在 DeepSWE v1.1(长周期软件工程)测试中,3.8 Flash 在自主端到端解决复杂工程问题方面的表现优于大多数大型前沿模型,而成本仅为后者的一小部分。
Additionally, 3.8 Flash exhibits the dependability required for critical enterprise autonomy, across specialized knowledge domains. In quantitative and professional fields that require advanced analysis and reporting, 3.8 Flash outperforms 3.7 Flash and other frontier models in benchmarks like Vals Finance Agent V2 and Harvey’s Legal Agent Benchmark. 3.8 Flash also achieves a 54.9% on HLE-Verified, demonstrating its ability to handle multi-step reasoning across STEM、humanities, and professional fields. 此外,3.8 Flash 在各个专业知识领域展现出了关键企业自主性所需的可靠性。在需要高级分析和报告的量化及专业领域,3.8 Flash 在 Vals Finance Agent V2 和 Harvey’s Legal Agent Benchmark 等基准测试中均优于 3.7 Flash 及其他前沿模型。3.8 Flash 在 HLE-Verified 测试中也达到了 54.9% 的准确率,证明了其处理 STEM、人文科学及专业领域多步推理的能力。
These performance gains stem from a core design choice: 3.8 Flash works harder. On complex tasks, it exhibits greater diligence — executing extra reasoning steps, and calling tools iteratively. At times, the model might use more tokens to maximize performance, especially at higher effort levels. 这些性能提升源于一个核心设计选择:3.8 Flash 工作更“努力”。在处理复杂任务时,它表现得更加严谨——执行额外的推理步骤并迭代调用工具。有时,为了最大化性能,模型可能会使用更多的 Token,尤其是在高努力程度(effort levels)设置下。
For applications where compute efficiency is the primary constraint, developers can utilize lower effort levels to minimize token overhead or continue to rely on Gemini 3.7 Flash, which remains fully supported for efficiency-first workloads. 对于以计算效率为首要约束的应用,开发者可以使用较低的努力程度设置来最小化 Token 开销,或者继续使用 Gemini 3.7 Flash,该版本将继续为追求效率的工作负载提供全面支持。
Gemini 3.8 Flash Cyber: expert cyber performance
Gemini 3.8 Flash Cyber:专家级的网络安全性能
Gemini 3.8 Flash Cyber, available to a set of trusted defenders via the Fairwind Program, provides a decisive advantage in today’s complex cybersecurity landscape, with the Flash speed and cost that enables quick iteration. Gemini 3.8 Flash Cyber 通过 Fairwind 计划向部分受信任的防御者开放,在当今复杂的网络安全环境中提供了决定性优势,并凭借 Flash 的速度和成本优势实现了快速迭代。
Autonomous vulnerability discovery 自主漏洞发现
On the standard industry benchmark for finding vulnerabilities, CyberGym, Gemini 3.8 Flash Cyber demonstrates frontier-level performance in autonomous vulnerability discovery. It surpasses both 3.5 Flash Cyber as well as significantly larger frontier models. 在行业标准的漏洞发现基准测试 CyberGym 上,Gemini 3.8 Flash Cyber 在自主漏洞发现方面展现了前沿水平的性能。它不仅超越了 3.5 Flash Cyber,还胜过了规模大得多的前沿模型。
To better capture real-world defensive needs which are not limited to just C/C++ codebases like in CyberGym, we also evaluated Gemini 3.8 Flash Cyber against a comprehensive internal benchmark in which the model has to discover a wide range of vulnerabilities across complex codebases spanning 20 programming languages. Here, the model showcases an impressive leap over our previous models and reaches a success rate exceeding 70%. 为了更好地满足现实世界的防御需求(这些需求不仅限于 CyberGym 中的 C/C++ 代码库),我们还使用一项全面的内部基准测试对 Gemini 3.8 Flash Cyber 进行了评估。在该测试中,模型必须在涵盖 20 种编程语言的复杂代码库中发现各种漏洞。结果显示,该模型较我们之前的版本有了令人印象深刻的飞跃,成功率超过了 70%。
Automated patching 自动修复
With Gemini 3.8 Flash Cyber, we focused specifically on equipping defenders with expert capabilities that give them an advantage over attackers. This is why we have invested in vulnerability fixing from the start, and prioritized it over offensive capabilities like exploitation. 在 Gemini 3.8 Flash Cyber 中,我们特别专注于为防御者配备专家级能力,使其在面对攻击者时占据优势。正因如此,我们从一开始就投入于漏洞修复,并将其优先级置于漏洞利用等攻击性能力之上。
CWE-Bench, run by Collinear, is a challenging external benchmark for patching capabilities. On this benchmark, Gemini 3.8 Flash Cyber is on the Pareto frontier: with a pass@1 of 47.2% compared to a leading frontier model at 47.8%, yet offered at a significantly lower cost. 由 Collinear 运营的 CWE-Bench 是一个极具挑战性的外部补丁能力基准测试。在该测试中,Gemini 3.8 Flash Cyber 处于帕累托前沿:其 pass@1 得分为 47.2%,而领先的前沿模型为 47.8%,但我们的模型成本显著更低。
Real-world impact: securing Google’s code 现实世界的影响:保障 Google 代码安全
We’re already using Gemini 3.8 Flash Cyber to secure code across Google. For example: 我们已经开始使用 Gemini 3.8 Flash Cyber 来保障 Google 全线的代码安全。例如:
- The Chrome Security team found that 3.8 Flash Cyber produced 2.6 times more correct patches to vulnerabilities in Chrome than the best commercial models that are much larger. Chrome 安全团队发现,3.8 Flash Cyber 在修复 Chrome 漏洞时产生的正确补丁数量,是目前市面上规模大得多的顶级商业模型的 2.6 倍。
- Wiz found that Gemini 3.8 Flash Cyber achieves +7.5-9.7% higher recall on their internal penetration test. Wiz 发现,Gemini 3.8 Flash Cyber 在其内部渗透测试中的召回率提高了 7.5% 到 9.7%。