Introducing Gemini 3.5 Flash Cyber
Introducing Gemini 3.5 Flash Cyber
隆重推出 Gemini 3.5 Flash Cyber
July 21, 2026 | Models | Raluca Ada Popa and Four Flynn 2026年7月21日 | 模型 | Raluca Ada Popa 与 Four Flynn
Google has invested in cybersecurity for years, pioneering automated vulnerability discovery to secure the world’s codebases. Tools like CodeMender, our code security agent, can automatically find and fix critical software vulnerabilities. But as AI agents become more capable at finding vulnerabilities faster than defenders can fix them, addressing this global threat requires a highly capable, affordable, and scalable approach. 多年来,Google 一直致力于网络安全领域,率先开发了自动化漏洞发现技术,以保护全球的代码库。我们的代码安全代理工具 CodeMender 能够自动发现并修复关键的软件漏洞。然而,随着 AI 代理在发现漏洞方面的能力提升速度超过了防御者的修复速度,应对这一全球性威胁需要一种功能强大、经济实惠且可扩展的解决方案。
Today, we’re expanding our longtime efforts to better prepare defenders by introducing Gemini 3.5 Flash Cyber, our lightweight cybersecurity model built on top of 3.5 Flash and fine-tuned to find, validate, and patch vulnerabilities quickly and efficiently, making it more effective at these tasks than Gemini’s mainline Flash models. 今天,我们通过推出 Gemini 3.5 Flash Cyber 来扩展我们长期的防御准备工作。这是我们基于 3.5 Flash 构建的轻量级网络安全模型,经过微调后能够快速高效地发现、验证和修补漏洞,使其在执行这些任务时比 Gemini 的主流 Flash 模型更为有效。
Flash’s performance and efficiency makes it an ideal foundation for our cybersecurity model efforts. By building on top of Flash, 3.5 Flash Cyber offers a cost-efficient and highly capable alternative to large, costly cybersecurity models. Flash 的性能和效率使其成为我们网络安全模型工作的理想基础。通过在 Flash 之上构建,3.5 Flash Cyber 为那些大型且昂贵的网络安全模型提供了一种经济高效且功能强大的替代方案。
Given the dual-use nature of this technology, we have taken an intentional approach to how we deploy 3.5 Flash Cyber. As part of a limited-access pilot program, 3.5 Flash Cyber will be exclusively available to governments and trusted partners via CodeMender soon, expanding over time. This will give frontline defenders a head start in finding and fixing critical vulnerabilities before they can be exploited, while mitigating against broader misuse. 鉴于该技术的双重用途性质,我们在部署 3.5 Flash Cyber 时采取了审慎的态度。作为有限访问试点计划的一部分,3.5 Flash Cyber 将很快通过 CodeMender 专门提供给政府和受信任的合作伙伴,并随时间推移逐步扩大范围。这将使一线防御者能够在关键漏洞被利用之前抢先发现并修复它们,同时降低被滥用的风险。
Separately, we’re also bringing CodeMender’s foundational capabilities directly to customers with generally available Gemini models through the Gemini Enterprise Agent Platform. 此外,我们还将通过 Gemini 企业代理平台(Gemini Enterprise Agent Platform),将 CodeMender 的基础能力直接提供给使用通用 Gemini 模型的客户。
The search space problem: The advantage of lightweight models in code security
搜索空间问题:轻量级模型在代码安全中的优势
Finding deep-seated flaws requires exploring an immense execution search space. Relying on a single, expensive call to a massive language model can create a bottleneck. 3.5 Flash Cyber is particularly suitable for finding vulnerabilities where the agent has to scan a large codebase and analyze a large number of codepaths. 发现深层次的缺陷需要探索巨大的执行搜索空间。仅依赖对大型语言模型的一次昂贵调用可能会造成瓶颈。3.5 Flash Cyber 特别适用于代理需要扫描庞大代码库并分析大量代码路径的漏洞发现场景。
CodeMender invokes 3.5 Flash Cyber multiple times, so agents can analyze vastly more code paths to discover and validate vulnerabilities. The sub-agents then produce a single, high-quality report. CodeMender 会多次调用 3.5 Flash Cyber,使代理能够分析更多的代码路径以发现和验证漏洞。随后,子代理会生成一份高质量的综合报告。
Thanks to its speed and affordability, 3.5 Flash Cyber can be easily integrated into frequent scans, time-sensitive launch processes or commit scanning pipelines at scale. 得益于其速度和经济性,3.5 Flash Cyber 可以轻松集成到频繁的扫描、对时间敏感的发布流程或大规模的提交扫描流水线中。
3.5 Flash Cyber benchmark results: an efficient alternative to larger cybersecurity models
3.5 Flash Cyber 基准测试结果:大型网络安全模型的高效替代品
We tested 3.5 Flash Cyber on a variety of benchmarks. In particular, we tested 3.5 Flash Cyber on the CyberGym benchmark, which evaluates AI agents against hundreds of real-world software vulnerabilities. Leveraging the low cost of 3.5 Flash Cyber by configuring CodeMender to call 3.5 Flash Cyber up to five times for a single, final report, the overall agent achieved competitive performance against significantly larger models on CyberGym. 我们在各种基准测试中对 3.5 Flash Cyber 进行了测试。特别是在 CyberGym 基准测试中,该测试评估了 AI 代理应对数百个真实世界软件漏洞的能力。通过配置 CodeMender 对 3.5 Flash Cyber 进行最多五次调用以生成一份最终报告,我们利用了其低成本的优势,使该代理在 CyberGym 上取得了与规模大得多的模型相竞争的性能。
We also stress-tested the model’s capabilities beyond CyberGym without safety guardrails. Google’s Big Sleep team independently built an evaluation focused on finding critical and hard to find vulnerabilities in some of the world’s most complex codebases like Chrome and Safari. Here, 3.5 Flash Cyber significantly surpassed mainline 3.5 Flash and 3.6 Flash. 我们还在没有安全护栏的情况下,对该模型在 CyberGym 之外的能力进行了压力测试。Google 的“Big Sleep”团队独立构建了一项评估,专注于在 Chrome 和 Safari 等全球最复杂的代码库中寻找关键且难以发现的漏洞。结果显示,3.5 Flash Cyber 显著超越了主流的 3.5 Flash 和 3.6 Flash。
3.5 Flash Cyber was also evaluated on Google Chrome’s production commit scanning pipeline. The vulnerabilities were not publicly disclosed, which ensured this benchmark remained free of contamination for Gemini and competitor models. 3.5 Flash Cyber 还在 Google Chrome 的生产提交扫描流水线上进行了评估。由于这些漏洞未公开披露,确保了该基准测试对于 Gemini 和竞争对手模型而言没有受到污染。
The results showed a significant uplift from 3.5 Flash Cyber compared to 3.5 Flash. Note: More recent competitor model versions after Opus 4.6 refuse to fulfill the tasks due to built-in safety guardrails, and therefore are not shown. 结果显示,与 3.5 Flash 相比,3.5 Flash Cyber 有显著提升。注:Opus 4.6 之后更新的竞争对手模型版本由于内置了安全护栏而拒绝执行这些任务,因此未予展示。
Moreover, 3.5 Flash Cyber consistently discovered more unique vulnerabilities compared with mainline 3.5 Flash and Claude Opus 4.6. When tested on the highly complex V8 JavaScript Engine across a fixed number of invocations, 3.5 Flash Cyber found 55 unique confirmed issues, compared to 47 found by mainline 3.5 Flash and 36 found by Opus 4.6, including 10 issues that the other two models tested did not catch. 此外,与主流的 3.5 Flash 和 Claude Opus 4.6 相比,3.5 Flash Cyber 始终能发现更多独特的漏洞。在对高度复杂的 V8 JavaScript 引擎进行固定次数的调用测试时,3.5 Flash Cyber 发现了 55 个独特的已确认问题,而主流 3.5 Flash 发现了 47 个,Opus 4.6 发现了 36 个,其中包括另外两个模型未能捕捉到的 10 个问题。
Basic cybersecurity models can get stuck in a loop, finding the same issue repeatedly while missing critical vulnerabilities. A strong model casts a wider net, finding a higher number of unique issues. 基础的网络安全模型可能会陷入循环,反复发现同一个问题,却错过了关键漏洞。而强大的模型则能撒下更广的网,发现更多独特的问题。
As we scale the number of invocations, we find that 3.5 Flash Cyber continues to discover new code paths and vulnerabilities. 随着我们增加调用次数,我们发现 3.5 Flash Cyber 能够持续发现新的代码路径和漏洞。
Real-world application and scaling defenses at Google
Google 的实际应用与防御扩展
Benchmarks are only part of the story. 3.5 Flash Cyber in CodeMender is already finding and fixing vulnerabilities in Google’s internal codebases including Chrome, Android, Cloud, Ads, and YouTube. 基准测试只是故事的一部分。CodeMender 中的 3.5 Flash Cyber 已经在 Google 的内部代码库(包括 Chrome、Android、Cloud、Ads 和 YouTube)中发现并修复了漏洞。
The speed of discovery made possible by a lightweight model has delivered measurable impact. 轻量级模型所带来的发现速度已经产生了可衡量的影响。
For example, Google’s Cloud Vulnerability Research team used 3.5 Flash Cyber to proactively secure our systems in record time. In just 2 hours, the model uncovered remote code execution vulnerabilities in public APIs and found a memory-corruption vulnerability in a sensitive production service. It then generated a 100% reliable remote-code execution exploit that bypassed standard mitigation techniques like Address Space Layout Randomization (ASLR) and Write XOR Execute (W^X). 例如,Google 云漏洞研究团队使用 3.5 Flash Cyber 以创纪录的时间主动保护了我们的系统。仅在 2 小时内,该模型就发现了公共 API 中的远程代码执行漏洞,并找到了一个敏感生产服务中的内存损坏漏洞。随后,它生成了一个 100% 可靠的远程代码执行漏洞利用程序,绕过了地址空间布局随机化 (ASLR) 和写异或执行 (W^X) 等标准缓解技术。
Early feedback from Wiz and Cloud CISO Security Engineering testers confirms the significant capability improvement of 3.5 Flash Cyber over the mainline 3.5 Flash model. 来自 Wiz 和云 CISO 安全工程测试人员的早期反馈证实,3.5 Flash Cyber 相比主流 3.5 Flash 模型在能力上有显著提升。
Empowering defenders at scale
大规模赋能防御者
Google’s leadership in software security gives us a unique advantage. For example, OSV.dev, a vulnerability database run by Google spanning over 700,000 open-source vulnerabilities, and more than 10 years of OSS-Fuzz. Google 在软件安全领域的领导地位赋予了我们独特的优势。例如,由 Google 运营的 OSV.dev 漏洞数据库涵盖了超过 70 万个开源漏洞,以及超过 10 年的 OSS-Fuzz 积累。