MAI-Cyber-1-Flash inside MDASH
MAI-Cyber-1-Flash inside MDASH
Models Introducing MAI-Cyber-1-Flash inside MDASH: World-class security at half the cost 模型介绍:MDASH 中的 MAI-Cyber-1-Flash——以一半的成本实现世界级安全防护
Mustafa Suleyman & Hayete Gallot | July 27, 2026 Mustafa Suleyman 与 Hayete Gallot | 2026年7月27日
Today we’re announcing MAI-Cyber-1-Flash inside of MDASH, our multi-agent vulnerability identification and remediation harness. Together they deliver world-class performance at 50% of the cost of leading models. 今天,我们宣布在 MDASH(我们的多智能体漏洞识别与修复框架)中集成 MAI-Cyber-1-Flash。两者结合,能以领先模型 50% 的成本提供世界级的性能表现。
Progress in AI has been startling and so has the new generation of cyber threats it’s unleashing. Attackers now wield increasingly powerful capabilities, probing an ever-growing mountain of code for just a single weakness that lets them in. As the cost of finding a flaw collapses, the old model of security, where you scan occasionally and patch eventually, is now obsolete. 人工智能的进步令人惊叹,但它所引发的新一代网络威胁同样不容小觑。攻击者现在拥有日益强大的能力,在海量代码中搜寻哪怕一个能让他们趁虚而入的弱点。随着发现漏洞的成本不断降低,那种“偶尔扫描、最终修补”的旧安全模式已然过时。
If we’re to unlock the true benefits of AI, we must first build outstanding cyber models that help all of us harden the software the world runs on. That’s the motivation behind MAI-Cyber-1-Flash, which has been built to find challenging vulnerabilities in complex codebases. It’s been deeply integrated into MDASH, honed by the best cybersecurity experts in the industry and hardened across the largest security estate on the planet. 如果我们想要释放人工智能的真正潜力,首先必须构建出色的网络模型,帮助我们所有人加固全球运行所依赖的软件。这就是 MAI-Cyber-1-Flash 背后的初衷,它专为在复杂代码库中发现高难度漏洞而生。它已深度集成到 MDASH 中,由业内顶尖的网络安全专家打磨,并在全球最大的安全资产规模上进行了强化。
This combined expertise delivers exceptional security protection, beating Mythos, Gemini and GPT on CyberGym, the gold standard benchmark for evaluating how systems reason over large codebases to find real vulnerabilities in the code. 这种结合了专业知识的方案提供了卓越的安全防护,在 CyberGym(评估系统如何通过大型代码库推理并发现真实漏洞的黄金标准基准测试)中击败了 Mythos、Gemini 和 GPT。
Picking the right model for the task
为任务选择合适的模型
Security is an always-on mission, and given the enormous volume of inbound attacks, token cost is now the real constraint for defenders. MAI-Cyber-1-Flash was designed to efficiently handle up to 90% of all tasks, enabling MDASH to use the larger and most costly models in our fleet (in this case GPT-5.4) for the 10% of exceptionally hard tasks that truly need them. 安全是一项全天候的任务,考虑到海量的入站攻击,Token 成本已成为防御者面临的真正制约因素。MAI-Cyber-1-Flash 的设计初衷是高效处理高达 90% 的常规任务,从而让 MDASH 能够将资源集中,仅在真正需要时(即那 10% 极具挑战性的任务)调用我们系统中更大、成本更高的模型(如 GPT-5.4)。
The result is that the unified system of MDASH with MAI-Cyber-1-Flash delivers 96% on CyberGym (+12 pt above Mythos). This combination delivers a 50% cost saving when compared against our best offering in MDASH today (GPT 5.4 + 5.4 mini + 5.3 codex). That’s the power of a well-tuned, multi-model system with access to uniquely rich historical training data. It ensures you always have the best model at the best price for every task. 结果显示,集成 MAI-Cyber-1-Flash 的 MDASH 统一系统在 CyberGym 上取得了 96% 的得分(比 Mythos 高出 12 个百分点)。与我们目前 MDASH 中最好的方案(GPT 5.4 + 5.4 mini + 5.3 codex)相比,这种组合实现了 50% 的成本节约。这就是一个经过精心调优、能够访问极其丰富的历史训练数据的多模型系统的威力。它确保您在处理每项任务时,始终能以最优价格获得最合适的模型。
In this new environment, being able to go from identifying a new vulnerability to addressing it in real-time is critical. And while AI remediation of software vulnerabilities is now a key security workflow, there are many jobs to be done by Security practitioners themselves. That’s why today we’re also launching Perception, our agentic security systems, that provides teams of agents for a variety of security workflows in MDASH, to continuously monitor, patch, and close new threat vectors. Perception will also soon use MAI-Cyber-1-Flash for many more security workflows, beyond the software vulnerability work. 在这一新环境下,能够从识别新漏洞到实时修复至关重要。虽然人工智能修复软件漏洞已成为关键的安全工作流,但仍有许多工作需要安全从业者亲自完成。因此,我们今天还推出了 Perception——我们的智能体安全系统。它为 MDASH 中的各种安全工作流提供智能体团队,以持续监控、修补并关闭新的威胁向量。Perception 也将很快在软件漏洞工作之外,将 MAI-Cyber-1-Flash 应用于更多的安全工作流中。
Three things matter today: Model. Data. Harness.
当下最重要的三件事:模型、数据、框架。
We have jointly optimized our world-class models, our unmatched historic data, and our expert-tuned harness to ensure that our customers have a uniquely powerful security offering. 我们共同优化了世界级的模型、无与伦比的历史数据以及经专家调优的框架,以确保我们的客户拥有独一无二的强大安全产品。
- Model. MAI-Cyber-1-Flash is a compact, code-heavy security model derived from the MAI-Thinking-1 lineage, which was built from scratch, in-house, on the highest quality data. Details in our technical report. 模型。 MAI-Cyber-1-Flash 是一款紧凑型、侧重代码的安全模型,源自 MAI-Thinking-1 系列,由我们内部基于最高质量的数据从零构建。详情请参阅我们的技术报告。
- Data. Our deepest advantage. Decades of building world-class security systems now give us trillions of daily signals across identity, endpoint, cloud, and network, and an unmatched record of real exploits and remediations. No one can manufacture this history. 数据。 我们最深厚的优势。数十年来构建世界级安全系统的经验,使我们每天能获得跨身份、终端、云和网络的数万亿条信号,并拥有无与伦比的真实漏洞利用与修复记录。没有人能凭空创造出这段历史。
- Harness. MDASH, our multi-agent vulnerability identification and remediation harness, is tuned by the best security experts in the industry, who have created 100+ agents using multiple leading models to find, validate, and remediate vulnerabilities. Agentic code scanning is a critical function in the Security Operating Center and feeds Project Perception, our new agentic security system. 框架。 MDASH 是我们的多智能体漏洞识别与修复框架,由业内顶尖的安全专家调优。他们利用多个领先模型创建了 100 多个智能体,用于发现、验证和修复漏洞。智能体代码扫描是安全运营中心(SOC)的一项关键功能,并为我们的新智能体安全系统“Project Perception”提供支持。
Built with safety first
安全至上
Because MAI-Cyber-1-Flash is Microsoft’s first cyber model, we built trust into every layer of the system, from model training to customer deployment. The model was developed with a security-first calibration, rigorously evaluated by Microsoft’s AI Red Team, tested through automated and expert-led adversarial exercises, and independently assessed by a third party. 由于 MAI-Cyber-1-Flash 是微软的首个网络模型,我们将信任构建在系统的每一层中,从模型训练到客户部署。该模型的开发采用了“安全优先”的校准方式,经过微软 AI 红队的严格评估,通过了自动化和专家主导的对抗性演练,并由第三方进行了独立评估。
Trust extends beyond the model itself. Through MDASH, customers get enterprise-grade controls including Role-Based Controls, tenant isolation, encryption, auditability, and sandboxed execution environments with no internet access. The result is a cyber model that delivers powerful capabilities to defenders while maintaining the governance, security, and control enterprises expect from Microsoft. 信任不仅限于模型本身。通过 MDASH,客户可以获得企业级的控制功能,包括基于角色的访问控制(RBAC)、租户隔离、加密、可审计性以及无互联网访问的沙箱执行环境。最终,我们打造出了一个既能为防御者提供强大能力,又能保持企业所期望的微软级治理、安全和控制水平的网络模型。
Our hill-climbing machine
我们的“爬山”机器(持续进化机制)
Cybersecurity is not just a data-rich domain; it is a live reinforcement learning loop. Every day, defenders investigate threats, triage alerts, hunt adversaries, remediate vulnerabilities, deploy protections, and learn from the outcome. Microsoft sees that loop end to end: vulnerabilities through Microsoft Security Response Center; attacks and defenses across identity, endpoint, cloud, data, browser, and applications; more than 100 trillion security signals every day; and operational insight from 1.6 million customers. 网络安全不仅是一个数据丰富的领域,它更是一个实时的强化学习循环。每天,防御者都在调查威胁、分类警报、搜寻攻击者、修复漏洞、部署防护并从结果中学习。微软能够端到端地洞察这一循环:通过微软安全响应中心(MSRC)处理漏洞;在身份、终端、云、数据、浏览器和应用程序中进行攻防;每天处理超过 100 万亿条安全信号;并从 160 万客户那里获得运营洞察。
Because we can connect actions to outcomes; what was exploitable, what was contained, what was blocked, and what actually worked; we have more than data. Our MAI reinforcement learning loop gives us the foundation to build cyber models that improve continuously and become expert cyber defenders. That’ll remain our commitment to our customers for years to come. 因为我们能够将行动与结果联系起来——什么被利用了、什么被遏制了、什么被拦截了、什么真正起作用了——我们拥有的不仅仅是数据。我们的 MAI 强化学习循环为我们构建能够持续改进并成为专家级网络防御者的模型奠定了基础。在未来的岁月里,这将始终是我们对客户的承诺。
Build the Future With Us
与我们共建未来
We’re a lean, talent-dense team of explorers, researchers, and full-stack engineers. We move fast, sweat the details, and operate at frontier scale with a roadmap to build the world’s most powerful AI models. Most importantly, we’re united by the belief that doing this right is the only way. 我们是一支精简且人才济济的团队,由探索者、研究人员和全栈工程师组成。我们行动迅速,注重细节,以前沿规模运营,并制定了构建全球最强大 AI 模型的路线图。最重要的是,我们坚信:以正确的方式完成这项工作是唯一的途径。