Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems

Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems

分而隐蔽,合而为害:针对基于技能的智能体系统的“技能级联攻击”

Abstract: A skill is a modular package of natural-language instructions, executable scripts, and reference resources that an agent can load at runtime to extend its capabilities for a specific task. Skill-based agent systems therefore enable flexible reuse of third-party capabilities, but the openness of this skill ecosystem also opens up a new attack surface. 摘要: “技能”是一种模块化包,包含自然语言指令、可执行脚本和参考资源,智能体可以在运行时加载这些技能以扩展其针对特定任务的能力。因此,基于技能的智能体系统能够灵活地复用第三方能力,但这种技能生态系统的开放性也带来了一个新的攻击面。

Prior work has focused on vulnerabilities within individual skills, but little attention has been paid to risks that arise from interactions across skills. In this paper, we introduce skill cascading attacks, a threat paradigm in which a malicious objective is distributed across multiple skills so that each modification looks benign in isolation, yet their combined execution is harmful. 以往的研究主要集中在单个技能内部的漏洞,但很少关注跨技能交互所产生的风险。在本文中,我们引入了“技能级联攻击”(skill cascading attacks),这是一种威胁范式,其中恶意目标被分散到多个技能中,使得每一项修改在孤立状态下看起来都是良性的,但它们的组合执行却具有危害性。

For instance, in a prescription-review pipeline, the first skill weakens signals of recently discontinued medications in the extracted history, the second downgrades the severity of any drug interaction tied to them, and the third suppresses the resulting low-priority alert in the final summary, so that a severe drug-interaction warning silently disappears before reaching the physician. 例如,在一个处方审查流程中,第一个技能会削弱提取出的病史中近期停用药物的信号,第二个技能会降低与这些药物相关的任何药物相互作用的严重程度,第三个技能则会在最终摘要中抑制由此产生的低优先级警报,从而导致严重的药物相互作用警告在到达医生面前之前就悄无声息地消失了。

To systematically study this safety blind spot, we develop SkillCascade, an automated multi-agent red-teaming framework, and release SkillCascade-Bench, a benchmark of 213 validated cascading test cases across multiple agent systems and domains. Across representative agents (e.g., OpenClaw, Claude Code, Codex) and LLM backbones, cascaded interactions reliably induce harmful behaviors while evading existing per-skill scanners and runtime monitors. 为了系统地研究这一安全盲点,我们开发了 SkillCascade,这是一个自动化的多智能体红队测试框架,并发布了 SkillCascade-Bench,这是一个包含 213 个跨多个智能体系统和领域的已验证级联测试用例的基准测试集。在代表性智能体(如 OpenClaw、Claude Code、Codex)和大语言模型骨干网络上,级联交互能够可靠地诱导产生有害行为,同时规避现有的针对单个技能的扫描器和运行时监控器。

Our findings highlight a gap between component-level integrity and system-level safety, and call for defenses that reason over cross-skill interactions rather than individual skills in isolation. 我们的研究结果强调了组件级完整性与系统级安全性之间的差距,并呼吁采取能够针对跨技能交互进行推理,而非仅关注孤立的单个技能的防御措施。