"Just use RBAC" is half right. Here's the other half for AI agents.

I maintain Aegis-DevOps, an open-source check that runs before an AI coding agent’s shell commands and blocks the ones your policy forbids. In the last two weeks I’ve heard the same objection twice. A Reddit thread said RBAC already does this. A maintainer turned down a plugin submission because the agent harness “already has built-in provisions for controlling tool use.” Both are half right. You should have RBAC, and you should use your agent’s permission settings. But neither one answers the question that matters when an agent is about to run kubectl delete: should this action, from this agent, happen right now?

我维护着 Aegis-DevOps,这是一个开源检查工具,它在 AI 编程代理执行 Shell 命令之前运行,并拦截那些被你策略禁止的命令。在过去两周里,我两次听到了同样的反对意见。一个 Reddit 帖子称 RBAC 已经能做到这一点。一位维护者拒绝了一个插件提交,理由是代理工具包“已经内置了控制工具使用的规定”。这两种说法都只对了一半。你应该拥有 RBAC,也应该使用代理的权限设置。但当代理即将运行 kubectl delete 时,这两者都无法回答那个关键问题:这个代理的这个操作,现在应该执行吗?

Disclosure: I drafted this article with help from an AI assistant and reviewed it myself. The command results below are from running aegis-devops 0.3.2 against its example policy. Agent permission settings match text Claude Code’s permission rules are a good example, and its documentation is unusually honest about them. A Bash rule matches the command text the model writes. The docs say a deny rule “isn’t a security boundary around the program”, and give examples: Bash(git push *) stops git push origin main but not git -C . push origin main. For inspecting the full command before it runs, they point you to a PreToolUse hook. That’s fine for what permission settings are for: deciding what the agent may run without asking you. It isn’t a policy engine, and nobody claims it is. A hook can do more because it can work out what the command does.

披露:本文是在 AI 助手的帮助下起草的,并由我本人进行了审阅。下方的命令结果是运行 aegis-devops 0.3.2 对照其示例策略得出的。代理权限设置匹配文本。Claude Code 的权限规则就是一个很好的例子,其文档对这些规则的描述非常坦诚。Bash 规则匹配模型编写的命令文本。文档指出,拒绝规则“并不是程序周围的安全边界”,并给出了示例:Bash(git push *) 可以阻止 git push origin main,但无法阻止 git -C . push origin main。为了在命令运行前检查完整指令,他们指向了一个 PreToolUse 钩子。这对于权限设置的初衷来说是没问题的:即决定代理在无需询问你的情况下可以运行什么。它不是一个策略引擎,也没人声称它是。钩子可以做得更多,因为它能分析出命令的具体意图。

These all reach the same rule in Aegis (“no deletes in the prod namespace”): Command the agent writes | Result kubectl delete deploy web -n prod | BLOCK kubectl -n prod delete deploy web | BLOCK kubectl delete deploy web -nprod | BLOCK /usr/bin/kubectl …, sudo kubectl …, env kubectl … | BLOCK bash -c ‘kubectl delete deploy web -n prod’ | BLOCK K=kubectl; $K delete deploy web -n prod | refused: can’t be checked statically, so the hook denies it

这些命令在 Aegis 中都会触发同一条规则(“禁止在生产命名空间中执行删除操作”): 代理编写的命令 | 结果 kubectl delete deploy web -n prod | 拦截 kubectl -n prod delete deploy web | 拦截 kubectl delete deploy web -nprod | 拦截 /usr/bin/kubectl ..., sudo kubectl ..., env kubectl ... | 拦截 bash -c 'kubectl delete deploy web -n prod' | 拦截 K=kubectl; $K delete deploy web -n prod | 拒绝:无法静态检查,因此钩子予以拦截

The last row matters as much as the others. A guard that can’t understand a command should say no, not “no rule matched, go ahead”. RBAC sees the developer, not the agent. RBAC and IAM decide what an identity may do. A coding agent on a laptop usually runs with the developer’s own kubeconfig and cloud credentials. As far as the API server is concerned, the agent is the developer, with every permission the developer has. You can fix that by giving agents their own identities, and you should. But even then, RBAC is a static grant. It can say “this service account may delete deployments in prod”. It can’t say “this agent may not, unless a human approved it, and not during the release freeze”, and it has no idea where a rule came from.

最后一行与其它行同样重要。一个无法理解命令的守卫应该说“不”,而不是“没有匹配到规则,请继续”。RBAC 看到的是开发者,而不是代理。RBAC 和 IAM 决定了一个身份可以做什么。笔记本电脑上的编程代理通常使用开发者自己的 kubeconfig 和云凭证运行。对于 API 服务器而言,代理就是开发者本人,拥有开发者所拥有的所有权限。你可以通过为代理分配独立身份来解决这个问题,而且你也应该这样做。但即便如此,RBAC 仍然是一种静态授权。它可以说“此服务账户可以删除生产环境中的部署”,但它无法说“除非有人批准,且不在发布冻结期内,否则该代理不得执行此操作”,而且它根本不知道规则的来源。

That last part is the reason I built Aegis. In an agent’s world, a “rule” can arrive inside a Jira ticket or a Slack message the agent was asked to read. So an Aegis rule only gets a vote if nobody has changed it since it was signed, and if its author was allowed to write that kind of rule. A line planted in a ticket gets no vote.

最后一点就是我构建 Aegis 的原因。在代理的世界里,“规则”可能出现在代理被要求阅读的 Jira 工单或 Slack 消息中。因此,只有当规则自签名后未被篡改,且其作者被授权编写此类规则时,Aegis 规则才有效。植入工单中的一行文字是无效的。

Use all three. The honest answer to “why not RBAC?” is “yes, and”: RBAC / IAM: the hard ceiling on what any identity can do. Keep it tight. Agent permission settings: what the agent may run without asking you. A pre-execution policy check: whether this specific action should happen, by rules you can trust.

三者并用。对于“为什么不用 RBAC?”这个问题的诚实回答是“不仅要用,还要配合其它手段”: RBAC / IAM:任何身份能做什么的硬性上限。请保持严格。 代理权限设置:代理在无需询问你的情况下可以运行什么。 执行前策略检查:通过你可以信任的规则,判断这个特定操作是否应该发生。

The policy check doesn’t have to stay on the laptop either. Once agents have their own identities, Aegis compiles the same policy into the platform layer: aegis compile aws writes Service Control Policies, and aegis compile kubernetes writes ValidatingAdmissionPolicies. That covers calls that never go through the hook, like a boto3 script. Both are previews, and I wrote up how the AWS part works.

策略检查也不必局限于笔记本电脑。一旦代理拥有了各自的身份,Aegis 就可以将相同的策略编译到平台层:aegis compile aws 会编写服务控制策略(SCP),而 aegis compile kubernetes 会编写验证准入策略(ValidatingAdmissionPolicies)。这涵盖了那些从未经过钩子的调用,例如 boto3 脚本。两者目前均为预览版,我已经撰写了关于 AWS 部分如何工作的说明。

What it doesn’t do yet. Readers keep finding real gaps, which is the point of writing these posts. This week one found that a namespace-scoped rule doesn’t match a command that relies on the namespace in your current kube context, and that kubectl apply -f - is allowed because the piped manifest is never read. Both are being fixed. Aegis is alpha. Don’t put it in front of anything you care about without reading the open gaps in the repo first.

它尚未实现的功能。读者不断发现真正的漏洞,这也是我撰写这些文章的意义所在。本周有人发现,命名空间范围的规则无法匹配依赖于当前 kube 上下文中命名空间的命令,并且 kubectl apply -f - 是被允许的,因为管道传输的清单从未被读取。这两个问题都在修复中。Aegis 目前处于 Alpha 阶段。在阅读仓库中列出的已知漏洞之前,请勿将其用于任何你重视的项目。

Try it: pip install aegis-devops && aegis init .aegis

Claude Code

claude plugin marketplace add moneytool/aegis-devops claude plugin install aegis-devops@aegis-devops

Codex, Copilot, VS Code, Cursor, Gemini CLI, OpenCode

aegis install codex # or copilot | vscode | cursor | gemini | opencode

尝试一下: pip install aegis-devops && aegis init .aegis

Claude Code

claude plugin marketplace add moneytool/aegis-devops claude plugin install aegis-devops@aegis-devops

Codex, Copilot, VS Code, Cursor, Gemini CLI, OpenCode

aegis install codex # 或 copilot | vscode | cursor | gemini | opencode

If you think the RBAC answer is enough for your setup, I’d like to hear why. Issues are open at github.com/moneytool/aegis-devops.

如果你认为 RBAC 对于你的设置已经足够,我很想听听原因。欢迎在 github.com/moneytool/aegis-devops 提交 Issue。