When Your Own Server Becomes the Attacker

When Your Own Server Becomes the Attacker

当你自己的服务器成为攻击者

In March 2019, a former Amazon engineer made a request to a Capital One server. Not a login, not a password guess. She asked the server to fetch a URL for her. The server said yes. That one request walked out with temporary AWS credentials, and those credentials unlocked the personal data of more than 100 million people. Social security numbers, bank account details, credit applications. One of the largest financial breaches in US history, and at the bottom of it was a bug so ordinary that you have probably written the ingredients for it yourself. It’s called SSRF, and I want to show you exactly how something this mundane turns into something this catastrophic.

2019 年 3 月,一位前亚马逊工程师向第一资本(Capital One)的一台服务器发送了一个请求。不是登录,也不是猜测密码,她只是要求服务器为她获取一个 URL。服务器答应了。仅仅这一个请求,就带走了临时的 AWS 凭证,而这些凭证解锁了超过 1 亿人的个人数据——包括社会安全号码、银行账户详情和信用卡申请信息。这是美国历史上最大的金融数据泄露事件之一,而其根源是一个极其普通的漏洞,你可能自己都写过类似的代码。它被称为 SSRF(服务器端请求伪造),我想向你展示这种看似平凡的东西是如何演变成如此灾难性的后果的。

The whole idea in one line: SSRF (Server-Side Request Forgery) is when you trick a server into making a request for you, to somewhere you could never reach on your own. That’s it. You don’t break in. You get the server to walk in for you.

核心概念一句话概括:SSRF(服务器端请求伪造)就是诱骗服务器为你发起请求,去访问你自己无法直接到达的地方。仅此而已。你不需要强行闯入,而是让服务器替你走进去。

Think about who you trust. Picture a secure office. You’re a stranger in the lobby. The server room, the records vault, the executive offices, all locked, all off-limits. Security won’t let you past the front desk. But the receptionist runs errands. “Could you grab the file from room 237?” Sure. They walk back, fetch it, hand it over. So you try your luck: “Could you bring me whatever’s in the server room?” And because the receptionist has access, and you asked politely, they do. You never touched the lock. The receptionist did the walking. That receptionist is your server, and SSRF is the art of asking it to fetch things it shouldn’t.

想想你信任谁。想象一个戒备森严的办公室。你是大厅里的陌生人。机房、档案室、高管办公室,全部上锁,严禁进入。保安不会让你通过前台。但接待员会跑腿。你问:“能帮我从 237 室拿个文件吗?”当然可以。他们走过去,取回文件,交给你。于是你试着碰碰运气:“能把机房里的东西给我拿来吗?”因为接待员有权限,而且你问得很有礼貌,他们就照做了。你从未触碰过锁,是接待员替你走了进去。那个接待员就是你的服务器,而 SSRF 就是诱导它去获取它本不该获取的东西的艺术。

The feature that becomes the bug. Here’s what makes SSRF sneaky: it almost never looks like a vulnerability. It looks like a feature request. “Let users add a profile picture by URL.” “Check stock levels from the warehouse’s internal API.” “Generate a PDF preview of any link.” “Fire a webhook to the customer’s endpoint.” Every one of those ends up as some version of this: const data = await fetch(userControlledUrl);

演变成漏洞的功能。SSRF 的隐蔽之处在于:它看起来几乎从不像是一个漏洞,而像是一个功能需求。比如:“允许用户通过 URL 添加头像”、“从仓库的内部 API 检查库存水平”、“生成任意链接的 PDF 预览”、“向客户的端点触发 Webhook”。每一个需求最终都会变成类似这样的代码:const data = await fetch(userControlledUrl);

The developer who wrote that line was picturing a normal URL. An image. A warehouse API. A customer’s webhook. What they weren’t picturing was someone typing: http://localhost/admin. And the server, having no opinion about where it’s pointed, fetches its own admin panel, the one that’s only supposed to be reachable from inside, and hands the attacker the response. The firewall guarding that panel never even got a say, because the request came from inside the house.

写下这行代码的开发者脑海中想的是一个正常的 URL:一张图片、一个仓库 API 或一个客户的 Webhook。他们没想过有人会输入:http://localhost/admin。而服务器对指向哪里没有判断力,它直接获取了自己的管理面板——那个本应只能从内部访问的面板——并将响应交给了攻击者。保护该面板的防火墙甚至连发言的机会都没有,因为请求来自“内部”。

When the server starts scanning for you. Reaching the admin panel is just the opening move. The deeper problem is that the server lives inside a private network you can’t see from the outside. So you make it look around. Point it at 192.168.0.1, then .2, then .3, and watch how the responses differ. A connection that refuses instantly means nothing’s there. One that hangs, or answers, means something is. Walk the whole range and the server quietly maps out the internal network for you, which machines exist, which ports are open, where the interesting things live. You’ve turned the victim’s own server into a scanner aimed at the victim’s own network.

当服务器开始为你扫描时。访问管理面板只是开场白。更深层的问题是,服务器位于你从外部无法看到的私有网络中。所以你让它四处查看。将其指向 192.168.0.1,然后是 .2,再是 .3,观察响应有何不同。连接立即被拒绝意味着那里什么都没有;连接挂起或有响应,则意味着那里有东西。遍历整个范围,服务器就会悄悄为你绘制出内部网络地图:哪些机器存在,哪些端口开放,有趣的东西在哪里。你已经把受害者的服务器变成了瞄准其自身网络的扫描仪。

The part that ended Capital One. Now the finale, the move that turns “interesting bug” into “front-page breach.” Every server running on AWS, Google Cloud, or Azure can reach a special internal address: 169.254.169.254. It’s the metadata service, a little endpoint the machine uses to learn about itself. Handy for the server. Devastating in the wrong hands, because on a misconfigured instance it will also hand over the server’s temporary cloud credentials. So the attacker points the SSRF here: http://169.254.169.254/latest/meta-data/iam/security-credentials/ and the server, asked politely, returns its own keys to the kingdom. With those, you’re no longer poking at one app. You’re inside the cloud account. That’s the whole Capital One story in one request: a URL field that trusted its input, a server that could reach the metadata endpoint, and 100 million records out the door.

导致 Capital One 事件的关键。现在是结局,这一招将“有趣的漏洞”变成了“头条新闻级的泄露”。在 AWS、Google Cloud 或 Azure 上运行的每台服务器都可以访问一个特殊的内部地址:169.254.169.254。这是元数据服务,机器用来了解自身信息的一个小端点。这对服务器很方便,但落入坏人之手则是毁灭性的,因为在配置错误的实例上,它还会交出服务器的临时云凭证。因此,攻击者将 SSRF 指向这里:http://169.254.169.254/latest/meta-data/iam/security-credentials/,服务器在被礼貌请求后,交出了自己的“王国钥匙”。有了这些,你就不再只是在探测一个应用程序,而是直接进入了云账户。这就是 Capital One 事件的全部经过:一个信任用户输入的 URL 字段,一个可以访问元数据端点的服务器,以及 1 亿条被窃取的数据记录。

How to not be the next case study. This is the part most write-ups skip, so here’s the version I’d actually want on a code review. 1. Allowlist, don’t blocklist. The instinct is to block the bad addresses. Don’t. Attackers will always find a spelling you didn’t think of: 127.0.0.1 can be written as 0x7f000001, or 2130706433, or hidden behind a domain that quietly resolves to an internal IP. Blocklists leak. Instead, decide the exact destinations your feature is allowed to hit, and refuse everything else. If the stock checker only ever needs one internal API, that’s the only thing it should be able to reach.

如何避免成为下一个案例。这是大多数文章都会跳过的部分,所以这是我在代码审查中真正想要看到的建议:1. 使用白名单,不要使用黑名单。本能反应是屏蔽坏地址,但千万别这么做。攻击者总能找到你没想到的写法:127.0.0.1 可以写成 0x7f000001 或 2130706433,或者隐藏在一个悄悄解析为内部 IP 的域名后面。黑名单总会泄露。相反,应明确规定你的功能允许访问的确切目的地,并拒绝其他所有请求。如果库存检查器只需要访问一个内部 API,那么它就只能访问那一个。

  1. Check the resolved IP, not the string. If you must allow user URLs, resolve the hostname first and inspect the actual IP it points to, then block the private ranges (127.x, 169.254.x, 10.x, 192.168.x, 172.16–31.x). A domain like my-innocent-site.com can resolve straight to 169.254.169.254 if the attacker owns the DNS.

  2. 检查解析后的 IP,而不是字符串。如果你必须允许用户输入 URL,请先解析主机名并检查它指向的实际 IP,然后屏蔽私有地址段(127.x, 169.254.x, 10.x, 192.168.x, 172.16–31.x)。如果攻击者控制了 DNS,像 my-innocent-site.com 这样的域名可以直接解析为 169.254.169.254。

  3. Kill the exotic schemes. Allow http and https, nothing else. file:// reads local files. gopher:// can forge raw requests to internal services like Redis. If your feature doesn’t need them, they’re pure downside.

  4. 禁用奇特的协议。只允许 http 和 https,其他一律禁止。file:// 可以读取本地文件,gopher:// 可以伪造对 Redis 等内部服务的原始请求。如果你的功能不需要它们,它们就纯粹是隐患。

  5. On AWS, enforce IMDSv2. The newer metadata service requires a session token before it hands anything over, which shuts down the exact move that hit Capital One. It’s close to a one-setting fix for the worst version of this bug. Turn it on.

  6. 在 AWS 上,强制使用 IMDSv2。较新的元数据服务在交出任何信息前需要会话令牌,这正好封堵了导致 Capital One 事件的那种攻击手段。这几乎是解决该漏洞最严重情况的“一键式”修复方案。请务必开启。

  7. Assume one layer fails. The app server shouldn’t be able to reach your sensitive internal systems in the first place. Segment the network so that even a successful SSRF runs into a wall instead of a vault.

  8. 假设某一层会失效。应用服务器本身就不应该能够访问你的敏感内部系统。对网络进行分段,这样即使 SSRF 攻击成功,攻击者撞上的也只是一堵墙,而不是金库。

One question before you close this tab: Go find the place in your codebase where the server fetches a URL it didn’t fully choose. The image importer, the webhook, the link preview, the PDF renderer. There’s almost always one. Now ask it two things: is the destination allowlisted, and can it reach 169.254.169.254? If you don’t like the answers, you’ve just found the same bug that cost Capital One 100 million records.

在关闭此页面前问自己一个问题:去代码库中找找服务器获取非完全自主选择的 URL 的地方。图片导入器、Webhook、链接预览、PDF 渲染器……几乎总会有一个。现在问它两个问题:目的地是否在白名单中?它能访问 169.254.169.254 吗?如果你对答案不满意,那么恭喜你,你刚刚发现了那个让 Capital One 损失 1 亿条记录的同款漏洞。