The 3-2-1-1-0 backup rule. We had everything except the zero

The 3-2-1-1-0 backup rule. We had everything except the zero

3-2-1-1-0 备份法则:我们拥有了一切,唯独缺了那个“零”

Backup looks simple until the day you need it. In simple terms, a backup means making a copy of your data so it can be recovered if the original is lost, corrupted, or deleted. But there is an important difference between having a backup and being able to recover your data from it. 备份看起来很简单,直到你需要它的那一天。简单来说,备份意味着制作一份数据副本,以便在原始数据丢失、损坏或删除时能够恢复。但“拥有备份”与“能够从备份中恢复数据”之间存在着重要的区别。

Imagine that every day an automated process backs up your database. At the end, the system shows a big green mark: the process finished without errors. You might think: “everything is protected.” But what does that green mark actually mean? Only that the process was able to run the steps it was supposed to run. By itself, it does not guarantee that the copy is intact, that the data is consistent, or that you can recover it when you really need to. 想象一下,每天都有一个自动化流程在备份你的数据库。最后,系统显示一个大大的绿色对勾:流程无错误完成。你可能会想:“一切都受到保护了。”但这个绿色对勾到底意味着什么?它仅仅意味着该流程能够执行预定的步骤。它本身并不能保证副本是完整的、数据是一致的,也不能保证你在真正需要时能够恢复它。

This is where the 3-2-1-1-0 rule comes in. It is a strategy to reduce the risk of losing data even when something goes wrong. The numbers describe how copies should be kept and protected: 这就是 3-2-1-1-0 法则的用武之地。这是一种旨在降低数据丢失风险的策略,即使在出现问题时也能奏效。这些数字描述了副本应如何保存和保护:

  • 3 copies of the data

  • 2 different types of media or storage

  • 1 copy kept off the primary environment

  • 1 isolated or immutable copy

  • 0 errors during the recovery test

  • 3 份数据副本

  • 2 种不同类型的介质或存储

  • 1 份副本保存在主环境之外

  • 1 份隔离或不可篡改的副本

  • 0 次恢复测试错误

The last number is the interesting one. Zero is not another copy or another place to store backups. It is a requirement: when it is time to recover the data, the restore has to work. Because a backup that has never been restored is, to a point, a hypothesis. 最后一个数字最值得玩味。“零”不是指另一个副本或另一个存储备份的地方。它是一项要求:当需要恢复数据时,恢复操作必须成功。因为一个从未被恢复过的备份,在某种程度上只是一个假设。

What fails when a number is missing

当缺少某个数字时会发生什么

The rule only holds if each number covers a different failure. If two numbers break the same way, you do not have five controls. You have repetition. 只有当每个数字涵盖不同的故障场景时,该法则才成立。如果两个数字以相同的方式失效,你并没有五个控制点,而只是在重复。

Without the 3, production and the “backup” live in the same place. A retention mistake, a broad policy, or a compromised identity deletes the original and the copies together. Three names in one location are not three copies. 没有“3”,生产环境和“备份”就处于同一位置。一个保留策略错误、宽泛的权限策略或身份被盗用,都会导致原始数据和副本同时被删除。在同一个地方存放三个名称并不等于三份副本。

Without the 2, the copies use the same media or the same storage service. The failure repeats: corruption, quota, API, region. Two identical disks in the same account fail the same way. 没有“2”,副本就会使用相同的介质或相同的存储服务。故障会重复发生:损坏、配额限制、API 故障、区域性中断。同一个账户下的两块相同磁盘会以同样的方式失效。

Without the 1 off the primary environment, the incident in that environment takes the backup with it. Compromised account, single IdP, unavailable region. The copy has to sit where that incident cannot reach. 没有“1”(主环境之外),主环境中的事故会连带摧毁备份。例如账户被入侵、单一身份提供商(IdP)故障或区域不可用。副本必须存放在事故无法触及的地方。

Without the 1 that is isolated or immutable, whoever reached production also reaches the backup delete. Isolation cuts the easy path. Immutability blocks the delete even when the path appears. Ransomware that hits production usually hits the backup repository next. Without that lock, the way back disappears with the original. 没有“1”(隔离或不可篡改),任何能触及生产环境的人也能删除备份。隔离切断了攻击路径;不可篡改性即使在路径暴露时也能阻止删除。勒索软件在攻击生产环境后,通常会紧接着攻击备份存储库。没有这种锁定,恢复之路就会随着原始数据一起消失。

Without the 0, the job is green and the box will not open. The first four numbers say where the copy lives. The zero says whether it comes back. That is failure engineering, not a poster. It is still incomplete until restore has been exercised. 没有“0”,任务显示绿色,但盒子却打不开。前四个数字说明了副本存放在哪里,而“零”决定了它能否回来。这是故障工程,而不是口号。在进行恢复演练之前,它始终是不完整的。

The green job lies by omission

绿色的任务通过遗漏来撒谎

A successful backup job answers a narrow question: the writer persisted what was requested, to the configured destination, inside the window. It does not answer whether the incremental chain is intact. It does not answer whether the database was captured in a usable state (with a flush, a consistent snapshot, a native backup), or only as a disk in the middle of a write. 一个成功的备份任务只回答了一个狭窄的问题:写入器是否在规定时间内将请求的数据持久化到了配置的目标位置。它无法回答增量链是否完整,也无法回答数据库是否以可用状态被捕获(通过刷新、一致性快照或原生备份),还是仅仅捕获了一个正在写入过程中的磁盘。

It does not answer whether the KMS key for restore still exists, whether the identity doing the restore can read the vault, whether the destination engine accepts that version, or whether the time to return fits what the business calls acceptable. 它无法回答用于恢复的 KMS 密钥是否仍然存在,执行恢复的身份是否有权读取存储库,目标引擎是否接受该版本,或者恢复时间是否符合业务的可接受范围。

Integrity means the backup is complete and coherent. Availability means you can reach it and apply it. Neither one is the job’s exit code. That is why the zero is not a quantity. It is not “one more vault”. It is the acceptance criterion: restore on purpose, watch the data come back, measure how long it took, and only then call the design 3-2-1-1-0. 完整性意味着备份是完整且连贯的。可用性意味着你可以访问并应用它。这两者都不是任务的退出代码。这就是为什么“零”不是一个数量,它不是“再多一个存储库”。它是验收标准:有目的地进行恢复,观察数据是否回来,测量耗时,只有做到这些,才能称之为 3-2-1-1-0 设计。

Test the job. Test recoverability. Both. One without the other leaves a green check on a box that will not open. 测试任务,测试可恢复性。两者缺一不可。否则,你只会得到一个显示绿色对勾、却永远打不开的盒子。

What a recovery test has to show

恢复测试必须展示什么

The zero does not ask for a lab. It asks for four answers, obtained on purpose, without panic: “零”并不要求建立实验室,它要求在非紧急情况下,通过有目的的测试获得四个答案:

  1. Did the data come back?

  2. Did it come back consistent, in a state the application will accept?

  3. Could someone on the team run the restore without inventing permissions, keys, or a destination on the spot?

  4. Does the time to return fit what the business can stand?

  5. 数据是否回来了?

  6. 数据回来时是否一致,且处于应用程序可接受的状态?

  7. 团队成员是否能在不临时编造权限、密钥或目标的情况下执行恢复?

  8. 恢复时间是否在业务可承受范围内?

If any answer is “the job was green”, the test has not happened yet. A dashboard measures the writer. Recovery measures the way back. 如果任何一个答案是“任务显示绿色”,那么测试就还没有真正发生。仪表盘衡量的是写入过程,而恢复衡量的是回归之路。

Does the backup work when you need it?

当你需要时,备份真的有效吗?

“Do we have a backup?” That sounds like the right question. In practice, it says little. Most environments can answer yes. The question that actually matters is this one: When was the last time someone restored that backup on purpose, in a real environment, and confirmed that the system came back up? “我们有备份吗?”这听起来是个正确的问题。但在实践中,它说明不了什么。大多数环境都能回答“有”。真正重要的问题是:上一次有人在真实环境中刻意恢复该备份,并确认系统成功启动是什么时候?

A dashboard can show that the backup finished successfully. It can show that the copy exists, that storage is healthy, and that the last job completed without errors. None of those signals answers whether you can recover the system when you actually need to. 仪表盘可以显示备份已成功完成。它可以显示副本存在、存储健康、且上一个任务无错误完成。但这些信号都无法回答:当你真正需要时,能否恢复系统。

If the only evidence that the backup works is a green dashboard, you know you have a copy. You might already have the 3, the 2, and both 1s. But you are still missing the zero. The zero is proof that the way back works. 如果证明备份有效的唯一证据是绿色的仪表盘,那么你只知道你拥有一个副本。你可能已经做到了 3、2 和两个 1,但你仍然缺少那个“零”。“零”才是证明回归之路畅通的证据。