Amazon Aurora Serverless v2 guia detalhado de arquitetura, escala, custo e operação
Amazon Aurora Serverless v2: A Detailed Guide to Architecture, Scaling, Cost, and Operation
Amazon Aurora Serverless v2 is an Aurora configuration where the database’s compute capacity tracks demand in real-time, rather than being fixed at instance creation. You define a range (minimum and maximum), and Aurora adjusts the capacity within that range without restarting the database or dropping connections. It is not a separate product, but rather an instance type within a standard Aurora cluster (db.serverless). Therefore, it inherits distributed storage, Multi-AZ, read replicas, and the rest of the Aurora ecosystem. Pricing, ACU limits, and engine version availability change over time. The figures in this article are references for reasoning; always confirm with the official AWS documentation and pricing pages.
Amazon Aurora Serverless v2 是一种 Aurora 配置,其数据库计算容量会随需求实时变化,而不是在创建实例时固定。您可以定义一个范围(最小值和最大值),Aurora 会在该范围内调整容量,而无需重启数据库或中断连接。它不是一个独立的产品,而是普通 Aurora 集群中的一种实例类型(db.serverless)。因此,它继承了分布式存储、多可用区(Multi-AZ)、只读副本以及 Aurora 生态系统的其余部分。价格、ACU 限制和引擎版本的可用性会随时间变化。本文中的数字仅供参考,请务必查阅 AWS 文档和定价页面以获取最新信息。
2. A bit of context: v1 vs. v2
| Feature | Serverless v1 | Serverless v2 |
|---|---|---|
| Model | Entire cluster scales | Instances scale individually |
| Scaling | Stepped, with pause to find “safe point” | Continuous, in small increments |
| Connections | Interrupts connections when scaling | Does not interrupt |
| Multi-AZ, Replicas, Global DB | Limited | Supported |
| Granularity | Capacity doubles | 0.5 ACU |
| Status | Legacy | Current model |
v2 was designed to eliminate the limitations of v1 and is the recommended model for new workloads.
2. 背景:v1 与 v2 的对比
| 特性 | Serverless v1 | Serverless v2 |
|---|---|---|
| 模型 | 整个集群扩展 | 实例独立扩展 |
| 扩展方式 | 阶梯式,需暂停以寻找“安全点” | 连续式,以小增量进行 |
| 连接 | 扩展时会中断连接 | 不会中断 |
| 多可用区、副本、全局数据库 | 受限 | 支持 |
| 粒度 | 容量翻倍 | 0.5 ACU |
| 状态 | 旧版 (Legacy) | 当前模型 |
v2 的设计旨在消除 v1 的局限性,是新工作负载的推荐模型。
3. How it works under the hood
3.1 The ACU (Aurora Capacity Unit)
Each ACU is equivalent to approximately 2 GiB of memory, with proportional CPU and networking. Capacity ranges from fractions of an ACU to hundreds of ACUs, with increments of 0.5 ACU. Think of an ACU as an instance “slice”: a database at 16 ACU has roughly 32 GiB of memory available.
3.1 ACU (Aurora 容量单位)
每个 ACU 大约相当于 2 GiB 内存,并配有相应的 CPU 和网络资源。容量范围从 ACU 的分数到数百个 ACU 不等,增量为 0.5 ACU。可以将 ACU 视为实例的“切片”:一个 16 ACU 的数据库大约拥有 32 GiB 的可用内存。
3.2 In-place scaling
Unlike changing instance classes, v2 adjusts CPU and memory within the database process itself. Practical consequences:
- No failover to scale.
- Ongoing connections and transactions are preserved.
- Scaling responds to CPU, memory, and network, not just the number of connections.
- Scaling speed depends on current capacity: the larger the instance, the faster it grows. A database stuck at 0.5 ACU scales slower than one already operating at 16 ACU. This is relevant for those using a very low minimum and expecting sudden spikes.
3.2 原地扩展 (In-place scaling)
与更改实例类型不同,v2 在数据库进程内部调整 CPU 和内存。实际影响包括:
- 扩展时无需故障转移 (Failover)。
- 正在进行的连接和事务得以保留。
- 扩展响应 CPU、内存和网络负载,而不仅仅是连接数。
- 扩展速度取决于当前容量:实例越大,增长越快。处于 0.5 ACU 的数据库扩展速度比已经在 16 ACU 下运行的数据库慢。对于设置了极低最小值并预期突发流量的用户来说,这一点非常重要。
3.3 Minimum and maximum range
You configure:
min_capacity: Capacity floor (can be 0 ACU with auto-pause, in compatible engine versions).max_capacity: Capacity ceiling and cost ceiling. The ACU in use never leaves this range. Beyond controlling costs, the floor and ceiling influence database parameters, such as buffer pool size and maximum connections, which are derived from the capacity configuration.
3.3 最小和最大范围
您可以配置:
min_capacity:容量下限(在兼容的引擎版本中,通过自动暂停可设为 0 ACU)。max_capacity:容量上限和成本上限。 使用的 ACU 永远不会超出此范围。除了控制成本外,下限和上限还会影响数据库参数(如缓冲池大小和最大连接数),这些参数均由容量配置衍生而来。
3.4 Auto-pause (scale-to-zero)
With min_capacity = 0, the database can pause after a period of inactivity you configure, and compute charges stop. When a new connection arrives, the database resumes, which takes some time (in the order of seconds). Points of attention:
- Storage is still charged.
- The first request after a pause suffers from resume latency.
- It is ideal for dev, test, and internal tools, and risky for production with waiting users.
- Pooled connections and frequent health checks can prevent pausing, keeping the database awake without you realizing it.
3.4 自动暂停 (Scale-to-zero)
当 min_capacity = 0 时,数据库可以在您配置的空闲时间后暂停,从而停止计算计费。当有新连接到达时,数据库会恢复,这需要一定时间(秒级)。注意事项:
- 存储费用仍会产生。
- 暂停后的第一个请求会经历恢复延迟。
- 此功能非常适合开发、测试和内部工具,但对于有用户等待的生产环境来说风险较高。
- 连接池和频繁的健康检查可能会阻止暂停,导致数据库在您未察觉的情况下保持运行。
3.5 Storage
Storage is the same as conventional Aurora: a distributed volume across 3 AZs, with 6 copies, growing automatically. You do not provision disk. Storage and I/O billing are independent of the instance type, and there are Standard and I/O-Optimized configuration options.
3.5 存储
存储与传统 Aurora 相同:跨 3 个可用区的分布式卷,包含 6 个副本,自动增长。您无需预置磁盘。存储和 I/O 计费与实例类型无关,并提供 Standard 和 I/O-Optimized 配置选项。
3.6 Promotion tiers in clusters with readers
In clusters with more than one instance, failover tiers (0 to 15) influence how readers scale:
- Tiers 0 and 1: The reader tracks the writer’s capacity to be ready to take over in case of failover.
- Tiers 2 to 15: The reader scales independently, according to its own load. This choice directly affects cost and failover readiness.
3.6 带有只读副本集群的提升优先级 (Promotion tiers)
在拥有多个实例的集群中,故障转移优先级(0 到 15)会影响只读副本的扩展方式:
- 优先级 0 和 1:只读副本会跟随写入节点的容量,以便在发生故障转移时准备好接管。
- 优先级 2 到 15:只读副本根据自身负载独立扩展。 此选择直接影响成本和故障转移的就绪性。
4. Supported features
In general, v2 works with the Aurora ecosystem:
- Multi-AZ and automatic failover
- Up to 15 read replicas
- Aurora Global Database
- RDS Proxy
- Performance Insights and Enhanced Monitoring
- Backups, snapshots, and Point-in-Time Recovery
- Encryption with KMS and IAM authentication
- Integration with Secrets Manager (
manage_master_user_password) - Blue/Green Deployments
- Data API (depending on version and region)
- Compatible with Aurora MySQL and Aurora PostgreSQL, in specific engine versions. It is worth checking the version matrix before creating the cluster.
4. 支持的功能
总体而言,v2 与 Aurora 生态系统兼容:
- 多可用区和自动故障转移
- 最多 15 个只读副本
- Aurora 全局数据库
- RDS Proxy
- Performance Insights 和增强监控
- 备份、快照和时间点恢复 (PITR)
- 使用 KMS 加密和 IAM 身份验证
- 与 Secrets Manager 集成 (
manage_master_user_password) - 蓝绿部署
- Data API(取决于版本和区域)
- 兼容特定引擎版本的 Aurora MySQL 和 Aurora PostgreSQL。在创建集群之前,建议先查看版本矩阵。
5. Cost model
5.1 How it is billed
- Compute: ACU-hour consumed, measured by the second, with a minimum charge per capacity period.
- Storage: per GB-month.
- I/O: per request in Standard mode. In I/O-Optimized mode, there is no I/O charge, but compute and storage prices are higher.
- Others: Backups beyond free retention, data transfer, Performance Insights beyond the free tier, etc.
5.1 计费方式
- 计算: 按消耗的 ACU 小时计费,按秒计量,每个容量周期有最低收费。
- 存储: 按 GB-月计费。
- I/O: Standard 模式下按请求计费。在 I/O-Optimized 模式下,不收取 I/O 费用,但计算和存储价格较高。
- 其他: 超出免费保留期的备份、数据传输、超出免费额度的 Performance Insights 等。
5.2 Break-even reasoning
The ACU-hour is, per unit of memory, more expensive than an equivalent provisioned instance. Illustrating with approximate values from us-east-1:
- 1 ACU ≈ US$ 0.12/h → running 24x7 ≈ US$ 87/month per ACU.
- A
db.r6g.largeinstance (16 GiB) costs around US$ 0.26/h ≈ US$ 190/month. - 16 GiB is equivalent to ~8 ACUs, which would cost ≈ US$ 0.96/h ≈ US$ 700/month if left on all the time. The conclusion is that under constant, high load, provisioned is much cheaper. If the database uses on average less than ~25–30% of the equivalent capacity, serverless usually comes out ahead, and with auto-pause in dev, it can cost a fraction. The correct reasoning is: cost = average ACUs consumed over time × ACU-hour price. That is why monitoring average ACU matters more than the peak.
5.2 盈亏平衡分析
按内存单位计算,ACU 小时的价格比同等的预置实例更贵。以 us-east-1 的近似值说明:
- 1 ACU ≈ 0.12 美元/小时 → 24x7 运行 ≈ 87 美元/月/ACU。
db.r6g.large实例 (16 GiB) 成本约为 0.26 美元/小时 ≈ 190 美元/月。- 16 GiB 相当于约 8 个 ACU,如果一直开启,成本约为 0.96 美元/小时 ≈ 700 美元/月。 结论是:在持续高负载下,预置实例要便宜得多。如果数据库平均使用量低于同等容量的 ~25–30%,Serverless 通常更具优势;而在开发环境中使用自动暂停,成本甚至可以忽略不计。正确的逻辑是:成本 = 时间内消耗的平均 ACU × ACU 小时价格。这就是为什么监控平均 ACU 比监控峰值更重要。
5.3 What not to forget
- Reserved Instances do not apply to ACU in the same way.
- Database Savings Plans may apply, according to current rules.
max_capacityis your compute spending limit, so set it consciously.- Readers with tier 0/1 follow the writer, and this multiplies the cost.
5.3 不容忽视的要点
- 预留实例 (Reserved Instances) 不以同样的方式适用于 ACU。
- 数据库节省计划 (Savings Plans) 可能适用,具体取决于当前规则。
max_capacity是您的计算支出上限,请谨慎设置。- 优先级 0/1 的只读副本会跟随写入节点,这会成倍增加成本。