FluidPD: In-Place Elasticity for SLO-Aware Prefill-Decode Disaggregated LLM Serving

Prefill-decode disaggregation is becoming a common architecture for LLM serving because it separates two phases with distinct execution patterns and SLO objectives. Existing systems typically combine a fixed prefill/decode worker ratio with request routing across workers.

预填充-解码(Prefill-decode)分离正成为大语言模型(LLM)服务的一种常见架构,因为它将具有不同执行模式和 SLO(服务等级目标)目标的两个阶段分离开来。现有系统通常结合了固定的预填充/解码工作节点比例,并对工作节点间的请求进行路由。

However, real-world workloads exhibit both short bursts and sustained shifts in the prefill-to-decode demand ratio. As a result, a configuration that is well provisioned at one time may quickly become mismatched, causing latency SLO violations even when idle capacity exists elsewhere.

然而,现实世界的工作负载在预填充与解码的需求比例上既表现出短时间的突发性,也存在持续性的偏移。因此,一个在某一时刻配置良好的系统可能会迅速变得不匹配,即使在其他地方存在空闲容量的情况下,也会导致延迟 SLO 的违规。

Existing autoscaling mechanisms can add capacity, but they react slowly, require spare GPUs, and do not directly address short-timescale phase imbalance. We present FluidPD, a P/D-disaggregated serving system that provides SLO-aware in-place elasticity.

现有的自动扩缩容机制虽然可以增加容量,但反应缓慢,需要备用 GPU,且无法直接解决短时间尺度下的阶段不平衡问题。我们提出了 FluidPD,这是一个提供 SLO 感知原地弹性(in-place elasticity)的 P/D 分离服务系统。

FluidPD introduces two complementary mechanisms. FluidToken handles transient imbalance by offloading a bounded portion of prefill computation to decode workers when decode-side slack is available. FluidRole handles sustained imbalance by reassigning running workers between prefill and decode roles in place, avoiding model reload and engine restart.

FluidPD 引入了两种互补的机制。FluidToken 通过在解码端有空闲资源时,将部分预填充计算卸载到解码工作节点,从而处理瞬时不平衡。FluidRole 则通过在原地重新分配运行中的工作节点在预填充和解码角色之间的任务,来处理持续性的不平衡,从而避免了模型重载和引擎重启。

Both mechanisms are guided by lightweight pressure indices that expose prefill and decode-side resource pressure before they appear as SLO violations. Across production Azure trace workloads, FluidPD improves overall SLO attainment over static SGLang by up to 94.6 percentage points, demonstrating that SLO-aware in-place P/D elasticity improves service quality without provisioning additional workers.

这两种机制均由轻量级的压力指标引导,这些指标能在 SLO 违规出现之前就揭示预填充和解码端的资源压力。在 Azure 生产环境的追踪工作负载测试中,FluidPD 将整体 SLO 达成率较静态 SGLang 提升了高达 94.6 个百分点,证明了 SLO 感知的原地 P/D 弹性可以在无需配置额外工作节点的情况下提升服务质量。