PD 分离 — Prefill 和 Decode 各走各的 GPU 池
Prefill (compute-bound) → H100 池 | Decode (memory-bound) → H200 池 | KV Cache 通过 NIXL/RDMA 跨池传输
vLLM Router — PD Router
is_prefill? → H100 Pool | is_decode? → H200 Pool
用户请求
Prefill Pool — 高算力
×4 H100
H100
989 TFLOPS
3.35 TB/s
适合 Prefill
GEMM 密集
H100
989 TFLOPS
3.35 TB/s
适合 Prefill
GEMM 密集
H100
...
H100
...
Decode Pool — 高带宽
×6 H200
H200
4.8★
Decode
H200
4.8★
Decode
H200
4.8★
Decode
H200
4.8★
Decode
H200
4.8★
Decode
H200
4.8★
Decode
新对话 / Prefill
(prompt 全部 token)
继续生成 / Decode
(逐 token 追加)
KV Cache 传输 (NIXL / RDMA)
节点内 NVLink <1μs | 跨节点 RDMA ~10μs
维度
Prefill (提示处理)
Decode (逐 token 生成)
计算类型
矩阵-矩阵乘法 (GEMM)
矩阵-向量乘法 (GEMV)
瓶颈
Compute-bound — 需要高算力
Memory-bound — 需要高带宽
最优 GPU
H100 (989 TFLOPS, 3.35 TB/s)
H200 (989 TFLOPS,
4.8 TB/s
★)
PD 分离收益
独立扩缩: 早高峰扩 Prefill 池
P99 TPOT -30~50% (不抢显存)
核心洞察: 同一个 GPU 型号不可能同时最优于两种负载 — PD 分离让硬件匹配负载特性