PD 分离 — Prefill 和 Decode 各走各的 GPU 池

Prefill (compute-bound) → H100 池  |  Decode (memory-bound) → H200 池  |  KV Cache 通过 NIXL/RDMA 跨池传输
vLLM Router — PD Router is_prefill? → H100 Pool  |  is_decode? → H200 Pool 用户请求 Prefill Pool — 高算力 ×4 H100 H100 989 TFLOPS 3.35 TB/s 适合 Prefill GEMM 密集 H100 989 TFLOPS 3.35 TB/s 适合 Prefill GEMM 密集 H100 ... H100 ... Decode Pool — 高带宽 ×6 H200 H200 4.8★ Decode H200 4.8★ Decode H200 4.8★ Decode H200 4.8★ Decode H200 4.8★ Decode H200 4.8★ Decode 新对话 / Prefill (prompt 全部 token) 继续生成 / Decode (逐 token 追加) KV Cache 传输 (NIXL / RDMA) 节点内 NVLink <1μs  |  跨节点 RDMA ~10μs 维度 Prefill (提示处理) Decode (逐 token 生成) 计算类型 矩阵-矩阵乘法 (GEMM) 矩阵-向量乘法 (GEMV) 瓶颈 Compute-bound — 需要高算力 Memory-bound — 需要高带宽 最优 GPU H100 (989 TFLOPS, 3.35 TB/s) H200 (989 TFLOPS, 4.8 TB/s ★) PD 分离收益 独立扩缩: 早高峰扩 Prefill 池 P99 TPOT -30~50% (不抢显存) 核心洞察: 同一个 GPU 型号不可能同时最优于两种负载 — PD 分离让硬件匹配负载特性