P/D disaggregation

[CHEATSHEET] Separate GPUs for Prefill and Decode.

Note:

In continuous batching, if prefill and decode and merged, then time is wasted for decode.Prefill GPU occupancy ~ 90%Decode GPU occupancy ~30% and limited by memory bandwidth.Prefill prefers TP and DP parallelism.Decode prefers PP.

Core idea of P/D disaggregation:

Schedule prefill and decode requests in separate GPUs.2P1D = 2 GPUs for prefill requests. 1 GPU for decode requests.Prefill requests can be batched together.Decode requests can be bat...
 •  0 comments  •  flag
Share on Twitter
Published on July 30, 2026 08:54
No comments have been added yet.