[CHEATSHEET] Use multiple GPUs and pipeline.
Idea: Merge Decode and Prefill phase of different requests (in batch dimension) and operate on this batched input. Use Paged KV cache.
This is Continuous Batching with 4 requests (R1, R2, R3, R4) all executing on 1 GPU.
The smaller uniform blocks are
decode phase.Larger 3 phases are
prefill phase.

At every instance, decode and prefill phases from different requests are merged in batch (B) dimension. So, one operation is executing at a time. This one opera...
Published on July 30, 2026 08:53