Advanced
Open
Pro
Continuous Batching or PagedAttention — Which Fixes This Symptom?
A team running a self-hosted model behind a naive request-at-a-time serving setup observes two separate symptoms under real concurrent load: (1) GPU utilization graphs show clear idle gaps even while requests are queued waiting to be served, and (2) the serving process occasionally fails to accept a new request with an out-of-memory error, even though the sum of currently active requests' actual token counts so far is well under the GPU's KV-cache capacity.
- Which of vLLM's two core techniques addresses symptom (1), and explain the mechanism that causes the idle gaps in the first place.
- Which technique addresses symptom (2), and explain why the OOM happens despite actual token usage being well under capacity.
- Would adopting only one of the two techniques (not both) fully resolve this team's problems? Explain why or why not.
Share this question