llm-d-router v0.10.0 flow control · one H200 · 2026-09-30
Priority holdback alongside eviction
Flow control can protect interactive requests from batch work by evicting in-flight batch requests, or by holding back part of the pool so batch never fills it. Is it worth running both?
It depends on how the EPP measures saturation, and neither setting protects against bursts larger than the reserve.
Request counts
Optional. Holdback removes evictions under steady load, but eviction alone already keeps TTFT near the floor.
Token counts
Required. Without it, TTFT p50 is 2.6 s under steady load; with it, 0.03 s.
Hybrid
Mixed. Fixes steady load, makes bursts worse (p95 1.5 s → 5.3 s).
Bursts, any mode
No help. Bursts larger than the reserve still evict or queue.
Resume off
No help. Bursts then redo about as much decode as the batch needs (+74% batch time), with or without holdback.
1Setup
One vLLM 0.30 replica on an H200 with --max-num-seqs 32, behind llm-d-router v0.10.0 with flow control and eviction enabled.
Batch requests (priority −1) go through a queue processor that resumes an evicted stream from its saved tokens. Interactive requests (priority 10) go straight to the gateway.
Holdback is priority-holdback-policy. A batch ceiling of 0.9 or 0.8 caps batch at 29 or 26 of 32 slots; interactive keeps ceiling 1.0 and evicts only once the pool is full.
Batch: 3 jobs × 64 requests × 1,500 output tokens. 3 runs per arm; charts show the mean with min–max whiskers.
2Request-count saturation
With eviction alone, batch refills every freed slot, so about 9 in 10 interactive requests evict a batch request. A 20% reserve absorbs steady arrivals (0 evictions at 1 and 2 req/s, 18 vs 219 at 4 req/s). Bursts of 32 overwhelm any reserve.
Qwen2.5-1.5B, short prompts. Holdback lowers TTFT p95 to the 0.13 s floor at 1 req/s, where waiting on an eviction showed up, and costs up to ~3 s of batch time. Holdback without eviction never gets TTFT near the floor (40% reserve: p50 1.05 s vs 0.22 s under bursts).
All numbers
Steady load: one interactive request every 1/rate seconds for 60 s. Bursts: 15 bursts of 32 requests, 6 s apart; this table includes the arms without eviction. Batch deadlines were 45, 75 and 100 s (jobs a, b, c).
steady 1 req/s
arm
evictions
TTFT p50
TTFT p95
batch done
job a met
job b met
job c met
eviction
54
0.12 s
0.57 s
62.4 s
87%
100%
100%
eviction + 10% holdback
11
0.12 s
0.13 s
64.6 s
96%
100%
100%
eviction + 20% holdback
0
0.12 s
0.13 s
64.8 s
100%
100%
100%
steady 2 req/s
arm
evictions
TTFT p50
TTFT p95
batch done
job a met
job b met
job c met
eviction
108
0.12 s
0.18 s
65.0 s
86%
100%
100%
eviction + 10% holdback
34
0.12 s
0.13 s
64.3 s
88%
100%
100%
eviction + 20% holdback
1
0.12 s
0.13 s
65.1 s
98%
100%
100%
steady 4 req/s
arm
evictions
TTFT p50
TTFT p95
batch done
job a met
job b met
job c met
eviction
219
0.12 s
0.18 s
70.1 s
80%
100%
100%
eviction + 10% holdback
68
0.11 s
0.12 s
65.9 s
78%
100%
100%
eviction + 20% holdback
18
0.12 s
0.13 s
72.9 s
89%
100%
100%
bursts of 32 every 6 s
arm
evictions
TTFT p50
TTFT p95
batch done
job a met
job b met
job c met
no eviction
0
1.80 s
6.72 s
75.6 s
100%
100%
100%
10% holdback
0
1.46 s
6.29 s
75.1 s
100%
100%
100%
40% holdback
0
1.05 s
4.93 s
80.5 s
100%
100%
100%
eviction
350
0.22 s
0.43 s
69.1 s
16%
100%
100%
eviction + 10% holdback
329
0.22 s
0.41 s
69.9 s
0%
100%
100%
eviction + 20% holdback
324
0.21 s
0.39 s
69.7 s
0%
100%
100%
3Token and hybrid saturation
Counting tokens, an interactive request weighs about a ninth of a batch request, so one eviction lets ~8 interactive requests in while vLLM has one free sequence. The rest queue inside vLLM, out of flow control's reach. Holdback keeps real slots free and fixes steady load; under bursts it does not help, and with hybrid counting it hurts.
concurrency-detector can count requests, estimated tokens, or both (hybrid: per endpoint, the larger ratio). Output tokens are estimated as min(1.5 × prompt, max_tokens), so these runs use long prompts: ~4.9k-token batch prompts and 500-token interactive prompts × 128 tokens on Qwen2.5-1.5B, with maxTokenConcurrency set to 32 batch requests' worth (204,800). Load is driven in-cluster with a per-run cache_salt, so TTFT is lower here than in section 2; compare only within this section.
All numbers
requests · steady 2/s
arm
evictions
TTFT p50
TTFT p95
batch done
eviction
125
0.03 s
0.16 s
85.5 s
eviction + 20% holdback
0
0.03 s
0.16 s
87.3 s
tokens · steady 2/s
arm
evictions
TTFT p50
TTFT p95
batch done
eviction
159
2.64 s
4.53 s
90.7 s
eviction + 20% holdback
0
0.03 s
0.31 s
85.4 s
hybrid · steady 2/s
arm
evictions
TTFT p50
TTFT p95
batch done
eviction
132
0.27 s
0.53 s
84.2 s
eviction + 20% holdback
0
0.03 s
0.13 s
86.0 s
requests · bursts
arm
evictions
TTFT p50
TTFT p95
batch done
eviction
451
0.18 s
0.34 s
90.5 s
eviction + 20% holdback
428
0.16 s
0.30 s
95.6 s
tokens · bursts
arm
evictions
TTFT p50
TTFT p95
batch done
eviction
44
4.70 s
9.73 s
96.1 s
eviction + 20% holdback
4
3.60 s
9.38 s
93.1 s
hybrid · bursts
arm
evictions
TTFT p50
TTFT p95
batch done
eviction
442
0.15 s
1.52 s
96.2 s
eviction + 20% holdback
231
0.64 s
5.33 s
95.7 s
4Without resume
Turning resume off (evicted requests restart from scratch after backoff) barely matters under steady load but nearly doubles batch work under bursts: ~288k tokens redone against 288k needed, and the batch takes 74% longer. Interactive TTFT does not change. Holdback does not rescue it: under bursts it trims redone work ~8% and still finishes later.
Request-count saturation, same long-prompt mix and in-cluster driver as section 3. Without a render_url the processor cannot resume, so an evicted request is retried from the start with exponential backoff; resume and immediate requeue cannot be switched off separately.
All numbers
steady 2/s · resume
arm
evictions
decode redone
TTFT p95
batch done
eviction
125
0k
0.16 s
85.5 s
eviction + 20% holdback
0
0k
0.16 s
87.3 s
steady 2/s · restart
arm
evictions
decode redone
TTFT p95
batch done
eviction, restart
134
7k
0.12 s
87.9 s
eviction + 20% holdback, restart
1
1k
0.19 s
87.7 s
bursts · resume
arm
evictions
decode redone
TTFT p95
batch done
eviction
451
0k
0.34 s
90.5 s
eviction + 20% holdback
428
0k
0.30 s
95.6 s
bursts · restart
arm
evictions
decode redone
TTFT p95
batch done
eviction, restart
483
288k
0.34 s
157.7 s
eviction + 20% holdback, restart
437
264k
0.31 s
168.9 s
5A larger model: Qwen2.5-14B
Holdback removes evictions here too (166 → 11 → 2) but TTFT p95 stays at ~0.92 s in every arm, and batch takes 4–5% longer. Interactive latency comes from sharing the GPU with long-context batch work, not from waiting on evictions.
Keep resume on. Without it, each burst eviction throws away work and the batch takes 74% longer; holdback does not win that back.
Keep request-count saturation with eviction as the baseline. It keeps interactive TTFT near the floor under both steady load and bursts.
Add holdback, sized to steady interactive concurrency, when evictions themselves cost something (lost prefix cache, churn). Expect 1–6% more batch time.
If saturation is counted in tokens on a server with a sequence cap, run holdback: token counts alone admit more requests than the server can run.
Neither setting absorbs bursts larger than the reserve. That needs request-count backpressure or a larger reserve.
7Caveats and open questions
Saturation is counted in requests except in section 3. The KV-cache saturation detector is untested. A different token cap would move the section 3 results.
The 45 s batch job misses deadlines only under eviction (100% met without it on the same build). Leading suspect, not yet confirmed: all 192 batch requests reach the gateway at once, so the processor's earliest-deadline-first order never applies and a resumed request rejoins the gateway's FCFS queue behind later-deadline work.
Why hybrid counting with holdback gets worse under bursts is not yet explained.
The 14B runs did not use a per-run cache_salt, so prefill comparisons are left out; its 20% holdback arm has 2 runs.
One interactive request out of 1,440 (tokens mode, bursts, eviction only) got a 503.
Every completed batch output was token-identical to a direct vLLM run (101 runs).