llm-d-async-rs · llm-d-router v0.10.0 · 2026-09-30
Holdback on top of eviction: it removes evictions under steady load, but only pays off when evictions are what hurt
Flow control can protect interactive traffic from batch work two ways: evict in-flight batch requests when interactive work is waiting, or hold back a share of the pool so batch never fills it. Is it worth running both?
It removes evictions under steady interactive load. Whether that helps depends on the model: on 1.5B it brings TTFT p95 to the floor; on 14B with long prompts it changes nothing.
- With eviction alone, batch refills every freed slot, so the pool is always full and about 9 in 10 interactive requests evict a batch request, at every rate tried.
- Holdback keeps a few slots free, so steady arrivals land there without evicting. With a 20% reserve: 0 evictions at 1 req/s, 1 at 2 req/s, 18 (vs 219) at 4 req/s. With 10%: about 70–80% fewer.
- That brings interactive TTFT p95 from 0.18–0.57 s down to the 0.13 s floor and lets the tightest batch job meet more deadlines, for at most ~3 s (4%) more batch time.
- Bursts of 32 overwhelm any reserve tried: 6–8% fewer evictions, no TTFT change, and the 45 s job gets worse. Holdback without eviction never gets TTFT near the floor.
- On 14B with ~4.8k-token batch prompts the evictions go away too (166 → 11 → 2) but TTFT p95 stays at ~0.92 s in every arm, and batch takes 4–5% longer. There, interactive tail latency comes from sharing the GPU with long-context batch work, not from waiting on evictions.
- So: holdback is cheap insurance when evictions are what interactive requests wait on (small models, short prompts, steady load). Size the reserve to steady interactive concurrency and let eviction handle bursts. Where TTFT is set by prefill and decode contention, it costs batch time and buys nothing.
Setup
- One vLLM 0.30 replica (Qwen2.5-1.5B-Instruct) on an H200,
--max-num-seqs 32. The EPP'sconcurrency-detectorcounts in-flight requests: saturation = in-flight / 32. - Batch goes through llm-d-async at priority −1; an evicted stream resumes from its saved tokens. Interactive requests go straight to the gateway at priority 10.
- Holdback is
priority-holdback-policywith batch ceilings 0.9, 0.8 and 0.6. The gate issaturation >= ceiling, so batch tops out at 29, 26 and 20 of 32 slots. Interactive keeps ceiling 1.0 and evicts only once the pool is full. - Batch: 3 jobs × 64 requests × 1,500 tokens, deadlines 45, 75 and 100 s after submit. Interactive: 256 output tokens. 3 runs per arm; charts show the mean with min–max whiskers.
Steady interactive load
One interactive request every 1/rate seconds for 60 s, alongside the batch.
steady 1 req/s
| arm | evictions | TTFT p50 | TTFT p95 | batch done | job a met | job b met | job c met |
|---|---|---|---|---|---|---|---|
| eviction | 54 | 0.12 s | 0.57 s | 62.4 s | 87% | 100% | 100% |
| eviction + 10% holdback | 11 | 0.12 s | 0.13 s | 64.6 s | 96% | 100% | 100% |
| eviction + 20% holdback | 0 | 0.12 s | 0.13 s | 64.8 s | 100% | 100% | 100% |
steady 2 req/s
| arm | evictions | TTFT p50 | TTFT p95 | batch done | job a met | job b met | job c met |
|---|---|---|---|---|---|---|---|
| eviction | 108 | 0.12 s | 0.18 s | 65.0 s | 86% | 100% | 100% |
| eviction + 10% holdback | 34 | 0.12 s | 0.13 s | 64.3 s | 88% | 100% | 100% |
| eviction + 20% holdback | 1 | 0.12 s | 0.13 s | 65.1 s | 98% | 100% | 100% |
steady 4 req/s
| arm | evictions | TTFT p50 | TTFT p95 | batch done | job a met | job b met | job c met |
|---|---|---|---|---|---|---|---|
| eviction | 219 | 0.12 s | 0.18 s | 70.1 s | 80% | 100% | 100% |
| eviction + 10% holdback | 68 | 0.11 s | 0.12 s | 65.9 s | 78% | 100% | 100% |
| eviction + 20% holdback | 18 | 0.12 s | 0.13 s | 72.9 s | 89% | 100% | 100% |
Bursty interactive load
15 bursts of 32 interactive requests, 6 s apart. Includes the arms without eviction for reference.
| arm | evictions | TTFT p50 | TTFT p95 | batch done | job a met | job b met | job c met |
|---|---|---|---|---|---|---|---|
| no eviction | 0 | 1.80 s | 6.72 s | 75.6 s | 100% | 100% | 100% |
| 10% holdback | 0 | 1.46 s | 6.29 s | 75.1 s | 100% | 100% | 100% |
| 40% holdback | 0 | 1.05 s | 4.93 s | 80.5 s | 100% | 100% | 100% |
| eviction | 350 | 0.22 s | 0.43 s | 69.1 s | 16% | 100% | 100% |
| eviction + 10% holdback | 329 | 0.22 s | 0.41 s | 69.9 s | 0% | 100% | 100% |
| eviction + 20% holdback | 324 | 0.21 s | 0.39 s | 69.7 s | 0% | 100% | 100% |
Expensive evictions: Qwen2.5-14B
On 1.5B an eviction costs almost nothing (17-token prompts, 96% prefix hits on resume). This repeats the steady 1 req/s comparison on Qwen2.5-14B with ~4.8k-token batch prompts × 1,500 tokens (96 requests, loose deadlines) and 500-token interactive prompts × 128 tokens. Evictions fall as on 1.5B, but TTFT p95 is ~0.92 s in every arm, and TTFT p50 is 0.22 s throughout.
steady 1 req/s
| arm | runs | evictions | TTFT p50 | TTFT p95 | batch done |
|---|---|---|---|---|---|
| eviction | 3 | 166 | 0.23 s | 0.92 s | 228.0 s |
| eviction + 10% holdback | 3 | 11 | 0.22 s | 0.92 s | 236.8 s |
| eviction + 20% holdback | 2 | 2 | 0.22 s | 0.92 s | 239.2 s |
Caveats
- Saturation counts requests, not tokens or KV cache. Token and hybrid detector modes are untested.
- Misses on the 45 s job happen only under eviction (the same image meets 100% without it). Leading suspect, not yet confirmed: the processor sends all 192 batch requests to the gateway at once, so its earliest-deadline-first ordering never applies, and a resumed request re-enters the gateway's FCFS batch queue behind later-deadline work.
- The 14B runs did not set a per-run
cache_salt, so prefix-cache state carried across runs; prefill and hit-rate comparisons are left out for that reason. The last 20% holdback run on 14B was lost to a tooling failure, so that arm has 2 runs. - Every completed batch output was token-identical to a direct vLLM run (53 runs).