Model Deployment Routing
A router only chooses when a role has more than one replica. Every routing choice measured here ran on a server role; on a prefill/decode pair, where each half is chosen separately, none has been run.
Contents
- Each router’s default
- Prefix affinity
- Switching to round robin
- llm-d-router takes no policy flag
- Per-worker routing metrics
Each router’s default
vllm-router and sglang-gateway route by cache_aware, and llm-d-router scores every
candidate for each request. The operator renders no policy flag for the first two, so each runs its
upstream default. The table describes the bundled vLLM router v0.1.15 and SGLang gateway
gateway-v0.3.1; the llm-d row describes the operator’s scoring profile.
spec.router.name |
Default | Replica selection |
|---|---|---|
vllm-router |
cache_aware |
The replica whose past requests share the longest prefix with this one, when that share is above --cache-threshold (0.3); below it, the replica with the smallest record. Once the busiest replica has more than --balance-abs-threshold (64) requests over the idlest and more than --balance-rel-threshold (1.5) times as many, the least-loaded replica instead |
sglang-gateway |
cache_aware |
The same algorithm, with the same three defaults |
llm-d-router |
A fixed scoring profile | The highest total of prefix-cache match (weight 3, not scored on a decode half), queue depth (2) and KV-cache utilization (2) |
Both cache_aware routers log the policy once at startup, as policy: CacheAware { cache_threshold: 0.3, … }.
Prefix affinity
Under cache_aware, requests that share a long prefix and arrive one after another all go to one
replica, and the others stay idle. That is the policy working, not a fault: the replica that served
the prefix holds it in its prefix cache, so the next request sharing it does not recompute it.
Such a request only goes elsewhere once the load gap above opens, and traffic that sends one request at a time never has more than one in flight to open it.
Measured on two server replicas, each on its own GPU, with serial chat requests that open with the same text of about 500 characters and end in a question of their own, streaming and non-streaming in turn. No request failed:
| Router | Engine | Requests on each replica |
|---|---|---|
vllm-router |
vLLM 0.29.0 |
354 and 0 |
sglang-gateway |
SGLang 0.5.18 |
406 and 0 |
llm-d-router |
vLLM 0.29.0 |
187 and 170 |
llm-d-router |
SGLang 0.5.18 |
375 and 0 |
How any of the three spreads concurrent traffic is not measured.
Switching to round robin
Add --policy round_robin to spec.router.extraArgs on vllm-router or sglang-gateway to
send each request to the next replica in turn, whatever it shares with earlier ones:
spec:
router:
name: vllm-router # or sglang-gateway
extraArgs:
- --policy
- round_robin--policy is in neither router’s refused list
, so admission
passes it and the operator appends it to the command line it renders. The router then logs
Starting router … | policy: RoundRobin.
Measured on the same two replicas and the same kind of traffic, with the flag set at creation:
under each router, 120 requests with none failing, 60 on each replica. vllm-router counted them
in vllm_router_policy_decisions_total{policy="round_robin"}, 60 per replica, and sglang-gateway
in smg_worker_selection_total{policy="round_robin"}, 120 in all.
router.extraArgs is editable
, so a
running deployment can switch as well. The router Pod is then replaced, and the new one starts with
no record of earlier prefixes. Only a flag set at creation has been run.
The value is each project’s own, and the router checks it, not admission, so a misspelled one stops the router at startup:
spec.router.name |
--policy values |
|---|---|
vllm-router |
random, round_robin, cache_aware, power_of_two, consistent_hash, rendezvous_hash |
sglang-gateway |
random, round_robin, cache_aware, power_of_two, prefix_hash, manual |
Only round_robin has been run. The three cache_aware thresholds pass through extraArgs the same
way, as do both routers’ --prefill-policy and --decode-policy for a prefill/decode pair; none of
those has been run either.
llm-d-router takes no policy flag
llm-d-router cannot be switched through extraArgs. Its scorers and their weights are in the
configuration document the operator renders and mounts, and --config-file, which would point the
router at another one, is refused there. No field sets them either.
Per-worker routing metrics
The deployment’s metrics snapshot
does not say which replica served a
request. Each router’s own /metrics does, for the series below, all seen exported in runs. A
worker label is the worker’s URL, which carries its Pod address.
spec.router.name |
Series | Meaning |
|---|---|---|
vllm-router |
vllm_router_processed_requests_total{worker} |
requests sent to each worker, in both modes |
vllm-router |
vllm_router_policy_decisions_total{policy,worker} |
picks each policy made, per worker |
vllm-router |
vllm_router_pd_prefill_requests_total{worker}, vllm_router_pd_decode_requests_total{worker} |
in P/D mode, requests sent to each prefill and each decode worker |
sglang-gateway |
smg_worker_requests_active{worker} |
requests in flight to each worker at the scrape |
sglang-gateway |
smg_worker_selection_total{worker_type,policy} |
picks per policy and per regular, prefill or decode worker type, not per worker |
llm-d-router |
llm_d_epp_per_endpoint_queue_size{model_server_endpoint} |
requests queued at each endpoint at the scrape |
llm-d-router |
llm_d_epp_disagg_decision_total{decision_type} |
requests sent through a prefill replica (prefill-decode) or straight to decode (decode-only) |
Neither sglang-gateway nor llm-d-router keeps a running count per replica. Behind them, compare
the engine Pods’ own counts instead, such as each Pod’s TTFT histogram _count.