Model Deployment Prefill and Decode
A deployment with Prefill and Decode roles is admitted as one set. Once both roles run, the
operator connects them through a router and a transfer leg. Prefill and
decode
describes the role fields and admission rules.
Contents
- How a pair is wired
- Two ways to configure a pair
- The router block
- The two router fields
- Direct transfer transport
- Different hardware per role
- Addressing a role
How a pair is wired
Without spec.router, kind adds the role discriminator to the engine’s shared-store connector and
nothing pairs the two roles.
With the managed llm-d-router, the connector a native vLLM role runs depends on whether
spec.kvCache is set; the two settings render two different connector documents, not one document
with a field toggled:
spec.kvCache |
Connector rendered | Block path |
|---|---|---|
| omitted | MooncakeConnector alone |
the direct prefill-to-decode transfer, and nothing else |
| set | MultiConnector wrapping MooncakeConnector and MooncakeStoreConnector |
the direct transfer, with the shared pool attached alongside it |
A pair needs no shared pool to hand blocks over, so omitting spec.kvCache is a supported shape.
Attaching a pool adds reuse across deployments.
In both shapes the decode Pod runs llm-d’s routing proxy as a restartable sidecar; it executes the prefill leg named by the router and then forwards the request to the decoder.
On Ascend the same two shapes render with that platform’s own connector names
(MooncakeConnectorV1 for the leg and AscendStoreConnector for the pool), and the sidecar relays
the handshake in its nixlv2 mode, the one whose per-request document matches what that connector
waits for.
That is the only router combination that renders the leg on Ascend. Behind vllm-router an Ascend
pair renders no leg: that router drives its pairs in a handshake vocabulary the Ascend connector
rejects, and behind no router nothing pairs the roles at all. The missing leg is quiet: the
deployment goes Ready, every request is answered normally, and the two roles never exchange a
block. A correctly answering deployment is exactly what makes the failure hard to see.
Both role containers then mount the host’s /usr/local/Ascend/driver tree read-only: the leg’s
transport builds Device RoCE endpoints and reads each NPU’s NIC address through the hccn_tool
that ships with the driver, while the engine image carries the driver libraries but not the tool.
Without the mount the leg renders but the engine dies at startup, unable to resolve a device IP; a node with no driver installation fails the Pod’s volume setup instead, naming the path. The mount ships with the transfer leg alone; an Ascend deployment without one carries no host path.
The leg reads each role’s parallel size from that role’s own command line: a managed role’s
extraArgs, or a take-over role’s whole command, never both. The one degree vLLM accepts as a
literal environment entry, VLLM_DP_SIZE, rides alongside either. A role declaring none renders
1/1, exactly as before.
The command line is the whole of what the leg can see. A degree arriving any other way is invisible
to it: the image’s own entrypoint (never part of the rendered argv), a config file, a flag inside a
sh -c string, --additional-config’s JSON value, an environment value pulled from a ConfigMap or
Secret. An invisibly widened role keeps its 1/1 half.
The connector’s startup assert compares the document against itself, not against the engine. An invisibly widened pair therefore fails loudly only when the document’s decode degree exceeds its prefill one. A pair widened symmetrically starts, answers, and pulls a wrong layout.
Admission holds the visible side of the contract: a degree the command line does not state legibly is refused, and so is a declared per-member width the role’s card request cannot hold; see Admission refusals .
SGLang renders its halves through the engine’s own disaggregation arguments rather than this connector path; the two engines’ legs differ by their handshake alone.
Two ways to configure a pair
Both are complete objects. The only difference is the kvCache block, and it is the difference
between “these two roles hand blocks to each other” and “these two roles hand blocks to each other
and share a pool with every other deployment bound to it”.
Without a shared pool, the ordinary shape for a pair serving one model:
apiVersion: worker.gpustack.ai/v1
kind: ModelDeployment
metadata:
name: qwen-pd
namespace: team-a
spec:
model:
name: Qwen/Qwen2.5-72B-Instruct
engine: # vLLM | SGLang; this example walks the vLLM pair
name: vLLM
version: "0.29.0"
router:
name: llm-d-router # required to pair the roles; takes either engine
roles:
- name: prefill
kind: Prefill
replicas: 2
instanceType: gpustack-nvidia-h20-linux-amd64
resources:
accelerator: 2
- name: decode
kind: Decode
replicas: 2
instanceType: gpustack-nvidia-h20-linux-amd64
resources:
accelerator: 2Each replica of a role becomes its own Kueue pod group: one group per replica, declaring a total
equal to the role’s size. Two roles naming the same instanceType are still separate groups,
each replica composing its own Workload. The example above therefore has four single-Pod
groups, admitted as a set.
A lost replica is replaced once its slot reads empty, and its siblings keep serving. A node drained, a replica preempted for higher-priority work, a kubelet evicting under pressure: each costs the one replica that left, not its siblings. Every replica is its own group with its own Workload, so a departure touches no other replica’s admission.
Why the slot is what a replacement waits for — the group is annotated as serving, so Kueue never releases the finalizer it holds on the departed Pod; only that replica’s Workload being deleted releases it. The replacement is created once no Pod for that replica’s ordinal reads on the API server, never beside a member still listed, because a group over its declared total reads as excess, and Kueue’s answer to the excess is to delete the newcomer.
The replica bounds the blast radius. A loss or an edit inside one replica’s group does not reach
another replica’s group: one Prefill replica turning over leaves its siblings and all of Decode
serving. The groups are still admitted together; an AdmissionCheck holds them until the whole set
has reserved quota, so Prefill still never starts without Decode.
Splitting instanceTypes buys different hardware; the isolation it used to buy, every replica has
by default.
With a shared pool, the same object plus one block:
spec:
# ... everything above, unchanged ...
kvCache:
poolRef:
name: team-a-dram # a KVCachePoolBinding in THIS namespaceThe router block
spec.router is what pairs the roles, and spec.kvCache is not a substitute for it. A
deployment declaring Prefill and Decode with no router is admitted and renders two roles that
nothing routes between: each gets the shared-store connector with its role discriminator, and no
request is ever split across them. Attaching a pool does not change that.
Conversely a router with no pool is complete. What each shape renders is the table under How a pair is wired .
spec.router.replicas is optional, and absent means one. More than one trades cache consistency for
availability: a router scoring on a prefix cache holds that state per replica; upstream reports radix
trees that do not synchronize across replicas and a hit rate falling by ten to twenty percent as a
result. Where replicas do exchange events, the exchange improves load estimation without making two
replicas route alike.
spec.router.extraArgs takes additional flags for the router process. A flag the operator derives
itself is refused rather than merged, so one setting has one source. The owned catalog is keyed by
router, and the refusal names the flag and the router.
For llm-d-router it is --endpoint-selector, --endpoint-target-ports, --config-file,
--secure-serving, --grpc-health-port and --metrics-endpoint-auth.
vllm-router and sglang-gateway take their whole configuration on the command line, so their
catalogs are wider: the two bind addresses and their ports, the four service-discovery flags and
each project’s own spelling of the disaggregation switch. Only vllm-router names a transfer
connector; the gateway has no such flag, because its transfer backend is the engine’s.
A router is also engine-matched, and a pair outside this table is refused naming both sides:
spec.router.name |
Engines | Rendered shape | Routing policy |
|---|---|---|---|
llm-d-router |
vLLM, SGLang |
An endpoint picker behind a proxy, configured by a mounted document | A fixed scoring profile |
vllm-router |
vLLM |
One process, configured entirely by its command line | cache_aware unless extraArgs names another
|
sglang-gateway |
SGLang |
One process, configured entirely by its command line | cache_aware unless extraArgs names another
|
llm-d-router takes both engines because upstream carries a handshake connector and a metrics
configuration for each. The other two are each one project’s router for that project’s own engine,
and admitting a cross pairing would be a claim this repository has not measured.
The two router fields
spec.router.requestTimeoutSeconds is how long the router waits for a reply. Leaving it out does
not mean one thing across the three. It renders nothing, so each router keeps its own upstream
default: one day under llm-d-router, whose proxy carries the timeout, against half an hour
under the two configured by their command line, a factor of forty-eight.
Setting it is what makes a declaration survive a change of router. Zero is refused: it would mean “wait forever” under the proxy and nothing in particular under the other two.
spec.router.disaggregationThresholdTokens is how many prompt tokens not already in a prefix
cache make a request worth splitting between a prefiller and a decoder; below it the decode replica
serves the whole request itself. Unset renders the value the picker already used.
Zero disables splitting entirely rather than meaning “always split” (the decider returns “do not disaggregate” on a zero threshold before reading anything else), and it is accepted because it is a value upstream defines. It is refused on the other two routers rather than ignored: they have no per-request decision to threshold, and a field that is legal to write and renders nothing is a shape this API has rejected before.
The proxy under llm-d-router also writes one access log: a JSON object per request on the
container’s standard output, carrying the response code, the response flags, the response code
details, the duration, the endpoint the picker selected, the method, the path, the request id and the
byte counts. It is not a field: the gap it fills is that nothing is logged at all, so there is no
value to choose. The other two routers log whatever their own flags say.
Two flags are fixed. The router runs with
--secure-serving=false and --metrics-endpoint-auth=false, where upstream defaults both to
true. The inversion is deliberate: a router manages east-west traffic, picking which replica
of this deployment serves a request already inside the cluster, while TLS and caller authentication
are north-south concerns owned by the gateway that admits traffic into the cluster.
Where those two flags sit, neither protects anything. --secure-serving puts TLS on the endpoint
picker’s gRPC server, whose only client is the proxy container in the same Pod dialing
127.0.0.1; --metrics-endpoint-auth guards the picker’s own /metrics, scraped in-cluster. So
there is no field for either, and neither is reachable through extraArgs.
Direct transfer transport
The point-to-point leg renders Mooncake’s tcp when the API field is unset:
spec:
kvTransfer:
protocol: RDMA # unset renders "tcp"The value is a property of one link, so it is deployment-wide: a per-role field could only express two ends naming different protocols for one connection, which fails at transfer time rather than at admission.
The API accepts Auto, TCP, RDMA, EFA, CANN and ROCM,
using the same Mooncake mapping as KVCacheBackend: Auto and TCP render tcp, CANN
renders ascend, and ROCM renders hip. MUSA and MACA are excluded because they are
same-node IPC transports and the two roles may run on different nodes.
The engine image must still carry a build for its transport. SGLang maps tcp to
--disaggregation-transfer-backend mooncake_tcp and every other mapped value to mooncake.
Which value works on which engine is the transport matrix
.
tcp is enforced, not only requested: the transfer engine picks its transport from the host,
and with no RDMA device a build with multi-node NVLink installs NVLink between hosts with no NVLink
path. So native vLLM also gets the defaulted
MC_FORCE_TCP=1. Both pins
are process-wide, so neither renders beside a store on another transport, and every client at
its engine’s supported minimum
honors them.
It is read on the direct-transfer leg, which every admitted router-and-engine pair renders on its
Prefill and Decode roles: a prefiller that cannot hand a decoder its blocks is not disaggregated
under any router. On every other shape the field is accepted and renders nothing.
What differs per pair is the handshake, not whether there is a leg: Mooncake’s bootstrap server
under native vLLM, SGLang’s own registry under SGLang, and on Ascend the decode sidecar’s relay
(the one router combination that renders a leg there
). There the field is
ignored: the tested vLLM-Ascend v0.23.0 and v0.26.0rc1 builds use ascend on this leg.
It is separate from the pool’s transport. KVCacheBackend.spec.transport feeds the engine’s store
client; this leg is engine to engine and never traverses the store, so the two declare separately: a
deployment with no kvCache block still has this leg to configure.
Editing it turns over every role
: the value renders
into both ends’ arguments, so every role’s replicas turn over one at a time. A prefiller and a decoder
can disagree on the protocol until both converge, the same window an engine.version edit opens.
Different hardware per role
A Kueue Workload carries one queueName, and that name is the one the role’s instanceType
publishes as its status.entrance. Every replica is its own group regardless, so two roles were
never going to share a Workload, on two instanceTypes or on one. The set is admitted together by an
admission check rather than by Kueue’s intra-group rule.
See One group per replica for what that costs an edit, and Authoring the InstanceType yourself for which queues carry the check.
Adding a second manufacturer uses the same mechanism. A queue’s accelerator quota is
credits.gpustack.ai/<manufacturer>, one resource name per manufacturer, and Kueue’s own webhook
refuses a second resource group repeating a covered resource within one queue. With a queue per role
there is no second group to repeat anything.
Two roles on different manufacturers cannot share KV through their pool. The deployment is admitted and both halves serve; what fails is the sharing, and it fails silently: reads from the shared store miss, and nothing on the deployment reports it. Three independent reasons stand between the halves:
- This operator’s own refusal is the one an administrator meets. On a pool left at its default transport, the renderer refuses the vLLM-Ascend half (the rule and its remediation are stated there). The check is one-sided: it fires for a single-manufacturer Ascend deployment just the same, and it is the only one of the three that produces a message; following its remediation clears only this refusal; the next two apply regardless.
- Two upstream limits then apply, described at the versions each engine is supported at (see the minimum per shape ); they describe that pair of releases, not a permanent property of either project. The operator renders no key that lets the Ascend half load from the shared store, and the two engines address it with incompatible keys, so every lookup misses and no error is raised.
This limit governs the shared pool alone: a pair needs no shared pool to hand blocks over .
For the direct transfer, what matters is the router in front, not the manufacturers. The leg renders per role, and on Ascend only one router combination carries it .
A mixed pair never forms a transfer either way: the two sides’ connectors speak different handshake vocabularies, so a relayed document from one names nothing the other reads. Two non-Ascend roles both render it (NVIDIA and AMD, say), and what the engines then do is upstream’s answer, unmeasured here.
Kueue assigns a ResourceFlavor per
PodSet
, so a role still takes whatever
its own pool assigns; what selects the hardware is the instanceType the role names.
Addressing a role
Each role gets a ClusterIP Service named <deployment>-<role>, beside the deployment-wide one, so a
decoder is reachable as a decoder. The managed router uses these stable role addresses for the
tokenizer and cache-event contracts; they also remain useful for addressing one half directly while
debugging.
With spec.router, the operator renders resources named <deployment>-router: a Deployment,
ConfigMap, Service, ServiceAccount, Role and RoleBinding. Removing spec.router prunes these resources.
What status.endpoint publishes in each shape is under Status
.