Upgrading to an Enforced Binding Dtype
From this release, a Binding’s spec.domain.dtype stops being a declaration and becomes the engine’s
--kv-cache-dtype, changing objects a running cluster already holds. The rule itself is stated once,
under
The dtype is handed to the engine
.
Contents
- Changes on upgrade
- Check every Binding first
- Find roles that pass the flag themselves
- If new Pods fail argument parsing
- Verify
Changes on upgrade
- Every pool-attached
ModelDeploymentrolls once. Its replicas gain an argument, which moves their spec hash, so each is replaced at the deployment’s first reconcile after the upgrade. - The
dtypereaches the engine verbatim. A spelling that engine rejects (bf16on vLLM,fp8orfloat16on SGLang) makes every new Pod fail argument parsing, so each replica the rollout replaces stops serving. - A role’s own
--kv-cache-dtypeis refused while it attaches a pool. A deployment already stored with one keeps running on its own value, which wins because it comes later on the command line, until its next update is refused. - A Pod opting into injection is refused when its container passes the flag itself, so its
owner reports
FailedCreateuntil the flag is removed. - A new Binding may not declare
auto. One stored with it keeps working and stays updatable.
Check every Binding first
List every Binding’s dtype, then compare each with the engines of the deployments that name it:
kubectl get kvcachepoolbindings -A \
-o custom-columns=NAMESPACE:.metadata.namespace,NAME:.metadata.name,DTYPE:.spec.domain.dtypeEach engine’s accepted spellings are listed once, under The dtype is handed to the engine . A Binding whose deployments run both engines needs a value in both rows there. A spelling its engine does not accept is what the next two sections recover from, and fixing it before upgrading is cheaper.
Find roles that pass the flag themselves
These deployments keep their own value after the upgrade, and their next edit is refused until the
flag is gone from extraArgs:
kubectl get modeldeployments -A -o json | jq -r '.items[]
| select(.spec.kvCache != null) | . as $md | .spec.roles[]
| select((.command // []) == [] and ((.extraArgs // []) | any(test("^--kv[-_]cache[-_]dtype"))))
| "\($md.metadata.namespace)/\($md.metadata.name) role \(.name)"'Remove the flag once the Binding’s dtype is the one the role should run.
If new Pods fail argument parsing
The dtype is immutable, so the Binding cannot be corrected in place. Turn the Setting off first, which renders and refuses nothing and recreates the replicas without the flag:
kubectl -n gpustack-system patch setting model-deployment-kv-cache-dtype-owned \
--type merge -p '{"spec":{"value":"false"}}'Then move the workloads to a Binding with a spelling the engine accepts, and turn the Setting back on:
- Create a new Binding on the same pool with the corrected
dtypeand a newdomain.name; the old name stays claimed until the old Binding is gone. - Recreate each deployment with
spec.kvCache.poolRef.namenaming the new Binding.kvCacheis an identity field, so it is a new deployment rather than an edit. - Delete the old deployments; the old Binding’s deletion completes once none of them holds it.
- Patch the Setting back to
"true".
The cache under the old domain is not carried over. It was written without a binding dtype, so it is not worth keeping.
Verify
Every replica of a pool-attached role carries the Binding’s value as two entries of its command,
because the operator owns that role’s whole argv and folds every argument into it:
kubectl -n <namespace> get pods -l app.kubernetes.io/instance=<deployment> \
-o jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.spec.containers[0].command}{"\n"}{end}' \
| grep -o -- '--kv-cache-dtype","[^"]*'An injected Pod carries the same two entries at the end of its args, where the webhook appends.