KV Cache Walkthrough
Create a store, a pool, a namespace grant and a workload in that order.
On a cluster with GPUStack installed, replace the node selector, namespace, instance type and model name in the manifests below, then apply them in that order.
Contents
- Creation order
- Step 1: the store
- Step 2: the pool and the grant
- Step 3: the workload
- Step 4: high availability
- Step 5: post-failover state
- Failure modes
Creation order
| Object | Scope | Creator | Purpose |
|---|---|---|---|
KVCacheBackend |
cluster | administrator | which nodes contribute memory, and how much |
KVCachePool |
cluster | administrator | how much of that store may be handed out at all |
KVCachePoolBinding |
namespace | administrator | that THIS namespace may draw on it, and up to what |
ModelDeployment |
namespace | user | a workload that attaches to the grant |
The order is fixed because each object names the one above it, so creating them the other way round leaves references that resolve to nothing. Why the two scopes split this way is on KV Cache Pool .
The Binding is the authorization point. A namespace gets access to a store when an administrator
creates a Binding in it. A user naming a pool is not a path this API has: poolRef is a
same-namespace reference, so the name it accepts is a Binding’s.
Step 1: the store
apiVersion: worker.gpustack.ai/v1
kind: KVCacheBackend
metadata:
name: mooncake-dram
spec:
type: Mooncake
connection:
managed:
leader: {}
members:
- nodeSelector:
kubernetes.io/os: linux # replace with the nodes that should contribute
medium: DRAM
capacityPerMember: 8Gispec.image is left unset on purpose: the cluster-wide kv-cache-backend-image
Setting
supplies this project’s own build, which is the one that can elect a leader
later. Naming a published upstream image here works until Step 4 and then does not. See
High availability
.
capacityPerMember is charged to each member Pod’s host memory request, so it is a claim on the node
and not a hint. One member Pod runs per node the selector matches.
leader: {} takes the field defaults, multiTenancy included (it defaults on), so this master keeps
a per-tenant quota ledger and Steps 2 and 3 read against it. A backend pinned to a store image from
before Mooncake 0.3.12 is the one exception and declares multiTenancy: false out loud. See
The project’s own build variants
.
Wait for it, then read what it reports:
$ kubectl get kvcb mooncake-dram -w
NAME TYPE PHASE ENDPOINT CAPACITY
mooncake-dram Mooncake Ready mooncake-dram-leader.gpustack-system.svc:50051 16Gi
CAPACITY is read from the store, never derived from the spec. Two nodes at 8Gi show 16Gi
because both members registered their segments, so this figure is the evidence that the store is
really assembled. A number smaller than expected means a member has not registered yet, whatever the
phase says: kubectl describe kvcb mooncake-dram and its MembersMounted condition name which.
Step 2: the pool and the grant
apiVersion: worker.gpustack.ai/v1
kind: KVCachePool
metadata:
name: shared-dram
spec:
backends: # exactly one
- mooncake-dram
quota:
total: 16Gi
---
apiVersion: worker.gpustack.ai/v1
kind: KVCachePoolBinding
metadata:
name: team-a
namespace: team-a
spec:
poolRef:
name: shared-dram
quota:
ceiling: 8Gi
domain: # every field here is immutable
blockSize: 64
dtype: bfloat16The domain is the cache’s compatibility key, and it is immutable for the reason it exists. Two
workloads sharing a domain share cached blocks, so a domain that could be edited would let a running
workload start reading blocks another tokenizer wrote. Pick it to match the model and engine
settings the deployments in this namespace will run; a second, different model gets a second Binding.
domain.name is left out, so it is default. The backend from Step 1 runs with multi-tenancy
(leader.multiTenancy defaults on), so the engines are handed the tenant id default, which is also
the tenant the store uses for a writer that names none. Name each domain once a second Binding shares
the master; the name is what keeps the two apart.
A quota ceiling is not a reservation. It is the most this namespace may hold at once, and going over it does not fail a write. See Full-quota behavior .
The ceiling is enforced because the Step-1 leader carries its tenant ledger, which the default
gave it. A backend declared leader.multiTenancy: false holds no ledger instead: its pool is
admitted with a warning, the Binding reports QuotaGranted=True with reason Unenforced and no
EFFECTIVE figure, and every write lands in the store’s default tenant.
A multi-tenant master refuses a tenant name absent from its ledger. An engine that ignores the
injected tenant then needs a second Binding whose domain is default, or that leaves name out.
See Tenant compatibility
.
Step 3: the workload
apiVersion: worker.gpustack.ai/v1
kind: ModelDeployment
metadata:
name: qwen-chat
namespace: team-a
spec:
model:
name: Qwen/Qwen2.5-7B-Instruct
engine:
name: vLLM
version: "0.29.0" # on the same store line as Step 1's default image
kvCache:
poolRef:
name: team-a # the Binding above, in this namespace
roles:
- name: server
replicas: 2
instanceType: <one from kubectl get instancetype>
resources:
accelerator: 1kvCache is optional. Omit it and the deployment runs with no shared cache at all, which is the
useful comparison to have run once before attributing anything to the pool.
The engine is configured by injection, not by anything you write here. The operator resolves the Binding, reads the backend’s endpoint and the domain, and injects the store’s environment into every role’s Pod. What is injected per engine, and every refusal, is on KV Cache Injection .
The store version must match the engine’s client. The engine version above decides the client,
and Step 1’s unset spec.image leaves the store on the Settings default. A pair on different lines
starts healthy and then fails every write, so the version above is a choice. See
The store version must match the engine’s client
.
Check the attachment from the deployment’s own status rather than from the Pods:
$ kubectl -n team-a get md qwen-chat -o jsonpath='{.status.conditions[?(@.type=="CacheAttached")]}'
Step 4: high availability
Everything so far runs one leader process that already holds a Kubernetes Lease. An update or a node failure takes the store’s metadata with it. Add standbys with one edit:
spec:
connection:
managed:
leader:
replicas: 3The healthy steady state now reads 3 desired / 1 ready, and that is not a broken Deployment.
Exactly one leader serves; the other two are standbys, deliberately not ready so the leader Service’s
endpoints never include a process that cannot serve. During a healthy failover 2 are briefly ready
as the old leader steps down. Both readings are normal.
Scaling from one to three keeps the first leader Pod and the member templates because the Lease
election and accounts were already present at one replica. See
the leader Deployment
for the
cost of changing leader.electionBackend.
leader.memberAddressing chooses how members find the master, and defaults to
Service. An explicit Lease value uses the member’s API access to read the current holder. See
High availability
for the measured failover limits.
These conditions apply at one replica too; the phase does not summarize them:
| Condition | True | False |
|---|---|---|
ElectionObserved |
the Lease names a holder, so an election happened | a ready leader and a holderless Lease — the image, the role binding, or a first campaign still running |
RolloutComplete |
every replica runs the current template and none of the previous one is left | a rollout in flight, or one that has stalled |
RolloutComplete exists because this workload disables the deployment deadline that would normally
answer it: that deadline requires every replica to be available, and only one ever is here.
Step 5: post-failover state
A failover keeps nothing in memory. A standby holds no data, so the replica that takes over knows none of the objects held in member memory and rebuilds from member remounts alone; a single leader that restarts does the same. High availability shortens the outage, and every one of those objects misses afterwards. Plan for a cold cache after every failover and every leader restart.
The store’s snapshot is not offered, and its flags are refused in leader.extraArgs, because
restoring a snapshot can make the cache serve another key’s bytes instead of a miss.
High availability
says why.
Failure modes
A published upstream store image, under high availability. No published kvcacheai/mooncake
image carries a leadership backend: the leader answers UNAVAILABLE_IN_CURRENT_MODE and runs as a
permanent standby, and members answer Invalid HA backend entry and CrashLoopBackOff. spec.image
and every members[].image need a build that has one. Leaving them unset is the simplest way.
Expecting a failover or a restart to keep the cache. Neither does: the new leader knows none of the objects held in member memory, so every lookup for one misses until an engine writes it again.
Reading 3/1 as a fault. It is the designed steady state under high availability, and the one
reading on this page most likely to be escalated as an outage.