GPUStack Operator

KV Cache Walkthrough

Create a store, a pool, a namespace grant and a workload in that order.

On a cluster with GPUStack installed, replace the node selector, namespace, instance type and model name in the manifests below, then apply them in that order.

Contents

Creation order

Object Scope Creator Purpose
KVCacheBackend cluster administrator which nodes contribute memory, and how much
KVCachePool cluster administrator how much of that store may be handed out at all
KVCachePoolBinding namespace administrator that THIS namespace may draw on it, and up to what
ModelDeployment namespace user a workload that attaches to the grant

The order is fixed because each object names the one above it, so creating them the other way round leaves references that resolve to nothing. Why the two scopes split this way is on KV Cache Pool .

The Binding is the authorization point. A namespace gets access to a store when an administrator creates a Binding in it. A user naming a pool is not a path this API has: poolRef is a same-namespace reference, so the name it accepts is a Binding’s.

Step 1: the store

yaml
apiVersion: worker.gpustack.ai/v1
kind: KVCacheBackend
metadata:
  name: mooncake-dram
spec:
  type: Mooncake
  connection:
    managed:
      leader: {}
      members:
        - nodeSelector:
            kubernetes.io/os: linux      # replace with the nodes that should contribute
          medium: DRAM
          capacityPerMember: 8Gi

spec.image is left unset on purpose: the cluster-wide kv-cache-backend-image Setting supplies this project’s own build, which is the one that can elect a leader later. Naming a published upstream image here works until Step 4 and then does not. See High availability .

capacityPerMember is charged to each member Pod’s host memory request, so it is a claim on the node and not a hint. One member Pod runs per node the selector matches.

leader: {} takes the field defaults, multiTenancy included (it defaults on), so this master keeps a per-tenant quota ledger and Steps 2 and 3 read against it. A backend pinned to a store image from before Mooncake 0.3.12 is the one exception and declares multiTenancy: false out loud. See The project’s own build variants .

Wait for it, then read what it reports:

console
$ kubectl get kvcb mooncake-dram -w
NAME            TYPE       PHASE   ENDPOINT                                         CAPACITY
mooncake-dram   Mooncake   Ready   mooncake-dram-leader.gpustack-system.svc:50051   16Gi

CAPACITY is read from the store, never derived from the spec. Two nodes at 8Gi show 16Gi because both members registered their segments, so this figure is the evidence that the store is really assembled. A number smaller than expected means a member has not registered yet, whatever the phase says: kubectl describe kvcb mooncake-dram and its MembersMounted condition name which.

Step 2: the pool and the grant

yaml
apiVersion: worker.gpustack.ai/v1
kind: KVCachePool
metadata:
  name: shared-dram
spec:
  backends: # exactly one
    - mooncake-dram
  quota:
    total: 16Gi
---
apiVersion: worker.gpustack.ai/v1
kind: KVCachePoolBinding
metadata:
  name: team-a
  namespace: team-a
spec:
  poolRef:
    name: shared-dram
  quota:
    ceiling: 8Gi
  domain:                                # every field here is immutable
    blockSize: 64
    dtype: bfloat16

The domain is the cache’s compatibility key, and it is immutable for the reason it exists. Two workloads sharing a domain share cached blocks, so a domain that could be edited would let a running workload start reading blocks another tokenizer wrote. Pick it to match the model and engine settings the deployments in this namespace will run; a second, different model gets a second Binding.

domain.name is left out, so it is default. The backend from Step 1 runs with multi-tenancy (leader.multiTenancy defaults on), so the engines are handed the tenant id default, which is also the tenant the store uses for a writer that names none. Name each domain once a second Binding shares the master; the name is what keeps the two apart.

A quota ceiling is not a reservation. It is the most this namespace may hold at once, and going over it does not fail a write. See Full-quota behavior .

The ceiling is enforced because the Step-1 leader carries its tenant ledger, which the default gave it. A backend declared leader.multiTenancy: false holds no ledger instead: its pool is admitted with a warning, the Binding reports QuotaGranted=True with reason Unenforced and no EFFECTIVE figure, and every write lands in the store’s default tenant.

A multi-tenant master refuses a tenant name absent from its ledger. An engine that ignores the injected tenant then needs a second Binding whose domain is default, or that leaves name out. See Tenant compatibility .

Step 3: the workload

yaml
apiVersion: worker.gpustack.ai/v1
kind: ModelDeployment
metadata:
  name: qwen-chat
  namespace: team-a
spec:
  model:
    name: Qwen/Qwen2.5-7B-Instruct
  engine:
    name: vLLM
    version: "0.29.0"                    # on the same store line as Step 1's default image
  kvCache:
    poolRef:
      name: team-a                       # the Binding above, in this namespace
  roles:
    - name: server
      replicas: 2
      instanceType: <one from kubectl get instancetype>
      resources:
        accelerator: 1

kvCache is optional. Omit it and the deployment runs with no shared cache at all, which is the useful comparison to have run once before attributing anything to the pool.

The engine is configured by injection, not by anything you write here. The operator resolves the Binding, reads the backend’s endpoint and the domain, and injects the store’s environment into every role’s Pod. What is injected per engine, and every refusal, is on KV Cache Injection .

The store version must match the engine’s client. The engine version above decides the client, and Step 1’s unset spec.image leaves the store on the Settings default. A pair on different lines starts healthy and then fails every write, so the version above is a choice. See The store version must match the engine’s client .

Check the attachment from the deployment’s own status rather than from the Pods:

console
$ kubectl -n team-a get md qwen-chat -o jsonpath='{.status.conditions[?(@.type=="CacheAttached")]}'

Step 4: high availability

Everything so far runs one leader process that already holds a Kubernetes Lease. An update or a node failure takes the store’s metadata with it. Add standbys with one edit:

yaml
spec:
  connection:
    managed:
      leader:
        replicas: 3

The healthy steady state now reads 3 desired / 1 ready, and that is not a broken Deployment. Exactly one leader serves; the other two are standbys, deliberately not ready so the leader Service’s endpoints never include a process that cannot serve. During a healthy failover 2 are briefly ready as the old leader steps down. Both readings are normal.

Scaling from one to three keeps the first leader Pod and the member templates because the Lease election and accounts were already present at one replica. See the leader Deployment for the cost of changing leader.electionBackend.

leader.memberAddressing chooses how members find the master, and defaults to Service. An explicit Lease value uses the member’s API access to read the current holder. See High availability for the measured failover limits.

These conditions apply at one replica too; the phase does not summarize them:

Condition True False
ElectionObserved the Lease names a holder, so an election happened a ready leader and a holderless Lease — the image, the role binding, or a first campaign still running
RolloutComplete every replica runs the current template and none of the previous one is left a rollout in flight, or one that has stalled

RolloutComplete exists because this workload disables the deployment deadline that would normally answer it: that deadline requires every replica to be available, and only one ever is here.

Step 5: post-failover state

A failover keeps nothing in memory. A standby holds no data, so the replica that takes over knows none of the objects held in member memory and rebuilds from member remounts alone; a single leader that restarts does the same. High availability shortens the outage, and every one of those objects misses afterwards. Plan for a cold cache after every failover and every leader restart.

The store’s snapshot is not offered, and its flags are refused in leader.extraArgs, because restoring a snapshot can make the cache serve another key’s bytes instead of a miss. High availability says why.

Failure modes

A published upstream store image, under high availability. No published kvcacheai/mooncake image carries a leadership backend: the leader answers UNAVAILABLE_IN_CURRENT_MODE and runs as a permanent standby, and members answer Invalid HA backend entry and CrashLoopBackOff. spec.image and every members[].image need a build that has one. Leaving them unset is the simplest way.

Expecting a failover or a restart to keep the cache. Neither does: the new leader knows none of the objects held in member memory, so every lookup for one misses until an engine writes it again.

Reading 3/1 as a fault. It is the designed steady state under high availability, and the one reading on this page most likely to be escalated as an outage.