# KV Cache Walkthrough

Create a store, a pool, a namespace grant and a workload in that order.

On a cluster with GPUStack installed, replace the node selector, namespace, instance type and model
name in the manifests below, then apply them in that order.

## Contents

- [Creation order](#creation-order)
- [Step 1: the store](#step-1-the-store)
- [Step 2: the pool and the grant](#step-2-the-pool-and-the-grant)
- [Step 3: the workload](#step-3-the-workload)
- [Step 4: high availability](#step-4-high-availability)
- [Step 5: post-failover state](#step-5-post-failover-state)
- [Failure modes](#failure-modes)

## Creation order

| Object | Scope | Creator | Purpose |
|---|---|---|---|
| `KVCacheBackend` | cluster | administrator | which nodes contribute memory, and how much |
| `KVCachePool` | cluster | administrator | how much of that store may be handed out at all |
| `KVCachePoolBinding` | namespace | administrator | that THIS namespace may draw on it, and up to what |
| `ModelDeployment` | namespace | user | a workload that attaches to the grant |

The order is fixed because each object names the one above it, so creating them the other way round
leaves references that resolve to nothing. Why the two scopes split this way is on
[KV Cache Pool](/gpustack-operator/main/docs/modules/kv-cache/pool/index.md#two-kinds-split-by-scope).

**The Binding is the authorization point.** A namespace gets access to a store when an administrator
creates a Binding in it. A user naming a pool is not a path this API has: `poolRef` is a
same-namespace reference, so the name it accepts is a Binding's.

## Step 1: the store

```yaml
apiVersion: worker.gpustack.ai/v1
kind: KVCacheBackend
metadata:
  name: mooncake-dram
spec:
  type: Mooncake
  connection:
    managed:
      leader: {}
      members:
        - nodeSelector:
            kubernetes.io/os: linux      # replace with the nodes that should contribute
          medium: DRAM
          capacityPerMember: 8Gi
```

`spec.image` is left unset on purpose: the cluster-wide `kv-cache-backend-image`
[Setting](/gpustack-operator/main/docs/reference/settings/index.md) supplies this project's own build, which is the one that can elect a leader
later. Naming a published upstream image here works until Step 4 and then does not. See
[High availability](/gpustack-operator/main/docs/modules/kv-cache/leader/index.md#high-availability).

`capacityPerMember` is charged to each member Pod's host memory request, so it is a claim on the node
and not a hint. One member Pod runs per node the selector matches.

`leader: {}` takes the field defaults, `multiTenancy` included (it defaults on), so this master keeps
a per-tenant quota ledger and Steps 2 and 3 read against it. A backend pinned to a store image from
before Mooncake 0.3.12 is the one exception and declares `multiTenancy: false` out loud. See
[The project's own build variants](/gpustack-operator/main/docs/modules/kv-cache/backend/index.md#the-projects-own-build-variants).

Wait for it, then read what it reports:

```console
$ kubectl get kvcb mooncake-dram -w
NAME            TYPE       PHASE   ENDPOINT                                         CAPACITY
mooncake-dram   Mooncake   Ready   mooncake-dram-leader.gpustack-system.svc:50051   16Gi
```

**`CAPACITY` is read from the store, never derived from the spec.** Two nodes at `8Gi` show `16Gi`
because both members registered their segments, so this figure is the evidence that the store is
really assembled. A number smaller than expected means a member has not registered yet, whatever the
phase says: `kubectl describe kvcb mooncake-dram` and its `MembersMounted` condition name which.

## Step 2: the pool and the grant

```yaml
apiVersion: worker.gpustack.ai/v1
kind: KVCachePool
metadata:
  name: shared-dram
spec:
  backends: # exactly one
    - mooncake-dram
  quota:
    total: 16Gi
---
apiVersion: worker.gpustack.ai/v1
kind: KVCachePoolBinding
metadata:
  name: team-a
  namespace: team-a
spec:
  poolRef:
    name: shared-dram
  quota:
    ceiling: 8Gi
  domain:                                # every field here is immutable
    blockSize: 64
    dtype: bfloat16
```

**The `domain` is the cache's compatibility key, and it is immutable for the reason it exists.** Two
workloads sharing a domain share cached blocks, so a domain that could be edited would let a running
workload start reading blocks another tokenizer wrote. Pick it to match the model and engine
settings the deployments in this namespace will run; a second, different model gets a second Binding.

**`domain.name` is left out, so it is `default`.** The backend from Step 1 runs with multi-tenancy
(`leader.multiTenancy` defaults on), so the engines are handed the tenant id `default`, which is also
the tenant the store uses for a writer that names none. Name each domain once a second Binding shares
the master; the name is what keeps the two apart.

**A quota ceiling is not a reservation.** It is the most this namespace may hold at once, and going
over it does not fail a write. See
[Full-quota behavior](/gpustack-operator/main/docs/modules/kv-cache/pool/index.md#full-quota-behavior).

**The ceiling is enforced because the Step-1 leader carries its tenant ledger, which the default
gave it.** A backend declared `leader.multiTenancy: false` holds no ledger instead: its pool is
admitted with a warning, the Binding reports `QuotaGranted=True` with reason `Unenforced` and no
`EFFECTIVE` figure, and every write lands in the store's default tenant.

**A multi-tenant master refuses a tenant name absent from its ledger.** An engine that ignores the
injected tenant then needs a second Binding whose domain is `default`, or that leaves `name` out.
See [Tenant compatibility](/gpustack-operator/main/docs/modules/kv-cache/injection/index.md#tenant-compatibility).

## Step 3: the workload

```yaml
apiVersion: worker.gpustack.ai/v1
kind: ModelDeployment
metadata:
  name: qwen-chat
  namespace: team-a
spec:
  model:
    name: Qwen/Qwen2.5-7B-Instruct
  engine:
    name: vLLM
    version: "0.29.0"                    # on the same store line as Step 1's default image
  kvCache:
    poolRef:
      name: team-a                       # the Binding above, in this namespace
  roles:
    - name: server
      replicas: 2
      instanceType: <one from kubectl get instancetype>
      resources:
        accelerator: 1
```

`kvCache` is optional. Omit it and the deployment runs with no shared cache at all, which is the
useful comparison to have run once before attributing anything to the pool.

**The engine is configured by injection, not by anything you write here.** The operator resolves the
Binding, reads the backend's endpoint and the domain, and injects the store's environment into every
role's Pod. What is injected per engine, and every refusal, is on
[KV Cache Injection](/gpustack-operator/main/docs/modules/kv-cache/injection/index.md).

**The store version must match the engine's client.** The engine version above decides the client,
and Step 1's unset `spec.image` leaves the store on the Settings default. A pair on different lines
starts healthy and then fails every write, so the version above is a choice. See
[The store version must match the engine's client](/gpustack-operator/main/docs/modules/kv-cache/backend/index.md#the-store-version-must-match-the-engines-client).

Check the attachment from the deployment's own status rather than from the Pods:

```console
$ kubectl -n team-a get md qwen-chat -o jsonpath='{.status.conditions[?(@.type=="CacheAttached")]}'
```

## Step 4: high availability

Everything so far runs one leader process that already holds a Kubernetes Lease. An update or a node
failure takes the store's metadata with it. Add standbys with one edit:

```yaml
spec:
  connection:
    managed:
      leader:
        replicas: 3
```

**The healthy steady state now reads `3 desired / 1 ready`, and that is not a broken Deployment.**
Exactly one leader serves; the other two are standbys, deliberately not ready so the leader Service's
endpoints never include a process that cannot serve. During a healthy failover `2` are briefly ready
as the old leader steps down. Both readings are normal.

**Scaling from one to three keeps the first leader Pod and the member templates** because the Lease
election and accounts were already present at one replica. See
[the leader Deployment](/gpustack-operator/main/docs/modules/kv-cache/leader/index.md#the-deployment-and-the-two-probes) for the
cost of changing `leader.electionBackend`.

**`leader.memberAddressing` chooses how members find the master**, and defaults to
`Service`. An explicit `Lease` value uses the member's API access to read the current holder. See
[High availability](/gpustack-operator/main/docs/modules/kv-cache/leader/index.md#high-availability) for the measured failover limits.

These conditions apply at one replica too; the phase does not summarize them:

| Condition | True | False |
|---|---|---|
| `ElectionObserved` | the Lease names a holder, so an election happened | a ready leader and a holderless Lease — the image, the role binding, or a first campaign still running |
| `RolloutComplete` | every replica runs the current template and none of the previous one is left | a rollout in flight, or one that has stalled |

`RolloutComplete` exists because this workload disables the deployment deadline that would normally
answer it: that deadline requires every replica to be available, and only one ever is here.

## Step 5: post-failover state

**A failover keeps nothing in memory.** A standby holds no data, so the replica that takes over knows
none of the objects held in member memory and rebuilds from member remounts alone; a single leader
that restarts does the same. High availability shortens the outage, and every one of those objects
misses afterwards. Plan for a cold cache after every failover and every leader restart.

**The store's snapshot is not offered, and its flags are refused in `leader.extraArgs`**, because
restoring a snapshot can make the cache serve another key's bytes instead of a miss.
[High availability](/gpustack-operator/main/docs/modules/kv-cache/leader/index.md#high-availability) says why.

## Failure modes

**A published upstream store image, under high availability.** No published `kvcacheai/mooncake`
image carries a leadership backend: the leader answers `UNAVAILABLE_IN_CURRENT_MODE` and runs as a
permanent standby, and members answer `Invalid HA backend entry` and CrashLoopBackOff. `spec.image`
and every `members[].image` need a build that has one. Leaving them unset is the simplest way.

**Expecting a failover or a restart to keep the cache.** Neither does: the new leader knows none of
the objects held in member memory, so every lookup for one misses until an engine writes it again.

**Reading `3/1` as a fault.** It is the designed steady state under high availability, and the one
reading on this page most likely to be escalated as an outage.

---

**See also** — [KV Cache Backend](/gpustack-operator/main/docs/modules/kv-cache/backend/index.md) (every field of the store, and what status reports) ·
[KV Cache Leader](/gpustack-operator/main/docs/modules/kv-cache/leader/index.md) (the election, why there is no snapshot, and the member addressing choice in full) ·
[KV Cache Pool](/gpustack-operator/main/docs/modules/kv-cache/pool/index.md) (quota, domains and what a full quota does) ·
[Model Deployment](/gpustack-operator/main/docs/modules/model-deployment/deployment/index.md) (roles, prefill/decode, rollout) ·
[Model Deployment Prefill and Decode](/gpustack-operator/main/docs/modules/model-deployment/prefill-decode/index.md) (router, transfer) ·
[KV Cache Injection](/gpustack-operator/main/docs/modules/kv-cache/injection/index.md) (what a Pod actually receives)

**Next** → [KV Cache Pool](/gpustack-operator/main/docs/modules/kv-cache/pool/index.md) — the quota this walkthrough set once and did not explain.
