# KV Cache Backend

A `KVCacheBackend` configures shared KV cache for inference workloads. With a managed backend,
the operator runs a leader for metadata and a member group with one store process per selected
node. An external backend connects to an existing service.

Mooncake calls the leader a master. Its flags, environment variables and metrics use that name.

## Contents

- [Connection and medium](#connection-and-medium)
- [The image](#the-image)
- [The metadata plane](#the-metadata-plane)
- [The members](#the-members)
- [Status and conditions](#status-and-conditions)
- [Growing and shrinking a group](#growing-and-shrinking-a-group)
- [The external mode](#the-external-mode)
- [Operating notes](#operating-notes)

## Connection and medium

`connection` chooses whether the operator manages the backend or connects to an external service.
For managed backends, `members[].medium` chooses host memory (`DRAM`) or accelerator memory (`VRAM`).

```yaml
apiVersion: worker.gpustack.ai/v1
kind: KVCacheBackend                 # cluster-scoped, short name kvcb
metadata:
  name: mooncake-dram
spec:
  type: Mooncake
  image: docker.io/kvcacheai/mooncake:0.3.13
  connection:
    managed:                         # or external: — exactly one
      leader: {}                     # replicas, allocationStrategy and multiTenancy default
      members:
        - nodeSelector:
            kubernetes.io/os: linux
          medium: DRAM               # what this group's SEGMENT is made of: DRAM or VRAM
          capacityPerMember: 4Gi
```

The example uses Mooncake `0.3.13`, which is compatible with the documented vLLM client. SGLang
requires a different version; check [Engine Versions](/gpustack-operator/main/docs/modules/model-deployment/engine-versions/index.md) before
choosing `spec.image`.

`connection.managed` and `connection.external` are both optional pointers and **exactly one** must be
set; neither and both are refused at admission with a message naming the two. Several member groups
are allowed; at most one of them may carry a [local disk tier](/gpustack-operator/main/docs/modules/kv-cache/local-disk-tier/index.md).

**`members[].medium` takes two values, `DRAM` and `VRAM`**: host memory or device memory. The
renderer treats the two differently. On a `DRAM` group the segment is host memory, so the operator
adds `capacityPerMember` to the member Pod's host memory request. On a `VRAM` group the segment is
device memory, so the Pod request carries nothing for it: the member claims its slice by allocating
it.

One binary is one medium, so a node contributing both does so as two groups selecting it. See
[The members](#the-members) for the renderings each value produces.

The field is **immutable**: a segment already mounted cannot change the kind of memory underneath
the data it holds, so an edit is refused and the choice is made when the group is declared.

An earlier shape offered five values. Four of them named things that are **not member groups**, and
each is reached another way:

| Former `medium` value | What it is | Location |
|---|---|---|
| `LocalDisk` | a tier on the members that already hold the memory replica | [`members[].localDisks`](local-disk-tier.md) |
| `NoF` | an NVMe-oF target coordinate, registered once, with no node affinity and no Pod | no API surface; it is not a member group |
| `CXL` | a DAX device the **leader process** allocates from | nowhere in this API: `enable_cxl`, `cxl_path` and `cxl_size` are refused in `leader.extraArgs`, because the first replaces `leader.allocationStrategy` and then brings the leader up advertising the allocator's size as capacity even where no DAX device exists |
| `DFS` | a distributed filesystem the **leader process** allocates from | the leader's own environment, which this API does not render |

> **Why the shape matters more than the names** — the leader routes an offload task to the client
> holding the key's memory replica. A member group with no memory segment is therefore never chosen,
> so a group declared as "the disk one" would report its disk capacity to the leader and never
> receive a single write. The object would say one thing, the running member another, and nothing
> would report a fault.

What an object written against that earlier shape can still do depends on the API server's
`CRDValidationRatcheting` gate, which follows the cluster's **effective** Kubernetes version:

- **v1.33 and later** — the gate is locked on. An update whose invalid field is **unchanged** is
  admitted, so removing a finalizer works and the object deletes normally.
- **v1.30 through v1.32** — the gate is on by default; a cluster that switches it off behaves like
  an older one.
- **v1.28 and v1.29** — the gate exists but is off by default; turned on, it admits the
  unchanged-field update.
- **Earlier** — the gate does not exist. Every write touching the object is refused, **finalizer
  removal included**, so the object cannot be deleted either. Remove such an object **before**
  applying a CRD that narrows the enum, not after.

No released version of this API carried those values, so the question does not arise on a cluster
that installed a release; this is a development-cluster shape.

The object is **cluster-scoped**: it names nodes, claims host memory and host paths, and on the RDMA
and EFA paths needs `hostNetwork` and a fabric device. Only a cluster administrator can legitimately
declare one.

A pool sets the quota ceiling on a backend; a Binding grants a namespace a share of that quota.
This mirrors Kueue's `ClusterQueue` and `LocalQueue` split
(<https://kueue.sigs.k8s.io/docs/concepts/>). Several pools can reference one backend. The grant
does not enforce access to the store; see [Limitations](/gpustack-operator/main/docs/modules/kv-cache/pool/index.md#limitations).

## The image

**`spec.image` is explicit and never derived from the operator's own image**, which breaks
deliberately with how the Device Manager image is derived from the worker image. Leave it unset and
the cluster-wide `kv-cache-backend-image` Setting supplies it. That Setting ships a default, so a
backend naming no image runs this project's own build; unset in both places (which takes an
administrator clearing the Setting) is refused at admission, naming both.

**Clearing that Setting later does not strand a backend admitted under it.** Admission re-asks for a
fallback only when an update moves `spec.image` itself. Every other update is admitted whatever the
Setting says now. That includes the reconciler's own removal of the finalizer, which would otherwise
leave an object that owns nothing and cannot be deleted.

> **Why** — the master's link-time dependencies differ per published vendor variant, so no single
> derivation is correct. On a CUDA-less host, for example, the base (CUDA 12) master build needs
> `libcuda.so.1` and `libcudart.so.12` and cannot load, while the `-rocm` and `-npu` builds need no
> accelerator library at all and load cleanly.

The client side is where the vendor lives, and it lives in the published engine wheel rather than in
this image: which transports a client can drive is a property of the wheel the engine image carries,
not of `spec.image`.

**The master image needs no accelerator runtime; a member image needs the runtime of the transport it
uses.** An `-npu`-built master on an all-NVIDIA cluster is legitimate, because the master is a pure
metadata service. A member on `ascend`, by contrast, needs CANN (`libascendcl.so`) in its container,
and a CANN-less image fails as a loader error whose own message reaches `status.phaseMessage`.

A member on `EFA` needs libfabric in its image, and one that can drive the node's adapter, not the
distro build. `mirrored-mooncake` installs AWS's, ahead of that copy in its loader cache.

Nothing has to be built to run this **without high availability**.
The example image above is published for amd64 and arm64 and carries **both** `mooncake_master` and
`mc_store_rest_server`, so one `spec.image` serves the leader and the members.
It runs on a host with no GPU: its `libcuda.so.1` is a stub and its `libcudart.so.12` is the real
library, and neither reaches a driver.

**[High availability](/gpustack-operator/main/docs/modules/kv-cache/leader/index.md#high-availability) needs a different image, and for both roles.** That
section carries which build, why no published one will do, and what each role does when handed one
that cannot.

> **Why the stub/real split matters** — the stub alone is enough for the master, but the Python client
> needs a versioned `cudaFreeHost` from a real runtime. An image carrying two stubs runs the master and
> fails every member.

**That split is what the `kv-cache-backend-image` default cannot cover.** Its value is this project's
own CPU build, carrying TCP, RDMA and EFA over DRAM, and one value cannot be right for every backend
at once: which build a member needs depends on the transport its backend asks for and the hardware
its group selects. A backend on a vendor fabric names its `spec.image`, which always wins over the
Setting.

The failure mode moved with the default, and that is what having one costs. Blank made a mismatch an
admission refusal naming both places to fix; a default makes a wrong image a loader error at runtime,
which is quieter and further from whoever can fix it. Clearing the Setting restores the refusal.

### The project's own build variants

Beside the `-cpu` default, this project publishes `mirrored-mooncake` in one build target per vendor
(`cuda`, `cann`, `rocm`), whose base images are dispatch-time build arguments. Tags carry the
toolchain version: `<mooncake-version>-<variant><toolchain>`, such as `0.3.13.post1-cuda13.0`.

Each variant is built on `0.3.13.post1`, the line vLLM's supported clients are on, and on
`0.3.10.post2`, which serves only engines below the [supported
minimum](../model-deployment/engine-versions.md). SGLang's clients are on the 0.3.12 line, which this
project does not build. Which line a backend needs is that table's question.

**A `0.3.10.post2` variant needs `leader.electionBackend: None` and
`leader.multiTenancy: false` written out.** It cannot run the default Kubernetes election.

The tenant ledger hangs on the master's `-enable_multi_tenants` switch, which Mooncake took in 0.3.12,
and the field defaults on. On an older image the default renders a flag the master does not recognize,
and it exits at startup. The explicit false renders no flag, which is the command line such an image
has always run.

**A VRAM group needs a build with VRAM segments compiled in (`USE_VRAM_SEGMENT=ON`), and the stock
`-cpu` default is not one.** VRAM segments exist only on the `0.3.13` line (the `0.3.10.post2`
variants carry the vendor transfer engine without them), so a VRAM group always names a
`0.3.13.post1` variant tag.

Nothing refuses a VRAM group that names no image: the default is an administrator-editable Setting,
so this paragraph is the guidance rather than a gate. A group names its image through
`members[].image`, which wins over the backend's `spec.image`, which wins over the Setting.

**The other direction fails at startup, on purpose.** A build with VRAM segments compiled in takes
device memory whenever it can reach a device, whatever medium the group declared. So a DRAM group is
rendered with every vendor's visibility variable set to `void`, which keeps the accelerators out of
its containers: such an image then finds no device and refuses to start, against the group that
asked for a medium it does not implement.

Without that, the failure was silent and landed elsewhere. The group consumed device memory no Pod
had requested, so neither the scheduler nor the quota chain knew it was gone, and what an operator
saw was an unrelated deployment crash-looping on a figure nobody had configured. Pick the image that
matches the medium; whether an image implements one is a fact about the image.

**A private registry needs `spec.imagePullSecrets`**, and an explicit policy needs
`spec.imagePullPolicy`. Both are backend-wide: they apply to the leader and to every member group,
including a group that names its own `image`. Left unset, the policy is **resolved from the image
tag by the same rule the API server would have applied** (`Always` for `:latest` or no tag,
`IfNotPresent` otherwise), and it is re-resolved whenever the image or the field moves.

> **Why they are fields and not Settings** — the cluster-wide `image-pull-policy` and
> `image-pull-secrets` Settings are values of the bundled-application chart install. They reach the
> subcharts and nothing a controller renders, so a `KVCacheBackend` that inherited them would be the
> only object in this API whose running workloads move when a chart value moves. The service accounts
> [high availability](/gpustack-operator/main/docs/modules/kv-cache/leader/index.md#high-availability) renders carry no registry credentials either (they
> grant Lease access and nothing else), so without these fields no image here could come from a
> private registry at all.

### The store version must match the engine's client

**A store and an engine-embedded Mooncake client interoperate only within one minor line.** The
criterion is the RPC wire signature, not the version string: the handshake answers `2.0.0` for every
0.3.x release, so a mismatched pair is not refused at startup. Every probe reads green, and every
write then fails at transfer time with `RPC_FAIL (-900)`.

Two posts of one minor line share their RPC signatures and interoperate. The 0.3.12 and 0.3.13
lines do not: the method names are unchanged, so the client reaches the handler and mis-decodes the
arguments. Multi-tenancy moves neither: with it off (a declared `multiTenancy: false`, the field
defaulting on) every request resolves to the default tenant.

**The client's version is a property of the engine image, not of anything on this CR.** Which
client each supported engine's runner image carries, and so which line its store runs, is in the
[Engine Versions](/gpustack-operator/main/docs/modules/model-deployment/engine-versions/index.md); an engine below its minimum there is
not supported.

For any other image, read the client off the image in hand rather than off an engine version. A
CUDA build resolves `mooncake-transfer-engine` against a lower bound at build time, so two images of
one vLLM version built on different days can carry different clients.

Direct P/D transfer (`MooncakeConnector`, no `spec.kvCache`) is engine to engine and exempt from this
matching.

Upstream has **no 0.3.12.post2**; the 0.3.12 line ends at 0.3.12.post1. High availability carries
its own per-version rule, a master from 0.3.12 on. See
[High availability](/gpustack-operator/main/docs/modules/kv-cache/leader/index.md#high-availability).

## The metadata plane

**The metadata plane is peer-to-peer and has no API field.** The member's `metadata_server` renders as
the literal `P2PHANDSHAKE`, unconditionally. It needs no etcd or Redis. The default leader election
does use a Kubernetes Lease and its API access, even with one leader replica.

Two axes get confused here, so both are stated. The metadata plane is how clients find one another.
The **HA backend store** (`-enable_ha` with `-ha_backend_type`) is how the leader elects, even at
one replica by default. The Kubernetes Lease is selected by
[`leader.electionBackend`](/gpustack-operator/main/docs/modules/kv-cache/leader/index.md#high-availability), and it moves nothing on the metadata plane.

**A manifest that tries to configure the metadata plane is not refused with a helpful message.**
There is no field, so there is nothing for a webhook to see:

- a strict client (`kubectl apply`'s default) is refused by the schema with
  `strict decoding error: unknown field "spec.metadata"`;
- a client with validation turned off has the block **silently pruned**, and the object is admitted
  and reconciled as though nothing had been written.

The second is indistinguishable from success at the point of apply. The protection against it is the
rule the page states: the metadata plane takes no configuration at all.

## The members

One member group renders **one DaemonSet** over `members[].nodeSelector`. A member is one node's
contribution to the store: it holds that node's host memory, or on a VRAM group its device memory,
plus the node's host paths and, on a host-fabric group, the number of fabric devices the group
requests.

The member's identity is the node, and a DaemonSet is the workload that gives one Pod to every
matching node. Two groups may select the same node (a DRAM group and a VRAM group is the shape the
second medium exists for), and each still renders its own DaemonSet.

The member's whole configuration renders as **environment variables**: no ConfigMap, no volume, no init
container.

| config key | environment variable |
|---|---|
| `local_hostname` | `MOONCAKE_LOCAL_HOSTNAME` (the **pod IP**, from the downward API) |
| `metadata_server` | `MOONCAKE_TE_META_DATA_SERVER` |
| `master_server_address` | `MOONCAKE_MASTER` |
| `protocol` | `MOONCAKE_PROTOCOL` |
| `global_segment_size` | `MOONCAKE_GLOBAL_SEGMENT_SIZE` |
| `local_buffer_size` | `MOONCAKE_LOCAL_BUFFER_SIZE` |
| `device_name` | `MOONCAKE_DEVICE` — **deliberately left unset**, see below |

**`MOONCAKE_TE_META_DATA_SERVER` carries an underscore inside `META_DATA`.** It is not
`MOONCAKE_TE_METADATA_SERVER`, and normalising it to the spelling that reads correctly **silently
degrades the metadata plane** rather than erroring.

**`MOONCAKE_DEVICE` is left unset on purpose, and the documented value `auto-discovery` is a trap.**
The client splits that key on commas into a device filter and nothing special-cases the string, so
setting it produces a filter matching a device no host has. **Empty means "use every device found".**

**`members[].extraArgs` is the exception to the table above: it renders into the container's
argv**, as `-D key=value`, not into an environment variable. `leader.extraArgs` does the same on the
leader, as `-key=value`.

**Either way the value is world-readable, on three paths.** It is stored verbatim on the
`KVCacheBackend` (which is cluster-scoped, so reading it needs no access to any workload), and it is
rendered into the container's argv, readable again from the Pod and from the DaemonSet or Deployment
carrying it. **Do not put a credential in `extraArgs`.** Nothing refuses one at admission.

`spec.transport.protocol` accepts `Auto`, `TCP`, `RDMA`, `EFA`, `CANN`, `ROCM`, `MUSA` and `MACA`,
and defaults to `Auto` whether or not the `transport` block is written at all. **`Auto` resolves to
`TCP`.** It is not a per-node probe that promotes itself. `MUSA` and `MACA` are intra-node IPC
transports, not host fabrics, so they take none of the fabric privileges below.

Which engine each value serves, and on which images, is [the transport
matrix](../model-deployment/engine-versions.md#which-transport-each-engine-can-use).

**`members[].transport.protocol` overrides that value for one group; left unset, the group inherits
the backend's.** The override exists for the one thing two media do not agree on: a VRAM group
reaching its peers over a fabric while the DRAM group beside it stays on `TCP`.

**`TCP` is complete, not a fallback, and it costs CPU.** Everything the store does works over it:
the cache fills, cross-instance prefix reuse works, and capacity scales with DRAM and the disk tier.
What it does not have is the zero-copy path the fabrics take. Every transfer traverses the kernel
network stack packet by packet, and the CPU that does so is CPU the node is not giving to anything
else.

So a `TCP` pool under load shows higher CPU on the member nodes and higher transfer latency than the
same pool on `RDMA`, and **that is the transport behaving as designed**. Read it as a reason to ask
for a fabric, not as a defect to open a report against.

The fabric privileges below render per group from the group's effective protocol, and an engine is
handed the protocol of the group it matched. An engine whose constraint no group in the pool
satisfies is refused at admission rather than started.

`CANN` is the one value an engine can **require**: vllm-ascend's store client currently raises on
any other protocol (verified at vLLM-Ascend v0.23.0 and v0.26.0rc1; upstream state, not a contract,
and it may change), so a pool serving Ascend engines declares `CANN` here, or on one member group,
rather than settling for the `TCP` default.

The member then needs a CANN-carrying image, per the variant table above. The project's own CPU
build compiles no `ascend` transport.

> **Why** — one group is one Pod template, which cannot express a per-node transport; and promoting to
> a host fabric would mean granting `hostNetwork` plus `IPC_LOCK` and `SYS_RESOURCE`. A privilege is
> requested, never inferred. Naming `RDMA` or `EFA` is also what accepts the security context that
> comes with it — which is those three things and **not** `privileged`. A `TCP` group sets none of
> them. `privileged` is reachable, but only by writing it into
> [`members[].securityContext`](#vram-group-accelerator-access), never by naming a protocol.

**Both host fabrics also grant the member one device, and the protocol names it.** Nothing is
declared: an `RDMA` group asks for `device.gpustack.ai/rdma.shared`, one of
[this operator's own RDMA keys](/gpustack-operator/main/docs/modules/rdma/network-topology/index.md#the-rdma-resource-keys);
an `EFA` group asks for
[the key AWS's EFA device plugin advertises](/gpustack-operator/main/docs/modules/rdma/network-topology/index.md#the-rdma-resource-keys).

The renderer **derives** the RDMA name rather than spelling it, so the page linked above is the one
to trust if the two ever disagree. The request is the permission (a bind mount of a device tree is
not), so the member that gets one can `open()` the adapter and the member that gets none could not
have.

**`members[].fabricInterfaceCount` sets how many interfaces each member asks for, on `RDMA` and
`EFA` only.** Left unset it counts as one, and an unset count renders the same member as a written
`1`. On `RDMA`, one asks for the shared key and more than one asks for that many exclusive
`device.gpustack.ai/rdma`. A written count is at least `1`.

On every other protocol the count has no effect: no device is requested. Admission **warns rather
than refuses** when a group writes one there, so the update that moves a group from `RDMA` back to
`TCP` still goes through with the count in it.

**On a cluster tracking the default branch, this changes what running RDMA members ask for.** No
release has ever carried this API, so there is no upgrade path to migrate; but a development cluster
whose `RDMA` backend predates the change will have its members re-rendered against
`device.gpustack.ai/rdma.shared` on the next reconcile. **Confirm the Device Manager is running on
those nodes first**, or the members roll and stay `Pending`.

**The cluster therefore needs the plugin that advertises the group's resource**, and a node without
it never runs that member: the Pod stays `Pending`. That refusal is the intended one. A cluster
whose nodes cannot serve the fabric is a cluster whose backend should say `TCP`, and an operator who
wants it back says so in `protocol` rather than by leaving a field empty.

**No member mounts `/dev/infiniband`.** Each fabric's plugin injects the verbs character device of
every device it grants, so mounting the tree beside that grant would add every adapter the member
was **not** granted, visible and unopenable.

**On `EFA` the mount breaks a partial grant.** A member granted fewer EFA devices than its node has
would see the rest through the tree, and `open()` on one returns `EPERM` from the device cgroup.
libfabric's EFA provider gives up its whole device list on the first `EPERM`, so the store reports
`No available EFA devices` and the container restarts in a loop. Without the mount it initializes.

Nothing is mounted from a host EFA install: the libfabric an `EFA` member runs on is in the image.

**The engines an `EFA` group serves need an EFA build of Mooncake as well**, which no runner image
carries. See [what an EFA leg needs from the engine
image](../rdma/operations.md#efa-engine-images).

EFA capability is a property of the instance size rather than of its family (the largest `i7ie`
sizes carry it while every smaller one does not), so check the size about to run, with
`fi_info -p efa` on the node or the instance type's own EFA field, never a family name.

**Cross-node EFA requires both nodes in one cluster placement group.** One availability zone is
necessary and not sufficient. Two nodes sharing a zone but in different placement groups complete
the libfabric handshake (the endpoint pair is reported established), and then every work request
hangs forever with the receiving adapter's byte counters at zero and no error on either side. The
silence is the problem: it reads as a hung probe, not as a placement fault.

Where nodes sit is outside this operator, so it belongs to whoever builds the cluster. A managed
node group typically gets a placement group of its own, which makes "members of one pool spread
across two groups" the default outcome rather than an unusual one. A pool whose members are meant
to reach each other should be pinned to node groups that share one.

**The medium and the transport are independent.** A `DRAM` group is as entitled to a fabric as a
`VRAM` one, and this rendering reads the medium nowhere: the protocol decides the network, the
capabilities and the device, while the medium decides only what the segment is made of and what it
is charged to.

**What a member's Pod requests follows the group's medium.** A DRAM member requests host memory for
`capacityPerMember + localBufferSize`. A VRAM member's segment is device memory, which **nothing
requests**: host memory carries `localBufferSize` only, and the device memory is claimed by
allocating it.

**The engine sharing that accelerator does not know the member took a slice of it.** Nothing in
Kubernetes accounts for device memory, so the two are held apart by **sizing `capacityPerMember`
against the engine's own memory fraction**, and by nothing else. Give the engine a
`gpu_memory_utilization` that leaves `capacityPerMember` free, and expect whichever starts second to
fail on allocation if you do not.

A member group asks for **no** accelerator extended resource, because the claim would not be true: in the tested Mooncake v0.3.13.post1,
a member allocates its whole segment on one device, so a resource request would take a whole
accelerator away from inference to account for a slice of one.

**One member's segment is therefore on one device**, whichever the container sees first. An
eight-accelerator node keeps all eight available to inference, and one member contributes a slice of
the first.

### VRAM group accelerator access

A VRAM group's container needs the vendor's **user-space driver**, which is what the store's client
loads to allocate a segment at all. That driver is the node's and is never in the image, so it has to
be brought in. Nothing is inferred; every route is written on the group:

| Field | Purpose |
|---|---|
| `members[].extraEnv` | The vendor runtime's own trigger, where one exists — `NVIDIA_VISIBLE_DEVICES: all` makes the NVIDIA toolkit inject the driver and every device. Ascend has no counterpart. |
| `members[].runtimeClassName` | The vendor container runtime named explicitly, for a cluster where it is not the default runtime. |
| `members[].hostPaths[]` | The driver tree taken from the node directly, for a cluster running no vendor runtime at all. |
| `members[].securityContext` | Privilege, when a mount alone is not enough to open what was mounted. |

**Privilege alone does not cover this.** It opens the node's device tree under `/dev`, which is
where the device nodes are and is **not** where the libraries are: Ascend's driver is under
`/usr/local/Ascend/driver` with DCMI beside it, and NVIDIA's `libcuda.so` is injected by the
container runtime. A privileged member with neither a runtime class nor the driver mounted starts,
reports healthy, and fails to allocate.

`securityContext` is **merged onto** what the group's protocol already earned, field by field, with
`capabilities.add` **unioned**. A `RDMA` or `EFA` group therefore keeps `IPC_LOCK` and `SYS_RESOURCE`
whatever it declares. Without them the transfer engine fails when it registers memory, long after
the container looked healthy. Declaring a capability adds it; a group that must hold neither of those
two declares a protocol that does not ask for them.

Each `hostPaths[]` entry is `{path, mountPath, type, readOnly}`. Name a `type`: left empty, the
kubelet checks nothing, so a missing path becomes an empty directory in the container and the member
starts anyway. A mount path is refused if it duplicates another entry's, if it is `/dev/infiniband`
(where a fabric's device plugin injects the granted device nodes, refused under every protocol since
that can change later), or if it is this group's own `localDisks[].path`.

Two worked groups. The NVIDIA one needs no mounts at all, because the container runtime injects the
driver once the variable tells it which devices to inject:

```yaml
members:
  - nodeSelector:
      kvcache-vram: "true"
    medium: VRAM
    image: gpustack/mirrored-mooncake:0.3.13.post1-cuda13.0
    capacityPerMember: 8Gi              # the engine's gpu_memory_utilization must leave this free
    extraEnv:
      - name: NVIDIA_VISIBLE_DEVICES
        value: all
```

That variable is honored only while the toolkit's
`accept-nvidia-visible-devices-envvar-when-unprivileged` is on, which is its default and which
NVIDIA's own hardening guidance turns **off**. On a cluster that turned it off, use
`runtimeClassName` instead.

The Ascend one has no such variable, and this cluster runs no vendor runtime, so the driver is
mounted from the node:

```yaml
members:
  - nodeSelector:
      kvcache-vram: "true"
    medium: VRAM
    image: gpustack/mirrored-mooncake:0.3.13.post1-cann9.1
    capacityPerMember: 32Gi
    hostPaths:
      - path: /usr/local/Ascend/driver
        mountPath: /usr/local/Ascend/driver
        type: Directory
        readOnly: true
      - path: /usr/local/dcmi
        mountPath: /usr/local/dcmi
        type: Directory
        readOnly: true
      - path: /usr/local/bin/npu-smi
        mountPath: /usr/local/bin/npu-smi
        type: File
        readOnly: true
```

**A VRAM group may also carry a [local disk tier](/gpustack-operator/main/docs/modules/kv-cache/local-disk-tier/index.md), and that combination is
measured on NVIDIA.** Objects written past a 256Mi device segment left memory for the tier and read
back with matching digests. No upstream test covers the pairing, so the reading is this project's own.

**The equivalent on Ascend (`cann`) and AMD (`rocm`) is still unmeasured.** Treat the pairing as
verified on NVIDIA and as unverified elsewhere, rather than as guaranteed everywhere.

**A writer against this tier must retry.** A put that fails on a full segment means eviction has
not caught up, not that the tier is full: the offload heartbeat runs on an interval, so a writer that
abandons on first failure reads "eviction in progress" as "cannot write".

**Reachability is a port range, never a list.** The transfer engine picks its data ports at random,
and the peer-to-peer plane is what binds them. Write firewall and NetworkPolicy rules between member
nodes, and from engine clients, as a **range**. The rendered Pod declares no fixed data-plane
`containerPort`, because a fixed list would be a false statement.

**The management port is fixed, and on a host fabric it lands on the node.** A member serves its HTTP
API on `8080 + <group index>` (`8080` for the first group, `8081` for a second). A `TCP` group
holds that port inside its own pod network namespace, but an `RDMA` or `EFA` group holds the host's,
so on every node such a group selects that port must be free. Reserve one port from `8080` upward per
member group.

> **Why it moves per group** — two host-network groups whose node selectors both match one node place
> two host-network Pods on it. On a single fixed port only the first binds; the second runs, never
> passes readiness, and reports nothing about why.

Members advertise their Pod IPs for client connections. Target those addresses in data-plane
network rules. With host networking, the Pod IP is the node address.

Every segment mount gets a new identity and transfer port. After a member restarts, clients
must discover its current endpoint rather than reuse a previous connection address.

## Status and conditions

```console
$ kubectl get kvcb
NAME            TYPE       PHASE   ENDPOINT                                        CAPACITY
mooncake-dram   Mooncake   Ready   mooncake-dram-leader.gpustack-system.svc:50051  12Gi
```

Five phases: `Provisioning`, `Ready`, `Degraded`, `Error`, `Deleting`. `Ready` carries no
`phaseMessage`; every other phase carries one.

Conditions report the axes: `LeaderAvailable`, `MembersMounted`, `CapacityObserved`, `PoolWrites`, `Deletable` and
`RolloutComplete`. Two more appear only where they have something to judge: `ElectionObserved`
when `leader.electionBackend` is `Kubernetes` (including at one replica), and
`TierWasEmpty` when a member group carries a [local disk tier](/gpustack-operator/main/docs/modules/kv-cache/local-disk-tier/index.md).

**Those last three do not move the phase, and that is deliberate.** A rollout in flight, an election
that has not happened, and a disk tier found holding data are all states in which the backend serves
normally. Reporting them as `Degraded` would put a storage arrangement in the same field as a
leader nobody can reach. Read the conditions for them; the phase will not tell you.

**A member that is starting is not a shortfall; a member that is stuck is one.** A Pod still pulling
its image is left alone. Holding it against the backend would report `Degraded` for the length of
every rollout. But one whose container will not start, or that no node will take, is never going to
arrive, so it reads `Degraded` even while the other members serve, and `phaseMessage` carries that
Pod's own reason.

**A member reads Ready only once its segment is mounted.** Its container carries a readiness probe
that connects to the entrypoint's REST port, and the entrypoint mounts the segment *before* it serves
that port. Readiness is therefore evidence of the mount, not of the process.

> **Why the probe is load-bearing** — without it the kubelet reports Ready as soon as the container
> runs. Every ready member Pod is held to the leader's listing, so that window would read as a
> shortfall and move a healthy backend to `Degraded` for the length of every rollout.

**A listing too large to publish is withheld, never truncated.** Past what `status.members` can carry,
the phase reads `Degraded` with reason `ListingTooLarge` and the previous listing is kept.

> **Why** — every entry is republished on each pass, so publishing past the object size the API server
> accepts would make every status write fail from then on while the read that produced it reported
> success. A truncated list would be worse: it reads exactly like a backend that lost members.

**Status is polled every 15 seconds, not only refreshed on events.** Everything above is read over
HTTP from the leader, and a store whose contents move while its Pods sit still produces no Kubernetes
event at all. An external backend produces none ever, since this operator owns no workload for it.
So `kubectl get kvcb -w` moves on its own.

> **15 seconds is an interval, not a maximum age.** The timer starts after a pass finishes, and a
> pass makes up to three sequential HTTP reads. More importantly, `status.members` is **deliberately
> retained** when a read of the segment listing fails (a stale list plus `MembersMounted=False` is
> more honest than an empty one), so it has no age bound at all while that read keeps failing. The
> condition is what says whether the list was refreshed; the list alone never does.

**Capacity is observed, not derived.** `status.capacity` is read from the leader's own counters,
never from what the spec declares. Nothing multiplies `capacityPerMember` by a replica count. A
backend with no disk tier reads `master_total_capacity_bytes`; one **with** a tier reads that plus
`master_total_file_capacity_bytes`. Adding two **observed** families is not the same thing as
adding up what members were asked to provide.

**`status.capacity.total` is capacity, not usage**, and a disk tier contributes the ceiling the
member declared (published as soon as the member registers, before anything is written there).

`PoolWrites` reports pool-wide write activity. `True` means the leader has seen a put end since
its current process started, or a member reports positive `allocator_used_bytes`. `False` means
the leader has seen both a put start and a put revoke while every member reports zero allocated
bytes. `Unknown` covers no put start, an unfinished put, or a missing reading.

The process lifetime is the observation window because these counters reset when the leader
restarts; a single status poll could miss a short write.

This condition does not identify which deployment attempted a write, and a previous failure in that
window does not prove that every current write fails.

To ask whether the **disk** tier is holding data, the figure to read is not on the CR at all. See
[Bucket writes](/gpustack-operator/main/docs/modules/kv-cache/local-disk-tier/index.md#bucket-writes).

**Capacity is absent (not zero) while the leader is starting.** `/metrics` is ungated: a leader
that is up but not serving answers 200 with a well-formed exposition whose gauges all read zero, and a
zero is indistinguishable at the parser from a genuinely empty cache. Publishing is therefore gated on
`service_ready`, not on the scrape succeeding.

`status.members[]` is read from the leader's segment listing, one entry per **listed** segment. Each
row carries the leader's `segmentID`, `clientID` and advertised `segmentName`; the list is keyed by the
unique segment ID because several members may legitimately share a name.

The leader is what allocation goes through, so a running member Pod it does not list holds nothing
and is counted in `MembersMounted`'s message instead. The two fields the listing cannot supply
(node name and medium) are joined in from the member Pod behind that segment, and left **empty**
rather than guessed when nothing matches.

**Two ready host-network member Pods that share an address make Pod attribution ambiguous.** Their
rows remain publishable because their segment and client IDs are distinct, but neither Pod exposes
the ID that maps a row back to it. `MembersMounted` goes `False` with reason
`AmbiguousMemberIdentity`; each row still carries the node and the medium its candidates **agree**
on, neither being a fact about one Pod, and leaves empty whichever of the two they dispute.

Two groups on one node do **not** collide by themselves. A `TCP` member advertises its own pod IP, so
each segment carries a distinct name even though both Pods answer to the node's name; the collision is
on host-network paths (`RDMA` and `EFA`), where both Pods hold the host's network namespace and
advertise the node's address.

The remedy is to give the groups node selectors that keep them on different nodes.

A failed listing scrape **keeps** the previous list and sets `MembersMounted=False`; a failed capacity
scrape **clears** the figures. That asymmetry is deliberate: capacity is two pointers and has an
"absent" that means *not observed*, while an empty list is a legible value meaning *no segments*, so
clearing it would publish a falsehood.

Mooncake masters before 0.3.12 do not report segment membership. They can serve requests and
report `Ready` while `MembersMounted` is `Unknown/SegmentListingNotServed` and `status.members`
is empty. Waiting does not add a capability missing from that version.

Other listing failures, including server errors or an unreachable master, report `ListingFailed`
and retain the previous member list.

`PoolWrites` on such a leader is still `True` once it has seen a put end, since that counter is on
`/metrics`; before that it is `Unknown` with the same reason, because its other verdicts read the
members' allocations off the listing.

Two states still read `MembersMounted=False` and `Degraded` there: a member Pod that will not start,
read from the Pod's own status, and a `status.capacity.total` of zero, reported as `NoSegments`
(the gauge sums the mounted segments, so zero is every member unmounted).

## Growing and shrinking a group

**Widening `members[].nodeSelector` adds members without restarting the ones already running.** The
DaemonSet places a Pod on each newly matching node; every existing Pod keeps its UID and its restart
count, leader included. The DaemonSet is left on `OnDelete`, and restarts are decided from a
comparison over the whole Pod template except the node selector: a widening matches that comparison,
while an image, argv, environment, resource or fabric change fails it and every member is recreated.

**A Pod runs the template it was created from, so a setting cannot protect the same edit that
removes it.** Anything rendered into the member Pod (the shutdown hook, its grace, the environment)
reaches a member only when that member is recreated. A departing member leaves with what it started
with.

To make such a setting apply to a shrink, do it in **two steps**: change only the setting and wait
for members to be recreated with it (their pod-spec-hash annotation moves), then narrow the selector
or remove the group. `scaleIn.gracePeriodSeconds` below is the case this bites today; the property
belongs to the Pod template, not to that field.

**Existing objects are not rebalanced.** How fast the cluster converges onto a new member depends on
`allocationStrategy`: `FreeRatioFirst` (the default) biases new writes toward the emptier member,
`Random` does not.

### Group identity by position

A group has **no name**. Its position in `members` is what the DaemonSet's name, its immutable
selector labels and its members' HTTP port are all derived from, so moving an entry in that list
leaves every one of those in place and changes only the spec underneath it.

**Moving a group to another position is refused at apply time.** The refusal names both positions.
Two shapes reach it:

| Edit | Outcome |
|---|---|
| swapping two entries | refused |
| removing a group ahead of others, which shifts the rest up | refused |
| appending a group | allowed |
| removing from the **end** of the list | allowed |
| editing a group in place, including widening its `nodeSelector` | allowed |

**To take a group out of service without removing it, narrow its `nodeSelector` until it matches no
node.** The group keeps its position, every later group keeps its DaemonSet, and nothing is rebuilt.
Positions are used instead of names because a DaemonSet's selector cannot change after creation: a
name free enough to make reordering possible would force every member DaemonSet to be deleted and
recreated, and the entire cache would go with them.

**The rule recognises a group that arrived unchanged at a position another group left; it does not
recognise every reorder.** Without a name there is nothing else to recognise a group by, so two shapes
are knowingly admitted:

- a reorder **combined with an edit** to the same group, which is indistinguishable from two ordinary
  edits;
- removing a group when a **later group is identical** to the one taking its place, `[A, B, C]`
  becoming `[A, C, C]`. That produces the same two lists as editing position 1 to match an unchanged
  position 2, which is how the second of two look-alike groups is taken out of service, so refusing
  it would forbid the operation recommended above. `[A, B, C]` to `[A, C]`, with no look-alike to
  arrive in the gap, is still refused.

Both admitted shapes leave one trace: the resulting `members` holds **two identical groups**. What
the rule buys is that the mechanical reorder (the one a rewritten manifest produces) is reported
instead of silently rebuilding members against another group's spec.

Shrinking a group discards the cache held by the removed members. Narrowing the selector or
removing a node immediately unmounts its segment; the operator does not migrate its cached data.
The termination grace period gives the process time to finish shutdown and does not preserve
that data.

**`scaleIn.gracePeriodSeconds` holds the process, not the tier.** A member with a
[local disk tier](/gpustack-operator/main/docs/modules/kv-cache/local-disk-tier/index.md) gets a `preStop` hook that deregisters the tier with the
leader and then waits out the grace. A group with no tier renders no hook and the setting is inert.

With the example image, deregistration takes effect when the hook runs. A key held only by that
tier becomes a clean miss immediately, and the process waits for the full configured period even
if no reads remain. Size the period for the departing member's local work.

This timing is specific to the example image. `spec.image` selects the backend implementation;
other images may deregister later or wait differently. The operator controls the value sent to
the endpoint.

```yaml
spec:
  connection:
    managed:
      scaleIn:
        gracePeriodSeconds: 30       # 0..3600
```

**The Pod's termination window is derived from it**, as `gracePeriodSeconds + 60`, rather than being
a second field beside it. That is what makes the relationship hold: two independent fields could be
set so the kubelet kills the container in the middle of the wait, and no validation makes that
impossible. It only makes it checkable.

The upper bound of 3600 is the member endpoint's own; above it the call is refused with a `400`, so a
larger value would render a hook that fails every time it runs. None of it makes a shrink lossless:
the memory segment is still dropped, and the disk tier stops answering as soon as the hook runs.

Migrating a member's data before it leaves (the store's drain job API) is **not** offered here. It
is stateful orchestration, and it reaches only the memory and NVMe-oF replicas: it cannot name the
segments of the disk-backed ones and skips those keys without counting them as blocked, so **a drain
over a backend with a disk tier reports success while leaving that tier's data where it was**. Its
own success signal does not tell you that happened.

## The external mode

`connection.external` points at a backend somebody else runs. The operator creates **nothing**
(no Deployment, no Service, no DaemonSet) and only observes.

```yaml
  connection:
    external:
      endpoints:
        - name: Client
          address: mooncake.example:50051
        - name: Admin
          address: mooncake.example:9003
```

Both roles are required, so the list holds exactly two entries: entries are keyed by `name`, which
has two values, and admission refuses a list missing either one. Each address is validated as
`host:port` at admission. A blank or portless one would otherwise be mirrored into status and
handed to an engine that cannot dial it.

**`Admin` is what this operator reads** (health, metrics and the segment listing), and **`Client` is
what an inference engine connects to**. The addresses are mirrored into `status.endpoints` unchanged.

Two behaviours differ from the managed mode:

- **An address that does not answer is reported as an `Error`, not as a backend still starting.** A
  managed leader is excused while its own Deployment has no ready replica; an external address was
  declared to name something that already runs, so there is no Pod to wait on and a mistyped endpoint
  would otherwise sit at `Provisioning` forever.
- `status.members[]` carries **no** node name or medium, because those Pods are not this operator's to
  look up. Capacity is the **sum of both pools**, since an external object names no medium to pick one
  by.
- **A redirect from the `Admin` address is never followed.** That address belongs to whoever wrote
  the spec, and honouring a `3xx` from it would read some other host with the operator's network
  identity, then copy an excerpt of the answer into a status readable by anyone who can read the
  object. The redirect is reported as the response it is.

### Two objects on one leader

**Two `KVCacheBackend` objects may name the same leader, and nothing in the operator notices.** For
a managed backend the object *is* the leader, so two objects are two leaders. For an external one the
object is a **declaration of addresses**, and one leader is reachable under more than one spelling:
by Service name in one object and by IP in another, with or without a trailing dot or an explicit
default port.

> **Why no check** — every identity the operator could compare is either editable or needs the leader
> reachable at admission. The comparison cheap enough to run (byte-identical addresses) catches the
> copy-paste case and misses the one a real deployment produces, which is the same leader spelled two
> ways. A check that catches the easy half invites the reader to trust it for the other half. This is
> tracked in [issue #288](https://github.com/gpustack/gpustack-operator/issues/288).

**Three consequences follow, and the third is the one to read before deciding the duplication is
safe.** The reuse-domain uniqueness rule is enforced between Bindings whose pools name the **same
backend object**, so two Bindings reaching one leader through two objects are both admitted on one
`domain.name`:

1. **The quota of a shared domain flips and never settles.** The leader keeps one ledger entry per
   tenant, and each pool's reconciler converges that entry toward its own Binding's `quota.ceiling`
   **on every pass**. Each pass reads the other's figure, finds it wrong, and writes its own back.
2. **The symptom of an undersized quota is a low hit rate and nothing else.** Exceeding a tenant's
   quota does not refuse the write: the store frees room by dropping that tenant's own older objects
   and retries, irreversibly and **without any counter moving**. So the flipping above never surfaces
   as an error. It surfaces as a cache that keeps losing content nobody asked it to lose.
3. **Two Bindings on one `domain.name` with a different `blockSize` or `dtype` corrupt each
   other's blocks.** The reuse identity an engine is handed is the domain **name alone**. Each
   Binding hands its own [`dtype`](/gpustack-operator/main/docs/modules/kv-cache/pool/index.md#engine-dtype) to its own
   engines, and `blockSize` reaches no engine at all. So two differently-shaped caches land under
   one identity, which is [the silent cache pollution](/gpustack-operator/main/docs/modules/model-deployment/deployment/index.md#inherited-reuse-domain)
   a wrong `blockSize` or `dtype` causes, reached here without either value being wrong.

If you point two objects at one leader, either keep their pools' Bindings on **different**
`domain.name` values, or make sure every Binding that shares a name also shares its `blockSize`,
`dtype` and `quota.ceiling`.

## Operating notes

**`replica_num` is engine-side.** How many replicas of a stored object Mooncake keeps is a per-`Put`
argument the caller supplies through its connector configuration. It is not controllable from this CR.

**One benign startup line, documented so nobody files it as a bug.** It appears on every client start
and is harmless:

```
E transfer_metadata.cpp:991] Local segment descriptor not found
```

**Client-side environment knobs**, observed at startup:

```
MC_TE_METRIC=1                        enable transfer-engine metrics (OFF by default)
MC_STORE_CLIENT_METRIC_BANDWIDTH      client bandwidth summary
MC_STORE_MEMCPY                       unset => auto-detected ("TCP-only environment, memcpy enabled")
MC_METADATA_SERVER / P2PHANDSHAKE     the transfer engine's own low-level metadata knob
```

**`MC_TE_METRIC` is rendered on every member group and cannot be turned off from the CR.** The other
three are set on the *workload* and not here. It is reserved in `members[].extraEnv` for the reason
every rendered name is: the hatch appends, and a container carrying one name twice leaves the winner
to the runtime.

The switch reports throughput and a task-latency distribution to the container log, on an interval
that prints nothing when nothing moved. The leader's Prometheus surface counts keys and bytes and
measures no data plane, so without it the receiving end of every transfer is unmeasured while the
sending end already reports.

`MC_METADATA_SERVER` is **not** the variable this operator renders. The member is configured
through the store client's key, `MOONCAKE_TE_META_DATA_SERVER`. Two names for the metadata plane is
exactly the near-miss that gets one of them typed into a template.

**A backend in use cannot be deleted.** While `status.usedBy` names a consumer the finalizer holds,
the object stays at `phase: Deleting` with the claimant named in its message, and the workloads keep
running. Clearing the last claim lets the teardown complete.

**A backend in use also cannot have `leader.multiTenancy` turned off.** The webhook refuses the edit
while `status.usedBy` names a consumer, and the refusal names them. Remove the consumers first. See
[KV Cache Pool](/gpustack-operator/main/docs/modules/kv-cache/pool/index.md#operating-notes) for what the withdrawal costs on their side.

> **Why** — the flag decides whether the master keeps a per-tenant ledger, and a consumer's quota is
> both written and released through that ledger. Withdrawing it under a live consumer therefore takes
> away the way that consumer is unwound, and the cost only appears when something is deleted. The
> refusal reads `status.usedBy` as it is written, not as it resolves: an entry naming an object that
> no longer exists refuses the edit too, and that list is where it is cleared.

**The teardown deletes the workloads first, and the object disappears last.** The leader Deployment,
its Service and every member DaemonSet go before the finalizer comes off. It waits for them to be
**gone**, not merely for the deletes to be accepted, so the object going away means the backend is
gone rather than scheduled to be.

> **Why not leave it to ownership** — they are owned dependents, so the collector would reach them
> either way. But between the finalizer coming off and it running, the leader is still serving on an
> address nothing accounts for. Ownership is the safety net, not the mechanism.

**That wait is bounded, one workload at a time.** Each gets the termination grace its own Pod template
declares plus a couple of minutes, timed from its own deletion timestamp. Past that it stops being
waited for, and a warning Event on the workload names the nodes its Pods are still terminating on.
Deleting a backend therefore finishes even when a node it ran on has stopped answering.

> **Why its own grace and not one number** — a member group with a disk tier derives its Pod's grace
> from `scaleIn.gracePeriodSeconds`, which reaches an hour, so any constant short enough to bound an
> unreachable node would abandon a group draining exactly as configured. That grace works as the clock
> because the kubelet treats it as a hard kill deadline: a Pod outliving it is not a slow one, it is
> one whose kubelet is not acting.

**Only objects carrying this backend's own note are deleted.** The names are derived, so an unrelated
object can hold one, and a delete has to be surer than a name. The member sweep finds its DaemonSets
by the same note it then checks, rather than by the identity labels. Discovering on one key and
judging on another is how an object goes missing from its own teardown.

---

**See also** — [KV Cache Leader](/gpustack-operator/main/docs/modules/kv-cache/leader/index.md) (the metadata process, its probes and its
Lease election) · [KV Cache Local Disk Tier](/gpustack-operator/main/docs/modules/kv-cache/local-disk-tier/index.md) (the optional disk layer on a member
group, and the bucket that is its write unit) · [KV Cache Pool](/gpustack-operator/main/docs/modules/kv-cache/pool/index.md) (how a namespace is granted
a quota on this store, and what a quota ceiling buys) · [Admission](/gpustack-operator/main/docs/modules/devices/admission/index.md) (the gates and the four-view status pattern) ·
[Settings & Environment Variables](/gpustack-operator/main/docs/reference/settings/index.md) (the `kv-cache-backend-image` Setting) ·
[Installation Modes](/gpustack-operator/main/docs/operate/installation-modes/index.md) (why the CRD is applied by the worker, not the chart)

**Next** → [Internals](/gpustack-operator/main/docs/contribute/internals/index.md) — startup ordering and the invariants that fail silently.
