KV Cache Backend
A KVCacheBackend configures shared KV cache for inference workloads. With a managed backend,
the operator runs a leader for metadata and a member group with one store process per selected
node. An external backend connects to an existing service.
Mooncake calls the leader a master. Its flags, environment variables and metrics use that name.
Contents
- Connection and medium
- The image
- The metadata plane
- The members
- Status and conditions
- Growing and shrinking a group
- The external mode
- Operating notes
Connection and medium
connection chooses whether the operator manages the backend or connects to an external service.
For managed backends, members[].medium chooses host memory (DRAM) or accelerator memory (VRAM).
apiVersion: worker.gpustack.ai/v1
kind: KVCacheBackend # cluster-scoped, short name kvcb
metadata:
name: mooncake-dram
spec:
type: Mooncake
image: docker.io/kvcacheai/mooncake:0.3.13
connection:
managed: # or external: — exactly one
leader: {} # replicas, allocationStrategy and multiTenancy default
members:
- nodeSelector:
kubernetes.io/os: linux
medium: DRAM # what this group's SEGMENT is made of: DRAM or VRAM
capacityPerMember: 4GiThe example uses Mooncake 0.3.13, which is compatible with the documented vLLM client. SGLang
requires a different version; check Engine Versions
before
choosing spec.image.
connection.managed and connection.external are both optional pointers and exactly one must be
set; neither and both are refused at admission with a message naming the two. Several member groups
are allowed; at most one of them may carry a local disk tier
.
members[].medium takes two values, DRAM and VRAM: host memory or device memory. The
renderer treats the two differently. On a DRAM group the segment is host memory, so the operator
adds capacityPerMember to the member Pod’s host memory request. On a VRAM group the segment is
device memory, so the Pod request carries nothing for it: the member claims its slice by allocating
it.
One binary is one medium, so a node contributing both does so as two groups selecting it. See The members for the renderings each value produces.
The field is immutable: a segment already mounted cannot change the kind of memory underneath the data it holds, so an edit is refused and the choice is made when the group is declared.
An earlier shape offered five values. Four of them named things that are not member groups, and each is reached another way:
Former medium value |
What it is | Location |
|---|---|---|
LocalDisk |
a tier on the members that already hold the memory replica | members[].localDisks
|
NoF |
an NVMe-oF target coordinate, registered once, with no node affinity and no Pod | no API surface; it is not a member group |
CXL |
a DAX device the leader process allocates from | nowhere in this API: enable_cxl, cxl_path and cxl_size are refused in leader.extraArgs, because the first replaces leader.allocationStrategy and then brings the leader up advertising the allocator’s size as capacity even where no DAX device exists |
DFS |
a distributed filesystem the leader process allocates from | the leader’s own environment, which this API does not render |
Why the shape matters more than the names — the leader routes an offload task to the client holding the key’s memory replica. A member group with no memory segment is therefore never chosen, so a group declared as “the disk one” would report its disk capacity to the leader and never receive a single write. The object would say one thing, the running member another, and nothing would report a fault.
What an object written against that earlier shape can still do depends on the API server’s
CRDValidationRatcheting gate, which follows the cluster’s effective Kubernetes version:
- v1.33 and later — the gate is locked on. An update whose invalid field is unchanged is admitted, so removing a finalizer works and the object deletes normally.
- v1.30 through v1.32 — the gate is on by default; a cluster that switches it off behaves like an older one.
- v1.28 and v1.29 — the gate exists but is off by default; turned on, it admits the unchanged-field update.
- Earlier — the gate does not exist. Every write touching the object is refused, finalizer removal included, so the object cannot be deleted either. Remove such an object before applying a CRD that narrows the enum, not after.
No released version of this API carried those values, so the question does not arise on a cluster that installed a release; this is a development-cluster shape.
The object is cluster-scoped: it names nodes, claims host memory and host paths, and on the RDMA
and EFA paths needs hostNetwork and a fabric device. Only a cluster administrator can legitimately
declare one.
A pool sets the quota ceiling on a backend; a Binding grants a namespace a share of that quota.
This mirrors Kueue’s ClusterQueue and LocalQueue split
(https://kueue.sigs.k8s.io/docs/concepts/
). Several pools can reference one backend. The grant
does not enforce access to the store; see Limitations
.
The image
spec.image is explicit and never derived from the operator’s own image, which breaks
deliberately with how the Device Manager image is derived from the worker image. Leave it unset and
the cluster-wide kv-cache-backend-image Setting supplies it. That Setting ships a default, so a
backend naming no image runs this project’s own build; unset in both places (which takes an
administrator clearing the Setting) is refused at admission, naming both.
Clearing that Setting later does not strand a backend admitted under it. Admission re-asks for a
fallback only when an update moves spec.image itself. Every other update is admitted whatever the
Setting says now. That includes the reconciler’s own removal of the finalizer, which would otherwise
leave an object that owns nothing and cannot be deleted.
Why — the master’s link-time dependencies differ per published vendor variant, so no single derivation is correct. On a CUDA-less host, for example, the base (CUDA 12) master build needs
libcuda.so.1andlibcudart.so.12and cannot load, while the-rocmand-npubuilds need no accelerator library at all and load cleanly.
The client side is where the vendor lives, and it lives in the published engine wheel rather than in
this image: which transports a client can drive is a property of the wheel the engine image carries,
not of spec.image.
The master image needs no accelerator runtime; a member image needs the runtime of the transport it
uses. An -npu-built master on an all-NVIDIA cluster is legitimate, because the master is a pure
metadata service. A member on ascend, by contrast, needs CANN (libascendcl.so) in its container,
and a CANN-less image fails as a loader error whose own message reaches status.phaseMessage.
A member on EFA needs libfabric in its image, and one that can drive the node’s adapter, not the
distro build. mirrored-mooncake installs AWS’s, ahead of that copy in its loader cache.
Nothing has to be built to run this without high availability.
The example image above is published for amd64 and arm64 and carries both mooncake_master and
mc_store_rest_server, so one spec.image serves the leader and the members.
It runs on a host with no GPU: its libcuda.so.1 is a stub and its libcudart.so.12 is the real
library, and neither reaches a driver.
High availability needs a different image, and for both roles. That section carries which build, why no published one will do, and what each role does when handed one that cannot.
Why the stub/real split matters — the stub alone is enough for the master, but the Python client needs a versioned
cudaFreeHostfrom a real runtime. An image carrying two stubs runs the master and fails every member.
That split is what the kv-cache-backend-image default cannot cover. Its value is this project’s
own CPU build, carrying TCP, RDMA and EFA over DRAM, and one value cannot be right for every backend
at once: which build a member needs depends on the transport its backend asks for and the hardware
its group selects. A backend on a vendor fabric names its spec.image, which always wins over the
Setting.
The failure mode moved with the default, and that is what having one costs. Blank made a mismatch an admission refusal naming both places to fix; a default makes a wrong image a loader error at runtime, which is quieter and further from whoever can fix it. Clearing the Setting restores the refusal.
The project’s own build variants
Beside the -cpu default, this project publishes mirrored-mooncake in one build target per vendor
(cuda, cann, rocm), whose base images are dispatch-time build arguments. Tags carry the
toolchain version: <mooncake-version>-<variant><toolchain>, such as 0.3.13.post1-cuda13.0.
Each variant is built on 0.3.13.post1, the line vLLM’s supported clients are on, and on
0.3.10.post2, which serves only engines below the supported
minimum
. SGLang’s clients are on the 0.3.12 line, which this
project does not build. Which line a backend needs is that table’s question.
A 0.3.10.post2 variant needs leader.electionBackend: None and
leader.multiTenancy: false written out. It cannot run the default Kubernetes election.
The tenant ledger hangs on the master’s -enable_multi_tenants switch, which Mooncake took in 0.3.12,
and the field defaults on. On an older image the default renders a flag the master does not recognize,
and it exits at startup. The explicit false renders no flag, which is the command line such an image
has always run.
A VRAM group needs a build with VRAM segments compiled in (USE_VRAM_SEGMENT=ON), and the stock
-cpu default is not one. VRAM segments exist only on the 0.3.13 line (the 0.3.10.post2
variants carry the vendor transfer engine without them), so a VRAM group always names a
0.3.13.post1 variant tag.
Nothing refuses a VRAM group that names no image: the default is an administrator-editable Setting,
so this paragraph is the guidance rather than a gate. A group names its image through
members[].image, which wins over the backend’s spec.image, which wins over the Setting.
The other direction fails at startup, on purpose. A build with VRAM segments compiled in takes
device memory whenever it can reach a device, whatever medium the group declared. So a DRAM group is
rendered with every vendor’s visibility variable set to void, which keeps the accelerators out of
its containers: such an image then finds no device and refuses to start, against the group that
asked for a medium it does not implement.
Without that, the failure was silent and landed elsewhere. The group consumed device memory no Pod had requested, so neither the scheduler nor the quota chain knew it was gone, and what an operator saw was an unrelated deployment crash-looping on a figure nobody had configured. Pick the image that matches the medium; whether an image implements one is a fact about the image.
A private registry needs spec.imagePullSecrets, and an explicit policy needs
spec.imagePullPolicy. Both are backend-wide: they apply to the leader and to every member group,
including a group that names its own image. Left unset, the policy is resolved from the image
tag by the same rule the API server would have applied (Always for :latest or no tag,
IfNotPresent otherwise), and it is re-resolved whenever the image or the field moves.
Why they are fields and not Settings — the cluster-wide
image-pull-policyandimage-pull-secretsSettings are values of the bundled-application chart install. They reach the subcharts and nothing a controller renders, so aKVCacheBackendthat inherited them would be the only object in this API whose running workloads move when a chart value moves. The service accounts high availability renders carry no registry credentials either (they grant Lease access and nothing else), so without these fields no image here could come from a private registry at all.
The store version must match the engine’s client
A store and an engine-embedded Mooncake client interoperate only within one minor line. The
criterion is the RPC wire signature, not the version string: the handshake answers 2.0.0 for every
0.3.x release, so a mismatched pair is not refused at startup. Every probe reads green, and every
write then fails at transfer time with RPC_FAIL (-900).
Two posts of one minor line share their RPC signatures and interoperate. The 0.3.12 and 0.3.13
lines do not: the method names are unchanged, so the client reaches the handler and mis-decodes the
arguments. Multi-tenancy moves neither: with it off (a declared multiTenancy: false, the field
defaulting on) every request resolves to the default tenant.
The client’s version is a property of the engine image, not of anything on this CR. Which client each supported engine’s runner image carries, and so which line its store runs, is in the Engine Versions ; an engine below its minimum there is not supported.
For any other image, read the client off the image in hand rather than off an engine version. A
CUDA build resolves mooncake-transfer-engine against a lower bound at build time, so two images of
one vLLM version built on different days can carry different clients.
Direct P/D transfer (MooncakeConnector, no spec.kvCache) is engine to engine and exempt from this
matching.
Upstream has no 0.3.12.post2; the 0.3.12 line ends at 0.3.12.post1. High availability carries its own per-version rule, a master from 0.3.12 on. See High availability .
The metadata plane
The metadata plane is peer-to-peer and has no API field. The member’s metadata_server renders as
the literal P2PHANDSHAKE, unconditionally. It needs no etcd or Redis. The default leader election
does use a Kubernetes Lease and its API access, even with one leader replica.
Two axes get confused here, so both are stated. The metadata plane is how clients find one another.
The HA backend store (-enable_ha with -ha_backend_type) is how the leader elects, even at
one replica by default. The Kubernetes Lease is selected by
leader.electionBackend
, and it moves nothing on the metadata plane.
A manifest that tries to configure the metadata plane is not refused with a helpful message. There is no field, so there is nothing for a webhook to see:
- a strict client (
kubectl apply’s default) is refused by the schema withstrict decoding error: unknown field "spec.metadata"; - a client with validation turned off has the block silently pruned, and the object is admitted and reconciled as though nothing had been written.
The second is indistinguishable from success at the point of apply. The protection against it is the rule the page states: the metadata plane takes no configuration at all.
The members
One member group renders one DaemonSet over members[].nodeSelector. A member is one node’s
contribution to the store: it holds that node’s host memory, or on a VRAM group its device memory,
plus the node’s host paths and, on a host-fabric group, the number of fabric devices the group
requests.
The member’s identity is the node, and a DaemonSet is the workload that gives one Pod to every matching node. Two groups may select the same node (a DRAM group and a VRAM group is the shape the second medium exists for), and each still renders its own DaemonSet.
The member’s whole configuration renders as environment variables: no ConfigMap, no volume, no init container.
| config key | environment variable |
|---|---|
local_hostname |
MOONCAKE_LOCAL_HOSTNAME (the pod IP, from the downward API) |
metadata_server |
MOONCAKE_TE_META_DATA_SERVER |
master_server_address |
MOONCAKE_MASTER |
protocol |
MOONCAKE_PROTOCOL |
global_segment_size |
MOONCAKE_GLOBAL_SEGMENT_SIZE |
local_buffer_size |
MOONCAKE_LOCAL_BUFFER_SIZE |
device_name |
MOONCAKE_DEVICE — deliberately left unset, see below |
MOONCAKE_TE_META_DATA_SERVER carries an underscore inside META_DATA. It is not
MOONCAKE_TE_METADATA_SERVER, and normalising it to the spelling that reads correctly silently
degrades the metadata plane rather than erroring.
MOONCAKE_DEVICE is left unset on purpose, and the documented value auto-discovery is a trap.
The client splits that key on commas into a device filter and nothing special-cases the string, so
setting it produces a filter matching a device no host has. Empty means “use every device found”.
members[].extraArgs is the exception to the table above: it renders into the container’s
argv, as -D key=value, not into an environment variable. leader.extraArgs does the same on the
leader, as -key=value.
Either way the value is world-readable, on three paths. It is stored verbatim on the
KVCacheBackend (which is cluster-scoped, so reading it needs no access to any workload), and it is
rendered into the container’s argv, readable again from the Pod and from the DaemonSet or Deployment
carrying it. Do not put a credential in extraArgs. Nothing refuses one at admission.
spec.transport.protocol accepts Auto, TCP, RDMA, EFA, CANN, ROCM, MUSA and MACA,
and defaults to Auto whether or not the transport block is written at all. Auto resolves to
TCP. It is not a per-node probe that promotes itself. MUSA and MACA are intra-node IPC
transports, not host fabrics, so they take none of the fabric privileges below.
Which engine each value serves, and on which images, is the transport matrix .
members[].transport.protocol overrides that value for one group; left unset, the group inherits
the backend’s. The override exists for the one thing two media do not agree on: a VRAM group
reaching its peers over a fabric while the DRAM group beside it stays on TCP.
TCP is complete, not a fallback, and it costs CPU. Everything the store does works over it:
the cache fills, cross-instance prefix reuse works, and capacity scales with DRAM and the disk tier.
What it does not have is the zero-copy path the fabrics take. Every transfer traverses the kernel
network stack packet by packet, and the CPU that does so is CPU the node is not giving to anything
else.
So a TCP pool under load shows higher CPU on the member nodes and higher transfer latency than the
same pool on RDMA, and that is the transport behaving as designed. Read it as a reason to ask
for a fabric, not as a defect to open a report against.
The fabric privileges below render per group from the group’s effective protocol, and an engine is handed the protocol of the group it matched. An engine whose constraint no group in the pool satisfies is refused at admission rather than started.
CANN is the one value an engine can require: vllm-ascend’s store client currently raises on
any other protocol (verified at vLLM-Ascend v0.23.0 and v0.26.0rc1; upstream state, not a contract,
and it may change), so a pool serving Ascend engines declares CANN here, or on one member group,
rather than settling for the TCP default.
The member then needs a CANN-carrying image, per the variant table above. The project’s own CPU
build compiles no ascend transport.
Why — one group is one Pod template, which cannot express a per-node transport; and promoting to a host fabric would mean granting
hostNetworkplusIPC_LOCKandSYS_RESOURCE. A privilege is requested, never inferred. NamingRDMAorEFAis also what accepts the security context that comes with it — which is those three things and notprivileged. ATCPgroup sets none of them.privilegedis reachable, but only by writing it intomembers[].securityContext, never by naming a protocol.
Both host fabrics also grant the member one device, and the protocol names it. Nothing is
declared: an RDMA group asks for device.gpustack.ai/rdma.shared, one of
this operator’s own RDMA keys
;
an EFA group asks for
the key AWS’s EFA device plugin advertises
.
The renderer derives the RDMA name rather than spelling it, so the page linked above is the one
to trust if the two ever disagree. The request is the permission (a bind mount of a device tree is
not), so the member that gets one can open() the adapter and the member that gets none could not
have.
members[].fabricInterfaceCount sets how many interfaces each member asks for, on RDMA and
EFA only. Left unset it counts as one, and an unset count renders the same member as a written
1. On RDMA, one asks for the shared key and more than one asks for that many exclusive
device.gpustack.ai/rdma. A written count is at least 1.
On every other protocol the count has no effect: no device is requested. Admission warns rather
than refuses when a group writes one there, so the update that moves a group from RDMA back to
TCP still goes through with the count in it.
On a cluster tracking the default branch, this changes what running RDMA members ask for. No
release has ever carried this API, so there is no upgrade path to migrate; but a development cluster
whose RDMA backend predates the change will have its members re-rendered against
device.gpustack.ai/rdma.shared on the next reconcile. Confirm the Device Manager is running on
those nodes first, or the members roll and stay Pending.
The cluster therefore needs the plugin that advertises the group’s resource, and a node without
it never runs that member: the Pod stays Pending. That refusal is the intended one. A cluster
whose nodes cannot serve the fabric is a cluster whose backend should say TCP, and an operator who
wants it back says so in protocol rather than by leaving a field empty.
No member mounts /dev/infiniband. Each fabric’s plugin injects the verbs character device of
every device it grants, so mounting the tree beside that grant would add every adapter the member
was not granted, visible and unopenable.
On EFA the mount breaks a partial grant. A member granted fewer EFA devices than its node has
would see the rest through the tree, and open() on one returns EPERM from the device cgroup.
libfabric’s EFA provider gives up its whole device list on the first EPERM, so the store reports
No available EFA devices and the container restarts in a loop. Without the mount it initializes.
Nothing is mounted from a host EFA install: the libfabric an EFA member runs on is in the image.
The engines an EFA group serves need an EFA build of Mooncake as well, which no runner image
carries. See what an EFA leg needs from the engine
image
.
EFA capability is a property of the instance size rather than of its family (the largest i7ie
sizes carry it while every smaller one does not), so check the size about to run, with
fi_info -p efa on the node or the instance type’s own EFA field, never a family name.
Cross-node EFA requires both nodes in one cluster placement group. One availability zone is necessary and not sufficient. Two nodes sharing a zone but in different placement groups complete the libfabric handshake (the endpoint pair is reported established), and then every work request hangs forever with the receiving adapter’s byte counters at zero and no error on either side. The silence is the problem: it reads as a hung probe, not as a placement fault.
Where nodes sit is outside this operator, so it belongs to whoever builds the cluster. A managed node group typically gets a placement group of its own, which makes “members of one pool spread across two groups” the default outcome rather than an unusual one. A pool whose members are meant to reach each other should be pinned to node groups that share one.
The medium and the transport are independent. A DRAM group is as entitled to a fabric as a
VRAM one, and this rendering reads the medium nowhere: the protocol decides the network, the
capabilities and the device, while the medium decides only what the segment is made of and what it
is charged to.
What a member’s Pod requests follows the group’s medium. A DRAM member requests host memory for
capacityPerMember + localBufferSize. A VRAM member’s segment is device memory, which nothing
requests: host memory carries localBufferSize only, and the device memory is claimed by
allocating it.
The engine sharing that accelerator does not know the member took a slice of it. Nothing in
Kubernetes accounts for device memory, so the two are held apart by sizing capacityPerMember
against the engine’s own memory fraction, and by nothing else. Give the engine a
gpu_memory_utilization that leaves capacityPerMember free, and expect whichever starts second to
fail on allocation if you do not.
A member group asks for no accelerator extended resource, because the claim would not be true: in the tested Mooncake v0.3.13.post1, a member allocates its whole segment on one device, so a resource request would take a whole accelerator away from inference to account for a slice of one.
One member’s segment is therefore on one device, whichever the container sees first. An eight-accelerator node keeps all eight available to inference, and one member contributes a slice of the first.
VRAM group accelerator access
A VRAM group’s container needs the vendor’s user-space driver, which is what the store’s client loads to allocate a segment at all. That driver is the node’s and is never in the image, so it has to be brought in. Nothing is inferred; every route is written on the group:
| Field | Purpose |
|---|---|
members[].extraEnv |
The vendor runtime’s own trigger, where one exists — NVIDIA_VISIBLE_DEVICES: all makes the NVIDIA toolkit inject the driver and every device. Ascend has no counterpart. |
members[].runtimeClassName |
The vendor container runtime named explicitly, for a cluster where it is not the default runtime. |
members[].hostPaths[] |
The driver tree taken from the node directly, for a cluster running no vendor runtime at all. |
members[].securityContext |
Privilege, when a mount alone is not enough to open what was mounted. |
Privilege alone does not cover this. It opens the node’s device tree under /dev, which is
where the device nodes are and is not where the libraries are: Ascend’s driver is under
/usr/local/Ascend/driver with DCMI beside it, and NVIDIA’s libcuda.so is injected by the
container runtime. A privileged member with neither a runtime class nor the driver mounted starts,
reports healthy, and fails to allocate.
securityContext is merged onto what the group’s protocol already earned, field by field, with
capabilities.add unioned. A RDMA or EFA group therefore keeps IPC_LOCK and SYS_RESOURCE
whatever it declares. Without them the transfer engine fails when it registers memory, long after
the container looked healthy. Declaring a capability adds it; a group that must hold neither of those
two declares a protocol that does not ask for them.
Each hostPaths[] entry is {path, mountPath, type, readOnly}. Name a type: left empty, the
kubelet checks nothing, so a missing path becomes an empty directory in the container and the member
starts anyway. A mount path is refused if it duplicates another entry’s, if it is /dev/infiniband
(where a fabric’s device plugin injects the granted device nodes, refused under every protocol since
that can change later), or if it is this group’s own localDisks[].path.
Two worked groups. The NVIDIA one needs no mounts at all, because the container runtime injects the driver once the variable tells it which devices to inject:
members:
- nodeSelector:
kvcache-vram: "true"
medium: VRAM
image: gpustack/mirrored-mooncake:0.3.13.post1-cuda13.0
capacityPerMember: 8Gi # the engine's gpu_memory_utilization must leave this free
extraEnv:
- name: NVIDIA_VISIBLE_DEVICES
value: allThat variable is honored only while the toolkit’s
accept-nvidia-visible-devices-envvar-when-unprivileged is on, which is its default and which
NVIDIA’s own hardening guidance turns off. On a cluster that turned it off, use
runtimeClassName instead.
The Ascend one has no such variable, and this cluster runs no vendor runtime, so the driver is mounted from the node:
members:
- nodeSelector:
kvcache-vram: "true"
medium: VRAM
image: gpustack/mirrored-mooncake:0.3.13.post1-cann9.1
capacityPerMember: 32Gi
hostPaths:
- path: /usr/local/Ascend/driver
mountPath: /usr/local/Ascend/driver
type: Directory
readOnly: true
- path: /usr/local/dcmi
mountPath: /usr/local/dcmi
type: Directory
readOnly: true
- path: /usr/local/bin/npu-smi
mountPath: /usr/local/bin/npu-smi
type: File
readOnly: trueA VRAM group may also carry a local disk tier , and that combination is measured on NVIDIA. Objects written past a 256Mi device segment left memory for the tier and read back with matching digests. No upstream test covers the pairing, so the reading is this project’s own.
The equivalent on Ascend (cann) and AMD (rocm) is still unmeasured. Treat the pairing as
verified on NVIDIA and as unverified elsewhere, rather than as guaranteed everywhere.
A writer against this tier must retry. A put that fails on a full segment means eviction has not caught up, not that the tier is full: the offload heartbeat runs on an interval, so a writer that abandons on first failure reads “eviction in progress” as “cannot write”.
Reachability is a port range, never a list. The transfer engine picks its data ports at random,
and the peer-to-peer plane is what binds them. Write firewall and NetworkPolicy rules between member
nodes, and from engine clients, as a range. The rendered Pod declares no fixed data-plane
containerPort, because a fixed list would be a false statement.
The management port is fixed, and on a host fabric it lands on the node. A member serves its HTTP
API on 8080 + <group index> (8080 for the first group, 8081 for a second). A TCP group
holds that port inside its own pod network namespace, but an RDMA or EFA group holds the host’s,
so on every node such a group selects that port must be free. Reserve one port from 8080 upward per
member group.
Why it moves per group — two host-network groups whose node selectors both match one node place two host-network Pods on it. On a single fixed port only the first binds; the second runs, never passes readiness, and reports nothing about why.
Members advertise their Pod IPs for client connections. Target those addresses in data-plane network rules. With host networking, the Pod IP is the node address.
Every segment mount gets a new identity and transfer port. After a member restarts, clients must discover its current endpoint rather than reuse a previous connection address.
Status and conditions
$ kubectl get kvcb
NAME TYPE PHASE ENDPOINT CAPACITY
mooncake-dram Mooncake Ready mooncake-dram-leader.gpustack-system.svc:50051 12Gi
Five phases: Provisioning, Ready, Degraded, Error, Deleting. Ready carries no
phaseMessage; every other phase carries one.
Conditions report the axes: LeaderAvailable, MembersMounted, CapacityObserved, PoolWrites, Deletable and
RolloutComplete. Two more appear only where they have something to judge: ElectionObserved
when leader.electionBackend is Kubernetes (including at one replica), and
TierWasEmpty when a member group carries a local disk tier
.
Those last three do not move the phase, and that is deliberate. A rollout in flight, an election
that has not happened, and a disk tier found holding data are all states in which the backend serves
normally. Reporting them as Degraded would put a storage arrangement in the same field as a
leader nobody can reach. Read the conditions for them; the phase will not tell you.
A member that is starting is not a shortfall; a member that is stuck is one. A Pod still pulling
its image is left alone. Holding it against the backend would report Degraded for the length of
every rollout. But one whose container will not start, or that no node will take, is never going to
arrive, so it reads Degraded even while the other members serve, and phaseMessage carries that
Pod’s own reason.
A member reads Ready only once its segment is mounted. Its container carries a readiness probe that connects to the entrypoint’s REST port, and the entrypoint mounts the segment before it serves that port. Readiness is therefore evidence of the mount, not of the process.
Why the probe is load-bearing — without it the kubelet reports Ready as soon as the container runs. Every ready member Pod is held to the leader’s listing, so that window would read as a shortfall and move a healthy backend to
Degradedfor the length of every rollout.
A listing too large to publish is withheld, never truncated. Past what status.members can carry,
the phase reads Degraded with reason ListingTooLarge and the previous listing is kept.
Why — every entry is republished on each pass, so publishing past the object size the API server accepts would make every status write fail from then on while the read that produced it reported success. A truncated list would be worse: it reads exactly like a backend that lost members.
Status is polled every 15 seconds, not only refreshed on events. Everything above is read over
HTTP from the leader, and a store whose contents move while its Pods sit still produces no Kubernetes
event at all. An external backend produces none ever, since this operator owns no workload for it.
So kubectl get kvcb -w moves on its own.
15 seconds is an interval, not a maximum age. The timer starts after a pass finishes, and a pass makes up to three sequential HTTP reads. More importantly,
status.membersis deliberately retained when a read of the segment listing fails (a stale list plusMembersMounted=Falseis more honest than an empty one), so it has no age bound at all while that read keeps failing. The condition is what says whether the list was refreshed; the list alone never does.
Capacity is observed, not derived. status.capacity is read from the leader’s own counters,
never from what the spec declares. Nothing multiplies capacityPerMember by a replica count. A
backend with no disk tier reads master_total_capacity_bytes; one with a tier reads that plus
master_total_file_capacity_bytes. Adding two observed families is not the same thing as
adding up what members were asked to provide.
status.capacity.total is capacity, not usage, and a disk tier contributes the ceiling the
member declared (published as soon as the member registers, before anything is written there).
PoolWrites reports pool-wide write activity. True means the leader has seen a put end since
its current process started, or a member reports positive allocator_used_bytes. False means
the leader has seen both a put start and a put revoke while every member reports zero allocated
bytes. Unknown covers no put start, an unfinished put, or a missing reading.
The process lifetime is the observation window because these counters reset when the leader restarts; a single status poll could miss a short write.
This condition does not identify which deployment attempted a write, and a previous failure in that window does not prove that every current write fails.
To ask whether the disk tier is holding data, the figure to read is not on the CR at all. See Bucket writes .
Capacity is absent (not zero) while the leader is starting. /metrics is ungated: a leader
that is up but not serving answers 200 with a well-formed exposition whose gauges all read zero, and a
zero is indistinguishable at the parser from a genuinely empty cache. Publishing is therefore gated on
service_ready, not on the scrape succeeding.
status.members[] is read from the leader’s segment listing, one entry per listed segment. Each
row carries the leader’s segmentID, clientID and advertised segmentName; the list is keyed by the
unique segment ID because several members may legitimately share a name.
The leader is what allocation goes through, so a running member Pod it does not list holds nothing
and is counted in MembersMounted’s message instead. The two fields the listing cannot supply
(node name and medium) are joined in from the member Pod behind that segment, and left empty
rather than guessed when nothing matches.
Two ready host-network member Pods that share an address make Pod attribution ambiguous. Their
rows remain publishable because their segment and client IDs are distinct, but neither Pod exposes
the ID that maps a row back to it. MembersMounted goes False with reason
AmbiguousMemberIdentity; each row still carries the node and the medium its candidates agree
on, neither being a fact about one Pod, and leaves empty whichever of the two they dispute.
Two groups on one node do not collide by themselves. A TCP member advertises its own pod IP, so
each segment carries a distinct name even though both Pods answer to the node’s name; the collision is
on host-network paths (RDMA and EFA), where both Pods hold the host’s network namespace and
advertise the node’s address.
The remedy is to give the groups node selectors that keep them on different nodes.
A failed listing scrape keeps the previous list and sets MembersMounted=False; a failed capacity
scrape clears the figures. That asymmetry is deliberate: capacity is two pointers and has an
“absent” that means not observed, while an empty list is a legible value meaning no segments, so
clearing it would publish a falsehood.
Mooncake masters before 0.3.12 do not report segment membership. They can serve requests and
report Ready while MembersMounted is Unknown/SegmentListingNotServed and status.members
is empty. Waiting does not add a capability missing from that version.
Other listing failures, including server errors or an unreachable master, report ListingFailed
and retain the previous member list.
PoolWrites on such a leader is still True once it has seen a put end, since that counter is on
/metrics; before that it is Unknown with the same reason, because its other verdicts read the
members’ allocations off the listing.
Two states still read MembersMounted=False and Degraded there: a member Pod that will not start,
read from the Pod’s own status, and a status.capacity.total of zero, reported as NoSegments
(the gauge sums the mounted segments, so zero is every member unmounted).
Growing and shrinking a group
Widening members[].nodeSelector adds members without restarting the ones already running. The
DaemonSet places a Pod on each newly matching node; every existing Pod keeps its UID and its restart
count, leader included. The DaemonSet is left on OnDelete, and restarts are decided from a
comparison over the whole Pod template except the node selector: a widening matches that comparison,
while an image, argv, environment, resource or fabric change fails it and every member is recreated.
A Pod runs the template it was created from, so a setting cannot protect the same edit that removes it. Anything rendered into the member Pod (the shutdown hook, its grace, the environment) reaches a member only when that member is recreated. A departing member leaves with what it started with.
To make such a setting apply to a shrink, do it in two steps: change only the setting and wait
for members to be recreated with it (their pod-spec-hash annotation moves), then narrow the selector
or remove the group. scaleIn.gracePeriodSeconds below is the case this bites today; the property
belongs to the Pod template, not to that field.
Existing objects are not rebalanced. How fast the cluster converges onto a new member depends on
allocationStrategy: FreeRatioFirst (the default) biases new writes toward the emptier member,
Random does not.
Group identity by position
A group has no name. Its position in members is what the DaemonSet’s name, its immutable
selector labels and its members’ HTTP port are all derived from, so moving an entry in that list
leaves every one of those in place and changes only the spec underneath it.
Moving a group to another position is refused at apply time. The refusal names both positions. Two shapes reach it:
| Edit | Outcome |
|---|---|
| swapping two entries | refused |
| removing a group ahead of others, which shifts the rest up | refused |
| appending a group | allowed |
| removing from the end of the list | allowed |
editing a group in place, including widening its nodeSelector |
allowed |
To take a group out of service without removing it, narrow its nodeSelector until it matches no
node. The group keeps its position, every later group keeps its DaemonSet, and nothing is rebuilt.
Positions are used instead of names because a DaemonSet’s selector cannot change after creation: a
name free enough to make reordering possible would force every member DaemonSet to be deleted and
recreated, and the entire cache would go with them.
The rule recognises a group that arrived unchanged at a position another group left; it does not recognise every reorder. Without a name there is nothing else to recognise a group by, so two shapes are knowingly admitted:
- a reorder combined with an edit to the same group, which is indistinguishable from two ordinary edits;
- removing a group when a later group is identical to the one taking its place,
[A, B, C]becoming[A, C, C]. That produces the same two lists as editing position 1 to match an unchanged position 2, which is how the second of two look-alike groups is taken out of service, so refusing it would forbid the operation recommended above.[A, B, C]to[A, C], with no look-alike to arrive in the gap, is still refused.
Both admitted shapes leave one trace: the resulting members holds two identical groups. What
the rule buys is that the mechanical reorder (the one a rewritten manifest produces) is reported
instead of silently rebuilding members against another group’s spec.
Shrinking a group discards the cache held by the removed members. Narrowing the selector or removing a node immediately unmounts its segment; the operator does not migrate its cached data. The termination grace period gives the process time to finish shutdown and does not preserve that data.
scaleIn.gracePeriodSeconds holds the process, not the tier. A member with a
local disk tier
gets a preStop hook that deregisters the tier with the
leader and then waits out the grace. A group with no tier renders no hook and the setting is inert.
With the example image, deregistration takes effect when the hook runs. A key held only by that tier becomes a clean miss immediately, and the process waits for the full configured period even if no reads remain. Size the period for the departing member’s local work.
This timing is specific to the example image. spec.image selects the backend implementation;
other images may deregister later or wait differently. The operator controls the value sent to
the endpoint.
spec:
connection:
managed:
scaleIn:
gracePeriodSeconds: 30 # 0..3600The Pod’s termination window is derived from it, as gracePeriodSeconds + 60, rather than being
a second field beside it. That is what makes the relationship hold: two independent fields could be
set so the kubelet kills the container in the middle of the wait, and no validation makes that
impossible. It only makes it checkable.
The upper bound of 3600 is the member endpoint’s own; above it the call is refused with a 400, so a
larger value would render a hook that fails every time it runs. None of it makes a shrink lossless:
the memory segment is still dropped, and the disk tier stops answering as soon as the hook runs.
Migrating a member’s data before it leaves (the store’s drain job API) is not offered here. It is stateful orchestration, and it reaches only the memory and NVMe-oF replicas: it cannot name the segments of the disk-backed ones and skips those keys without counting them as blocked, so a drain over a backend with a disk tier reports success while leaving that tier’s data where it was. Its own success signal does not tell you that happened.
The external mode
connection.external points at a backend somebody else runs. The operator creates nothing
(no Deployment, no Service, no DaemonSet) and only observes.
connection:
external:
endpoints:
- name: Client
address: mooncake.example:50051
- name: Admin
address: mooncake.example:9003Both roles are required, so the list holds exactly two entries: entries are keyed by name, which
has two values, and admission refuses a list missing either one. Each address is validated as
host:port at admission. A blank or portless one would otherwise be mirrored into status and
handed to an engine that cannot dial it.
Admin is what this operator reads (health, metrics and the segment listing), and Client is
what an inference engine connects to. The addresses are mirrored into status.endpoints unchanged.
Two behaviours differ from the managed mode:
- An address that does not answer is reported as an
Error, not as a backend still starting. A managed leader is excused while its own Deployment has no ready replica; an external address was declared to name something that already runs, so there is no Pod to wait on and a mistyped endpoint would otherwise sit atProvisioningforever. status.members[]carries no node name or medium, because those Pods are not this operator’s to look up. Capacity is the sum of both pools, since an external object names no medium to pick one by.- A redirect from the
Adminaddress is never followed. That address belongs to whoever wrote the spec, and honouring a3xxfrom it would read some other host with the operator’s network identity, then copy an excerpt of the answer into a status readable by anyone who can read the object. The redirect is reported as the response it is.
Two objects on one leader
Two KVCacheBackend objects may name the same leader, and nothing in the operator notices. For
a managed backend the object is the leader, so two objects are two leaders. For an external one the
object is a declaration of addresses, and one leader is reachable under more than one spelling:
by Service name in one object and by IP in another, with or without a trailing dot or an explicit
default port.
Why no check — every identity the operator could compare is either editable or needs the leader reachable at admission. The comparison cheap enough to run (byte-identical addresses) catches the copy-paste case and misses the one a real deployment produces, which is the same leader spelled two ways. A check that catches the easy half invites the reader to trust it for the other half. This is tracked in issue #288 .
Three consequences follow, and the third is the one to read before deciding the duplication is
safe. The reuse-domain uniqueness rule is enforced between Bindings whose pools name the same
backend object, so two Bindings reaching one leader through two objects are both admitted on one
domain.name:
- The quota of a shared domain flips and never settles. The leader keeps one ledger entry per
tenant, and each pool’s reconciler converges that entry toward its own Binding’s
quota.ceilingon every pass. Each pass reads the other’s figure, finds it wrong, and writes its own back. - The symptom of an undersized quota is a low hit rate and nothing else. Exceeding a tenant’s quota does not refuse the write: the store frees room by dropping that tenant’s own older objects and retries, irreversibly and without any counter moving. So the flipping above never surfaces as an error. It surfaces as a cache that keeps losing content nobody asked it to lose.
- Two Bindings on one
domain.namewith a differentblockSizeordtypecorrupt each other’s blocks. The reuse identity an engine is handed is the domain name alone. Each Binding hands its owndtypeto its own engines, andblockSizereaches no engine at all. So two differently-shaped caches land under one identity, which is the silent cache pollution a wrongblockSizeordtypecauses, reached here without either value being wrong.
If you point two objects at one leader, either keep their pools’ Bindings on different
domain.name values, or make sure every Binding that shares a name also shares its blockSize,
dtype and quota.ceiling.
Operating notes
replica_num is engine-side. How many replicas of a stored object Mooncake keeps is a per-Put
argument the caller supplies through its connector configuration. It is not controllable from this CR.
One benign startup line, documented so nobody files it as a bug. It appears on every client start and is harmless:
E transfer_metadata.cpp:991] Local segment descriptor not found
Client-side environment knobs, observed at startup:
MC_TE_METRIC=1 enable transfer-engine metrics (OFF by default)
MC_STORE_CLIENT_METRIC_BANDWIDTH client bandwidth summary
MC_STORE_MEMCPY unset => auto-detected ("TCP-only environment, memcpy enabled")
MC_METADATA_SERVER / P2PHANDSHAKE the transfer engine's own low-level metadata knob
MC_TE_METRIC is rendered on every member group and cannot be turned off from the CR. The other
three are set on the workload and not here. It is reserved in members[].extraEnv for the reason
every rendered name is: the hatch appends, and a container carrying one name twice leaves the winner
to the runtime.
The switch reports throughput and a task-latency distribution to the container log, on an interval that prints nothing when nothing moved. The leader’s Prometheus surface counts keys and bytes and measures no data plane, so without it the receiving end of every transfer is unmeasured while the sending end already reports.
MC_METADATA_SERVER is not the variable this operator renders. The member is configured
through the store client’s key, MOONCAKE_TE_META_DATA_SERVER. Two names for the metadata plane is
exactly the near-miss that gets one of them typed into a template.
A backend in use cannot be deleted. While status.usedBy names a consumer the finalizer holds,
the object stays at phase: Deleting with the claimant named in its message, and the workloads keep
running. Clearing the last claim lets the teardown complete.
A backend in use also cannot have leader.multiTenancy turned off. The webhook refuses the edit
while status.usedBy names a consumer, and the refusal names them. Remove the consumers first. See
KV Cache Pool
for what the withdrawal costs on their side.
Why — the flag decides whether the master keeps a per-tenant ledger, and a consumer’s quota is both written and released through that ledger. Withdrawing it under a live consumer therefore takes away the way that consumer is unwound, and the cost only appears when something is deleted. The refusal reads
status.usedByas it is written, not as it resolves: an entry naming an object that no longer exists refuses the edit too, and that list is where it is cleared.
The teardown deletes the workloads first, and the object disappears last. The leader Deployment, its Service and every member DaemonSet go before the finalizer comes off. It waits for them to be gone, not merely for the deletes to be accepted, so the object going away means the backend is gone rather than scheduled to be.
Why not leave it to ownership — they are owned dependents, so the collector would reach them either way. But between the finalizer coming off and it running, the leader is still serving on an address nothing accounts for. Ownership is the safety net, not the mechanism.
That wait is bounded, one workload at a time. Each gets the termination grace its own Pod template declares plus a couple of minutes, timed from its own deletion timestamp. Past that it stops being waited for, and a warning Event on the workload names the nodes its Pods are still terminating on. Deleting a backend therefore finishes even when a node it ran on has stopped answering.
Why its own grace and not one number — a member group with a disk tier derives its Pod’s grace from
scaleIn.gracePeriodSeconds, which reaches an hour, so any constant short enough to bound an unreachable node would abandon a group draining exactly as configured. That grace works as the clock because the kubelet treats it as a hard kill deadline: a Pod outliving it is not a slow one, it is one whose kubelet is not acting.
Only objects carrying this backend’s own note are deleted. The names are derived, so an unrelated object can hold one, and a delete has to be surer than a name. The member sweep finds its DaemonSets by the same note it then checks, rather than by the identity labels. Discovering on one key and judging on another is how an object goes missing from its own teardown.