KV Cache Leader
A KVCacheBackend runs a leader process for cache metadata. The operator manages its Deployment,
Service, probes and Kubernetes Lease election.
This API says leader; the artifact says master, and every rendered flag, environment variable and metric keeps the vendor’s spelling.
Contents
The Deployment and the two probes
The leader is a Deployment plus a ClusterIP Service publishing two ports (50051 for engine clients
and 9003 for the admin surface, which serves the Prometheus exposition and the HTTP admin API on one
port).
extraArgs exists to reach settings this API does not name, so this page does not enumerate what
may be set through it. An upstream flag can also reach a setting this API does name, spelled
differently. The object then reports the value the operator rendered while the process runs the
value from extraArgs, and nothing reports the divergence. Accepting that gap is the
administrator’s trade-off.
The CXL switch was one instance of exactly that shape, and the webhook now refuses it by name. The
CXL row in KV Cache Backend
carries which keys and why. It is recorded as the shape
to expect, not as a live hazard.
port is refused in leader.extraArgs, and it would have moved nothing. It is the store’s
deprecated spelling of rpc_port, which this operator always renders, so the rendered one wins: the
key reads as a port that moved without moving one.
metrics_host is reachable on purpose, and on an IPv6 cluster it is required. It moves the
address the admin surface binds to, not the port. The store’s default 0.0.0.0 is IPv4 only, so
where the Pod’s address is IPv6 nothing answers the probes and the leader never becomes ready; ::
listens on IPv6 and, on a dual-stack host, on both.
Any other concrete address breaks both probes, because the kubelet reaches them at the Pod’s
own. The leader then stays not-ready rather than reporting the cause, so 0.0.0.0 and :: are the
only two values worth setting.
default_kv_lease_ttl is reachable on purpose, and it is the one flag this operator renders and
still lets you override. A lease protects a cached object from eviction, and the store’s own ten
seconds expires before a second engine replica asks for a block the first one just read. The
operator renders 5m instead. An entry in leader.extraArgs renders after it, and wins.
A lease is granted when an object is read, never when it is written, so this value protects the recently-read set rather than everything written. Raise it where replicas share prefixes over longer turns; the cost arrives only once the set read within the window outgrows the store itself.
leader.extraEnv is the same hatch for the environment: every entry renders after the derived
variables, and a name the renderer derives (the pod’s own identity variables) is refused with the
same message a member gets. The schema keys the list by name, so one name cannot carry two values.
replicas defaults to 1, and 5 is the ceiling in the webhook and in the schema alike: only
one leader ever serves, so further replicas are spare processes rather than capacity. More than one
requires electionBackend: Kubernetes; the webhook refuses None at that count.
With electionBackend: None the Deployment runs one replica whatever replicas says. The
webhook refuses a larger count, and this clamp also protects clusters without the webhook.
electionBackend defaults to Kubernetes, even at one replica. The first leader campaigns for
a Lease and starts with its election flags and API token. Raising replicas from one to three adds
standbys without changing that leader’s Pod template or rolling the members.
Set electionBackend: None only for a single leader whose image cannot run the Kubernetes Lease
backend. Switching between None and Kubernetes changes the leader and member Pod templates, so
it restarts them and loses their DRAM cache contents.
Increasing the leader count normally takes one update. If the live Deployment still has an
unelected template, the operator first enables election at one replica, then adds standbys after
the template rolls out. Until that step completes, RolloutComplete reports
False/ReplicasPending.
Returning to one replica leaves election enabled unless you set its backend to None.
The update strategy follows the replica count, and the two cases are opposites. At one replica
the Deployment uses Recreate: an update stops the old master before starting the new one, so expect
a gap with no master on every image or flag change. Members keep their segments across it and
re-register. Above one replica it rolls instead, with maxSurge: 1 and maxUnavailable: replicas,
because Recreate would take every standby down together with the leader and leave nothing to elect.
A rollout still has a window with no serving master, and standby replicas can shorten it. A floor of zero available replicas is what lets the old leader go, so it can go before a replacement has taken the Lease. The window depends on election and activation; replacements already running as standbys do not have to wait for scheduling or image pulls.
The two probes deliberately take different paths.
| probe | path | gated? |
|---|---|---|
| readiness | GET /get_all_segments |
yes — 503 until the service plane is active |
| liveness | GET /health |
no, and it must not be |
/health answers 200 in every state, so using it for readiness is the same as having no readiness
probe. Using a gated route for liveness would kill a leader that is slow to activate.
The health document has four fields that matter:
{"status":"ok","role":"leader","ha_state":"serving","service_ready":true}status is a hard-coded constant. It reads "ok" on a leader that is serving nothing.
service_ready is the only verdict in the document, and every readiness decision rests on it.
With electionBackend: None, a single leader reports service_ready: true from its first answer.
Under the default Kubernetes election, readiness waits for the process to win the Lease.
High availability
The default leader.electionBackend: Kubernetes elects through a Kubernetes Lease at every
replica count. The Lease is named <backend>-leader in the operator’s namespace. Scale an already
elected leader by changing only its count:
spec:
connection:
managed:
leader:
replicas: 3leader.memberAddressing is optional and only selects how members find the
winner. It does not turn the election on. With electionBackend: None, members
use the leader Service even if memberAddressing: Lease is set.
A published kvcacheai/mooncake image cannot do this, on either side. Leadership backend
availability is a compile-time switch and every option ships off:
| Role | Behavior |
|---|---|
| leader | answers UNAVAILABLE_IN_CURRENT_MODE, runs as a permanent standby |
| member | answers Invalid HA backend entry, exits, CrashLoopBackOffs |
Use an image built from pack/mirrored-mooncake
for
spec.image and for every members[].image, on Mooncake 0.3.12 or later: an electing leader is
also rendered -pod_name and -pod_namespace to label the winner, and a 0.3.11 master exits on both.
A lease-less image can run one leader with electionBackend: None. It cannot serve a backend with
multiple leader replicas. Set None when creating a backend with such an image; an omitted field
selects Kubernetes and renders election flags that the image cannot use.
A member group on RDMA, ROCM or CANN runs under high availability on the build that carries
its transport: every mirrored-mooncake target (the default build and the cuda, cann and
rocm variants alike) compiles the Lease backend in. The vendor axis and the leadership axis are
orthogonal.
The matching rule is unchanged from the backend page
: the default build
covers RDMA over DRAM (the transport has no compile switch to leave off, and rdma-core is
installed), and a ROCM or CANN group names its variant.
What none of those three has is a real-machine run under an election: the Lease backend’s presence is asserted per build target, the fabric data path is not.
MUSA and MACA keep their own rule because this project builds no variant for either, by
intent: a group on one of them runs an image you built, so whether it also carries the leadership
backend is a property of your build rather than of anything here.
EFA needs no vendor runtime either, only libfabric, so mirrored-mooncake compiles it into the
default build. The image build proves the transport installed by running a target-mode bench told
--protocol=efa and refusing the transport map’s “Invalid protocol”: the device-less builder’s
“No EFA devices found” and an EFA-capable builder’s clean run both pass. An EFA member group runs
under high availability, on nodes that have the AWS EFA driver installed.
Which address a member follows under that election is the memberAddressing choice documented
below; issue #279
, which tracked it, is
closed. The leader Service address is the default and the Lease coordinates are the option.
enable_oplog is refused in leader.extraArgs, and not as a policy choice: the store’s
operation log requires the etcd backend, which cannot be compiled together with the Lease backend, so
the flag produces a leader that refuses to start. Standbys rebuild from member remounts alone
instead.
etcd_endpoints is refused as well, and it is out of reach twice over. The store reads it only
where ha_backend_connstring is empty, which the election never leaves empty, and the etcd backend it
names is the one this image cannot carry.
rpc_address and rpc_interface are refused there too, because the election renders the
first. The store folds it with the RPC port into the string it campaigns with, so that one value is
both the election’s identity and the address written into the Lease for members to dial. Each replica
advertises its own Pod IP; a value supplied by hand would point every member at one host, chosen
without knowing whether the Pod answers there.
The healthy steady state reads N desired / 1 ready. Exactly one leader serves; the rest are
standbys, and a standby is deliberately not ready: that is what keeps it out of the leader
Service’s endpoints, so an engine never connects to a process that cannot serve. Read without that
context, 3/1 is what a broken Deployment looks like. kubectl get deploy during a healthy
failover briefly shows 2 ready as the old leader steps down; both readings are normal.
To see which Pod is serving, read the label the store itself sets when it wins:
$ kubectl get pod -l mooncake.io/store-role=leader
Both roles get a ServiceAccount, and they are different accounts. The operator renders a
ServiceAccount, Role and RoleBinding per role in its own namespace, and names them on the Pods:
| role | grant | why |
|---|---|---|
| leader | leases: create, get, update; pods: patch |
takes the Lease, and labels its own Pod |
| member | leases: get |
follows an explicit Lease address, and nothing more |
The member’s is narrower on purpose: a shared account would let any member take the Lease from the leader it is following.
A member’s MOONCAKE_MASTER defaults to the leader Service address. A standby is not ready, so
the Service resolves to the serving leader. With memberAddressing: Lease and an election running,
the member instead gets k8s://<namespace>/<lease> and reads the current holder itself. Without an
election, both values render the Service address.
| Value | Member receives | Cost |
|---|---|---|
Lease |
the Lease’s coordinates, read by the client itself | the member must reach the API server, so its image must carry the leadership backend |
Service (default) |
<backend>-leader.<namespace>.svc:50051 |
endpoint propagation after an election |
Service pays endpoint propagation after an election. Lease avoids it by having the member read
the holder itself, which requires access to the API server. One failover comparison found no
meaningful timing difference; the Service endpoint transition was inferred rather than observed.
Retest if election timing changes. Changing this value rolls every member group when HA is active.
Standby leaders can shorten the outage; they do not keep the cache. A standby holds no data. The replica that takes over, like a restarted single leader, rebuilds the members’ segment list from their remounts and learns none of the keys in them, so a DRAM-only key misses until it is written again.
Keys fully written to a member’s local disk tier may recover after a disk segment and its objects are registered again. The directory must still be available and the scan must complete.
A member restart also loses its DRAM bytes. In the tested Mooncake version the restarted member gets a new client ID, and a stale replica record from before the restart can block that new ID from registering its disk keys. Disk recovery is therefore conditional, even when the files survive.
What a second replica buys is time. On a single-node test cluster a failover left the store unusable for about 16 seconds and a single-leader restart for about 30. Neither figure is a bound on when each disk key first becomes a hit, and on a real cluster a single leader also waits for scheduling and image pulls.
Run one leader by default. Add replicas when the shorter service gap justifies the standby processes. The default single leader already uses the Lease and the two accounts described above.
The store’s snapshot is not offered, and its flags are refused in leader.extraArgs. A
snapshot records where each key sits in member memory, and restoring one does not check that the
memory still holds that key. An engine reset frees the memory, and a standby’s startup snapshot
can become stale before it takes over. Restoring that index can return another key’s bytes to the
engine, producing an incorrect KV block.
enable_snapshot and enable_snapshot_restore are refused with that reason. Every other
snapshot_* key is refused because it is read only under one of those two, and memory_allocator
because it moves the store off the allocator these rules were traced under.
A missing grant fails differently on each side, and one of them is silent. A leader that cannot
reach the Lease retries every second forever. Liveness is ungated, so nothing restarts and the
Deployment sits at 0/N ready with no message naming the cause. A member that cannot read it gives
up after twenty tries and CrashLoopBackOffs. Check both accounts exist before reading 0/N as a
store problem.