High Availability Operations
Every control-plane component the chart deploys elects a leader, so extra replicas stand by. They buy failover, not throughput. A highly available install raises each replica count and turns on its disruption budget; everything else is one pod per node (device managers, NFD worker, both CSI node DaemonSets).
Configure it in values.yaml: no HA values file ships. Every knob below sits at its chart’s default,
and the caveats are inline.
Contents
Node count prerequisites
A DoNotSchedule spread needs at least as many schedulable nodes as your largest replica count; below
that, surplus replicas stay Pending forever. This page assumes three replicas on three nodes; with
fewer, lower the counts or use ScheduleAnyway, a preference. The components tolerate control-plane
taints, so three control-plane nodes spread fine.
They tolerate those taints and no others, a cordon’s included, so a pod evicted by a drain never
lands back on the node being drained. Where your nodes carry a taint of your own, add it to each
component’s tolerations.
Earlier releases let the Kueue controller manager, the NFD master and the NFD gc tolerate every taint;
an install that relied on that must set their three lists before upgrading. Upgrade with
--reset-then-reuse-values, not --reuse-values, or the new lists never reach the release; see
Upgrading the Chart
.
Per-component settings
| Component | Replicas | PodDisruptionBudget | Node spread |
|---|---|---|---|
| Worker (control plane) | worker.replicas |
worker.podDisruptionBudget.enabled + .minAvailable |
worker.topologySpreadConstraints, or worker.affinity |
| Kueue controller manager | kueue.controllerManager.replicas |
kueue.controllerManager.podDisruptionBudget.enabled + .minAvailable |
kueue.controllerManager.topologySpreadConstraints |
| NFD master | node-feature-discovery.master.replicaCount |
node-feature-discovery.master.podDisruptionBudget.enable + .minAvailable |
node-feature-discovery.master.affinity only |
| NFS CSI controller | csi-driver-nfs.controller.replicas |
— none — | csi-driver-nfs.controller.topologySpreadConstraints (vendored patch) |
| S3 CSI controller | csi-driver-s3.controller.replicas |
— none — | — none — |
Those “none” and “only” cells are upstream chart limitations: each changes what three replicas get you, and each is spelled out below.
Example values
worker:
replicas: 3
podDisruptionBudget:
enabled: true
minAvailable: 2
topologySpreadConstraints:
- maxSkew: 1
topologyKey: kubernetes.io/hostname
whenUnsatisfiable: DoNotSchedule
kueue:
controllerManager:
replicas: 3
podDisruptionBudget:
enabled: true
minAvailable: 2
topologySpreadConstraints:
- maxSkew: 1
topologyKey: kubernetes.io/hostname
whenUnsatisfiable: DoNotSchedule
labelSelector:
matchLabels:
app.kubernetes.io/name: kueue
control-plane: controller-manager
node-feature-discovery:
master:
replicaCount: 3
podDisruptionBudget:
enable: true
minAvailable: 2
csi-driver-nfs:
controller:
replicas: 2
strategyType: RollingUpdate
topologySpreadConstraints:
- maxSkew: 1
topologyKey: kubernetes.io/hostname
whenUnsatisfiable: DoNotSchedule
labelSelector:
matchLabels:
app: csi-nfs-controller
csi-driver-s3:
controller:
replicas: 2
strategyType: RollingUpdatehelm upgrade gpustack-operator gpustack/gpustack-operator \
--namespace gpustack-system --reset-then-reuse-values --values ha.yamlWorker (control plane)
The worker runs the aggregated extension API server and the scheduling-chain controllers in one
process. Only one replica reconciles, because leader election is always on. Every replica serves the
extension API and the admission webhooks, so with one, losing its node takes kubectl get instancetypes
and Pod admission down until it reschedules.
Node spread belongs in worker.topologySpreadConstraints. Its default preferred pod anti-affinity is
deliberate: replicas stay off one node without any becoming unschedulable, so three still come up on two
nodes. An entry omitting labelSelector gets the worker’s own; worker.affinity replaces the
default rather than adding to it.
Kueue controller manager
Kueue’s managerConfig elects a leader, so standbys do not reconcile. Each replica still serves Kueue’s
admission webhook, whose failurePolicy: Fail means one replica losing its node blocks Pod creation in
every namespace Kueue manages until it reschedules. HA buys the most here.
- The spread constraints carry no selector of their own. Unlike the worker’s, they render as given, so
a
DoNotSchedulespread needslabelSelectorspelled out:app.kubernetes.io/name: kueuepluscontrol-plane: controller-manager, as in the example. Omit it and the spread counts every pod in the namespace. - The Kueue chart has no affinity key, so spread constraints are its only placement control.
NFD master
The NFD master turns detections into node labels: while it is down, no node is (re)classified and the chain stalls for new or changed nodes. Labelled nodes and admitted workloads are unaffected.
- NFD spells the budget key
enable, notenabled: a strayenabled: trueis schema-valid and does nothing. - Above one replica, NFD’s chart adds
-enable-leader-electionfor you. Standbys watch without writing. - NFD’s templates render no topology spread constraints, so spreading the master means
node-feature-discovery.master.affinity, which replaces NFD’s preference for control-plane nodes; re-state it if wanted.
The garbage collector stays at one replica: a stalled GC only delays cleanup of a departed node’s objects, which nothing reads.
The two CSI controllers
Losing a CSI controller delays volume provisioning, resizing and snapshotting; volumes already mounted keep working, because the mounting side is the node DaemonSet. These are the least urgent of the four, with the weakest chart support.
Both charts render no PodDisruptionBudget, and honour controller.affinity only when it carries
nodeSelectorTerms: a pod anti-affinity is schema-valid, then silently dropped.
The S3 chart renders no topology spread at all, so two replicas may land on one node and a drain can take both. Raise its count for process-level failover; it gives no node-level redundancy.
The NFS chart renders controller.topologySpreadConstraints through a vendored patch, and above one
replica the spread is required, not just prudent: its pods run on the host network and bind the
liveness health port there, so two pods on one node leave the second crash-looping on the bind and the
Deployment never fully ready. As with Kueue, the example’s labelSelector must be spelled out.
Also set strategyType: RollingUpdate: both default to Recreate, which takes every replica down
before the new one starts and gives up, at every upgrade, the failover the replica was added for.
The non-redundant topology
When the worker runs outside the cluster it manages but near it (image mode, !LoopbackKubeInside && LoopbackKubeNearby), its admission webhooks register against one node IP URL instead of a Service.
That URL is a single endpoint, so extra replicas receive no traffic; the source calls this “launch
multiple instances, only one takes working”. Keep that topology at one replica. Chart mode is
unaffected: its webhooks are always in-cluster and Service-backed.
Verify
NS=gpustack-system
# Every control-plane Deployment reports its full replica count Ready.
kubectl -n "$NS" get deploy
# The budgets exist and are satisfied (ALLOWED DISRUPTIONS ≥ 1).
kubectl -n "$NS" get pdb
# Replicas really are on distinct nodes — one line per pod, node in the second column.
kubectl -n "$NS" get pods -o wide \
--field-selector status.phase=Running \
-o custom-columns='POD:.metadata.name,NODE:.spec.nodeName'
# The NFD master elected a leader (only above one replica).
kubectl -n "$NS" logs deploy/node-feature-discovery-master | grep -i "leader"Then the failover: drain the node running the worker’s leader; a standby takes over:
kubectl drain <node> --ignore-daemonsets --delete-emptydir-data
kubectl -n "$NS" rollout status deploy/gpustack-operator-worker
kubectl get instancetypes # served throughout by the surviving replicas
kubectl uncordon <node>