Command Reference
The gpustack-operator binary provides long-running services and one-shot commands. The services are what a
deployment runs; the one-shots are what you run by hand on a node, and the first two sections choose
between them.
Contents
- Command summary
- Choosing a one-shot command
- Global flags
- worker
- worker-gateway
- device-manager serve
- model-manager
- device-manager detect
- device-manager monitor
- device-manager preflight
Command summary
| Command | Alias | Shape | Who runs it |
|---|---|---|---|
worker |
w |
service | the operator Deployment, one per cluster |
worker-gateway |
wg |
service | a sidecar beside worker |
device-manager serve |
dm serve |
service | the Device Manager DaemonSet, one per node |
model-manager |
mm |
service | the Model Manager DaemonSet, one per node |
device-manager detect |
dm detect |
one-shot | a person, on a node |
device-manager monitor |
dm monitor |
one-shot | a person, on a node |
device-manager preflight |
dm preflight |
one-shot | a person, on a node |
Choosing a one-shot command
The one-shot commands answer different questions; only preflight changes node state.
| Question | Command | Node effect |
|---|---|---|
| What accelerators are here, in full? | detect |
nothing — a pure read |
| What are they doing right now? | monitor |
nothing — a pure read |
| Can they actually be sliced, and managed while sliced? | preflight |
starts containers, stages libraries, asks a driver mode on |
detect and monitor are reads no flag can turn into an action. preflight is the one that acts:
it starts probe containers, stages the injected libraries onto the host, and, where a
manufacturer’s sharing mode is off, asks the driver to turn it on to see whether it takes, putting
it back either way. --dry-run is how it is made a read.
The probe containers it starts are themselves not privileged: they get the allocator’s injection and nothing more, because the isolation that injection establishes is the thing being measured.
Why
detectis not folded intopreflight—detectemits the whole ledger, typed as theDevicesCRD’s ownspec.groups, so it can be diffed against what the cluster recorded.preflight’s detection block is a verdict and a count, deliberately not a second ledger. A reader who wants both runs both.
Global flags
Every command accepts these. They come from klog and are omitted from the per-command tables below.
| Flag | Effect |
|---|---|
-v |
log verbosity, 0 unless you pass one. The DaemonSet runs at -v=2; -v=3 adds the per-decision lines, including why a container runtime did or did not resolve |
--vmodule |
per-file verbosity, pattern=N comma separated |
--logtostderr |
log to standard error rather than files (default true) |
--add_dir_header, --skip_headers, --log_backtrace_at |
log line shape |
--print-cmdline |
print the effective command line, including what was read from the environment |
--version |
the binary’s version |
worker
It is the control plane: it profiles node capacity, runs the controllers that materialize the Kueue scheduling chain, serves the aggregated extension APIs and the admission webhooks, and installs the applications a cluster needs.
Runs as the operator Deployment. Flags below are the ones that change behaviour rather than tune a
connection; see Settings
for what is configured through the Setting CR instead.
| Flag | Default | Effect |
|---|---|---|
--manufacturer |
all manufacturers | which manufacturers to detect |
--disable-applications |
none | applications to skip installing, or * for all. See Installation Modes
|
--secure-port |
31443 |
the HTTPS port |
--bind-address |
0.0.0.0 |
the address to serve on |
--cert-dir |
— | where tls.crt and tls.key are |
--disable-auths |
false |
skip authentication and authorization |
--kube-leader-election |
true |
leader election, for multi-instance deployments |
--kube-leader-election-id |
worker.gpustack.ai |
the election’s unique ID |
--kube-leader-lease, --kube-leader-renew-timeout |
15s, 10s |
how long leadership is held and renewed within |
--kube-conn-qps, --kube-conn-burst, --kube-conn-timeout |
200, 400, 5m |
loopback API client limits |
--informer-cache-resync-period |
1h |
informer resync period |
--gopool-worker-factor |
100 |
goroutine pool size, per CPU core |
--audit-log-path, --audit-policy-file, --audit-webhook-config-file |
— | request auditing |
--cors-allowed-origins |
all | allowed cross-domain origins, as anchored regular expressions |
--model-manager-service-account |
gpustack-operator-model-manager |
the ServiceAccount, in the operator namespace, the model-manager plugin runs as; only it may write a NodeModelStore’s status |
gpustack-operator worker --secure-port=31443 --disable-applications=kueueworker-gateway
Aggregates resources from upstream clusters and mirrors them for worker to read. Runs beside
worker; by default the two speak over a unix socket rather than a port.
| Flag | Default | Effect |
|---|---|---|
--bind-unix-path |
/var/lib/gpustack/gpustack-operator-worker-gateway.sock |
the socket to serve on. When set, --bind-address and --secure-port are ignored |
--worker-conn-mode |
gpustack-api |
how the worker is reached: gpustack-api or loopback |
--worker-conn-gpustack-api-port |
30080 |
the GPUStack server’s port in gpustack-api mode |
--worker-conn-gpustack-api-scheme |
derived from the port | http or https |
--worker-readiness-check-timeout |
30s |
one readiness check of a worker’s API services. Must not exceed --kube-conn-timeout, which otherwise bounds it |
--secure-port, --bind-address, --cert-dir |
— | serving, when not on a socket |
--kube-conn-*, --informer-cache-resync-period, --gopool-worker-factor |
as worker |
client and cache tuning |
gpustack-operator worker-gateway --worker-conn-mode=loopbackdevice-manager serve
It is the per-node agent: it detects accelerators, keeps the Devices ledger current, samples
utilization, registers with the kubelet as a device plugin and injects the vendor runtime into
allocated containers.
Runs as the Device Manager DaemonSet. The --no-* mode flags are how a node is restricted to a
subset of the allocation modes its hardware would otherwise offer.
| Flag | Default | Effect |
|---|---|---|
--manufacturer |
all manufacturers | which manufacturers to detect |
--no-sliced |
false |
do not create logically sliced devices |
--no-partitioned |
false |
do not create hardware-partitioned devices |
--no-shared |
false |
do not create shared devices |
--no-pci-check |
false |
skip the PCI check the detect pass makes |
--no-fast-failed |
false |
publish a first pass that found nothing rather than holding it back and detecting again. Only applies when --manufacturer names exactly one |
--monitor-period |
15s |
how often utilization is sampled; only the latest sample is kept |
--kube-socket |
/var/lib/kubelet/device-plugins/kubelet.sock |
the kubelet socket to register on |
--secure-port |
32443 |
the HTTPS port |
--bind-address, --cert-dir, --kube-conn-*, --informer-cache-resync-period, --gopool-worker-factor |
as worker |
serving, client and cache tuning |
gpustack-operator device-manager serve --manufacturer=nvidia --no-partitionedmodel-manager
The per-node model cache. A CSI node plugin serving inline ephemeral volumes of the driver
model.csi.gpustack.ai: it authorizes each mount against the Pod’s namespace, downloads and
verifies a Hugging Face artifact’s files once per node, and writes that node’s NodeModelStore
status. See Node Model Store
.
Runs as the Model Manager DaemonSet beside the node-driver-registrar sidecar. Its configuration
comes from its node’s NodeModelStore.spec, not from flags.
| Flag | Default | Effect |
|---|---|---|
--node-name |
— | the node it runs on, also its NodeModelStore’s name |
--kubelet-dir |
/var/lib/kubelet |
kubelet’s root directory, whose pods directory holds the mount targets |
--cache-root |
/var/lib/gpustack/models |
the node’s cache directory |
--csi-socket |
<kubelet-dir>/plugins/model.csi.gpustack.ai/csi.sock |
the Unix socket the CSI services listen on |
--secure-port |
32444 |
the HTTPS port: metrics, readiness and liveness |
--bind-address, --cert-dir, --kube-conn-*, --informer-cache-resync-period, --gopool-worker-factor |
as worker |
serving, client and cache tuning |
gpustack-operator model-manager --node-name="$(hostname)" --cache-root=/mnt/modelsdevice-manager detect
Runs one detect pass and prints the result. It is a pure read: nothing is written, nothing is started, and no flag changes that.
The output is the Devices CRD’s own spec.groups, so it can be compared directly against what the
cluster recorded; kubectl get devices <node> -o yaml beside it is a diff.
| Flag | Default | Effect |
|---|---|---|
--manufacturer |
all manufacturers | which manufacturers to detect |
--no-pci-check |
false |
skip the PCI check |
--no-fast-failed |
false |
publish a first pass that found nothing rather than detecting again |
docker run --rm --privileged -v /dev:/dev -v /sys:/sys \
gpustack/gpustack-operator:<tag> \
gpustack-operator device-manager detect --manufacturer=amd- id: radeon-rx-7800-xt
manufacturer: amd
name: AMD Radeon RX 7800 XT
memory: 16368
cores: 60
driverVersion: 6.16.13
runtimeVersion: "7.2"
computeCapability: gfx1101
family: GC 11.0.0
accelerators:
- id: GPU-5c88007d760374f3
index: 0
physicalIndexes:
- 1
- 128
topology:
pciBusId: "0000:04:00.0"
numaAffinity: "0"
cpuAffinity: 0-13
status:
unhealthy: false
logicalSliced:
coresPercentageOvercommit: true
count: 128A vendor’s user-space libraries have to be reachable. AMD has no container runtime that injects them, so a run without
/opt/rocmmounted finds no accelerator on a node that has two.preflight’s host cross-check is what tells that apart from a node with none; see below .
device-manager monitor
Takes one utilization sample and prints it. A pure read, like detect, and it accepts the same
flags.
docker run --rm --privileged -v /dev:/dev -v /sys:/sys \
gpustack/gpustack-operator:<tag> \
gpustack-operator device-manager monitor --manufacturer=amd- manufacturer: amd
timestamp: 2026-08-29T08:01:43.565038175+08:00
accelerators:
- id: GPU-5c88007d760374f3
memory: 16368
memoryUsage: 174
memoryUtilization: 1
coresUtilization: 0
temperature: 44
powerUsage: 8
unhealthy: falseThe same figures reach a cluster through the Instance metrics surfaces; see Instance Metrics .
device-manager preflight
It answers whether the allocation modes this node’s allocators offer actually work here. detect
cannot answer that question, because a capability can be declared and still not work.
This is the only command an operator runs by hand that acts on the node. See Preflight for the procedure (what to mount, what it starts and removes, how to read each row).
| Flag | Default | Effect |
|---|---|---|
--manufacturer |
all manufacturers | which manufacturers to ask about. Every one asked about is reported, including those nothing is read for |
--dry-run |
false |
print the container steps instead of taking them, and write nothing to the host. Each answer reports the complete invocation — and, because nothing was written, names the staging it would have done: the library tree, and whatever the manufacturer’s responder renders. Both have to exist before the printed command runs |
--probe-image |
derived per family | the image the probe containers run. Required for a family with no default, and the way to run in an air-gapped environment |
--host-root |
/host |
where the host’s root filesystem is mounted into this container |
--runtime |
resolved | the host runtime to drive, overriding what was resolved. One of docker, nerdctl, ctr; anything else is refused before the pass starts. An escape hatch: a runtime that the host does not carry drops every container step to being emitted |
--no-pci-check |
false |
skip the PCI check |
Row states and depths
Each row carries a state, a depth and, where a container ran, the evidence it ran on.
| State | Meaning | Exit code |
|---|---|---|
ok |
the capability works, at the depth stated | 0 |
not-declared |
the accelerator does not offer it, so there is nothing to check | 0 |
unavailable |
it is offered and this pass did not establish it | non-zero |
| Depth | Method |
|---|---|
declared |
read from the driver |
simulated |
the allocator produced the injection, and it was not run |
measured |
a container ran with that injection and its output was read |
A measured example
Taken from a run on two RX 7800 XT cards:
- accelerator: GPU-5c88007d760374f3
capability: sliced-runtime-loaded
state: ok
depth: measured
detail: the slicing runtime the injection mounts is loaded in the container
command: /usr/sbin/chroot /host docker run --rm --label gpustack.ai/preflight=true … ubuntu:24.04 …
evidence: |
gpustack-preflight-maps-begin
…
7205a176b000-7205a176d000 r--p 00000000 fc:00 13006695 /usr/local/vrocm/libvrocm.soThe verdict is the mapping line, not the shim’s own greeting: a library that failed to load prints a
message naming the same path, so only /proc/self/maps says it is actually there.
The same row under --dry-run carries the command and no evidence, and its depth is
simulated: the allocator’s intent, proven, with nothing run to confirm it took.
Not every row on that node reaches measured. sliced-quota-in-force on the same two cards is
simulated, because the vendor reader there reports nothing until an allocation has been charged and
the probe allocates nothing: the injection sets a cap and no container was observed under it. A row
is measured only where something was read back, never where the intent alone is known.
Runtime resolution
Resolved from the kubelet’s own CRI endpoint wherever the host names one, because that is what
starts a container on this node in production. A host that names none (a bare machine, or a
distribution keeping that configuration elsewhere) falls through to probing docker, nerdctl,
ctr.
The sources are ordered, and the first that names an endpoint answers, so a standard kubelet
configuration wins over a distribution drop-in rather than conflicting with it. The conflict that
does stop a run is within one source: a glob matching two directories that disagree, where
neither is the kubelet’s own by position. The steps are emitted instead, naming the endpoints in
conflict, and --runtime is how you settle it.
ctr is never driven: ctr run has no --device, and its only alternative grants every device, so
what it measured would not be the isolation the injection established. A node resolving to it gets
the step emitted for nerdctl instead, against the socket and namespace ctr resolved, or for
docker where nerdctl cannot pass this manufacturer’s vendor runtime either.
docker run --rm --privileged --network=host --runtime nvidia \
-e NVIDIA_VISIBLE_DEVICES=all \
-v /:/host -v /dev:/dev -v /sys:/sys \
gpustack/gpustack-operator:<tag> \
gpustack-operator device-manager preflight --manufacturer=nvidia --dry-run--network=host is not optional: the host’s CLI is entered through a chroot, which changes the root
and not the network namespace, so an image pull would read the host’s resolver and find nothing
answering.