Settings & Environment Variables
GPUStack Operator reads configuration from two places: the Setting resource, which you can change
at runtime, and environment variables, which the process reads once at startup.
First-deploy seeding vs. runtime changes. On first deploy,
settings.Initializecreates the delegated Secretgpustack-settingsin the system namespace and seeds every Setting fromGPUSTACK_<UPPER_SNAKE_NAME>, or the built-in default when that variable is unset. Later restarts only backfill missing Settings and never overwrite a stored value: aGPUSTACK_*variable is an initial seed only, so after the first deploy change theSettingresource, not the environment.
Contents
Online-adjustable settings
A fixed catalog of named values, served as the namespaced Setting aggregated API resource
(gpustack.ai/v1, short name set, category gpustack) and read at runtime with kubectl; the operator
picks a new value up on its next reconcile. Fixed means the resource serves
get,list,watch,apply,update,patch and no create/delete: you edit a Setting’s value, not its existence.
# List every setting and its current value
kubectl -n gpustack-system get settings # or: kubectl -n gpustack-system get set
# Inspect / change one
kubectl -n gpustack-system get setting node-management-manual -o yaml
kubectl -n gpustack-system edit setting node-management-manual
kubectl -n gpustack-system patch setting instance-type-derived-from-node --type merge -p '{"spec":{"value":"false"}}'| Setting | Env seed | Default | Effect |
|---|---|---|---|
container-registry |
GPUSTACK_CONTAINER_REGISTRY |
(blank) | Registry to pull images from, for all built-in applications’ deployments. |
container-namespace |
GPUSTACK_CONTAINER_NAMESPACE |
(blank) | Namespace to pull images from, for all built-in applications’ deployments. |
image-pull-secrets |
GPUSTACK_IMAGE_PULL_SECRETS |
(blank) | Image pull secret for pulling images, for all built-in applications’ deployments — the subcharts the chart or the worker installs. It does NOT reach a workload a controller renders from a custom resource: a KV cache backend
names its own spec.imagePullSecrets. |
image-pull-policy |
GPUSTACK_IMAGE_PULL_POLICY |
IfNotPresent |
Image pull policy for all built-in applications’ deployments, on the same boundary as the secret above: a controller-rendered workload carries its own. |
instance-general-resources-overcommit |
GPUSTACK_INSTANCE_GENERAL_RESOURCES_OVERCOMMIT |
true |
Overcommit an Instance’s general resources: when enabled, a general unit requests 800m CPU / 128Mi RAM and one-eighth local storage, an accelerated unit 100m CPU / 128Mi RAM and one-eighth local storage — so e.g. a 1C/4Gi + 128Gi type requesting 2 accelerators and 64Gi storage resolves to 200m/256Mi + 8Gi. A CPU limit whose whole cores cost more than the limit itself (e.g. 100m, or 1500m whose half core rounds up) requests the limit instead — a request may never exceed its limit. |
instance-ssh-server-image |
GPUSTACK_INSTANCE_SSH_SERVER_IMAGE |
gpustack/ssh-server:v1.3.0 |
Image of the SSH server used when deploying Instances. |
kv-cache-backend-image |
GPUSTACK_KV_CACHE_BACKEND_IMAGE |
gpustack/mirrored-mooncake:0.3.13.post1-cpu |
Image every role of a KV cache backend
runs when the object does not name one itself. The default is this project’s own build — a CPU build carrying TCP and EFA over DRAM, and the only image that can run leader.electionBackend
. The per-vendor variant builds, their tags and which one a VRAM group needs are under The project’s own build variants
. It does not fit every backend: a backend on another transport or runtime names its own spec.image, which always wins over this. Clearing this Setting restores the admission refusal for a backend naming no image anywhere. Why one default cannot fit every backend, and what having one costs, is under The image
. |
model-deployment-kv-cache-dtype-owned |
GPUSTACK_MODEL_DEPLOYMENT_KV_CACHE_DTYPE_OWNED |
true |
Hand every KVCachePoolBinding’s spec.domain.dtype to the engines attached through it as --kv-cache-dtype, verbatim, and refuse a pool-attached role or injected Pod naming that flag itself, and a new Binding declaring auto — see The dtype is handed to the engine
. It is the escape from a Binding whose spelling the engine rejects, which makes every new Pod fail argument parsing: false renders and refuses nothing, as before the flag was owned. Flipping it recreates every pool-attached replica at its next reconcile. The recovery is under Upgrading to an enforced Binding dtype
. |
model-deployment-router-image |
GPUSTACK_MODEL_DEPLOYMENT_ROUTER_IMAGE |
gpustack/llm-router:v0.1.0 |
Image a managed router runs when the ModelDeployment does not name one in spec.router.image. It carries every router this operator supports, and which binary runs is decided by the rendered command rather than by the image — spec.router.name picks the binary, and this setting only says where the binaries come from. Its default is safe for every cluster at once in a way kv-cache-backend-image’s is not: it runs no model, so it links no accelerator runtime and cannot be paired with the wrong one. The default is this project’s own build rather than an upstream tag, because the three routers are compiled from three separate sources and one of them carries a patch this repository ships, so no upstream image holds them together. An image named on the object is used verbatim and is never redirected to the cluster mirror, because it is the user’s own reference. |
model-deployment-router-proxy-image |
GPUSTACK_MODEL_DEPLOYMENT_ROUTER_PROXY_IMAGE |
gpustack/mirrored-envoy:distroless-v1.33.2 |
Proxy fronting a managed router’s endpoint picker. It has no field on the API to override it: this operator renders the proxy’s configuration against one proxy’s configuration schema, so swapping the binary would mean swapping that configuration too. The setting exists for registry redirection and for pinning a release back, not for running a different proxy. |
model-deployment-routing-sidecar-image |
GPUSTACK_MODEL_DEPLOYMENT_ROUTING_SIDECAR_IMAGE |
gpustack/mirrored-llm-d-router-disagg-sidecar:v0.10.0 |
Sidecar a decoder runs to accept a remote prefill handoff. Same terms as the proxy above: this operator renders its arguments, so the setting is for redirection and pinning rather than for a different implementation. |
model-deployment-tcp-tw-reuse |
GPUSTACK_MODEL_DEPLOYMENT_TCP_TW_REUSE |
false |
Render net.ipv4.tcp_tw_reuse=1 on the prefill half of every SGLang prefill/decode pair, which otherwise runs out of local ports
under sustained load. Allow the sysctl on the kubelet of every node that can run such a Pod before turning it on: without that the Pod fails with SysctlForbidden, and so does every replacement. The steps, which Pods it reaches and how to confirm the refusal are under SGLang TIME-WAIT port reuse
. |
model-prefetch-warmup-image |
GPUSTACK_MODEL_PREFETCH_WARMUP_IMAGE |
gpustack/mirrored-python:3.12-alpine |
Image the ModelPrefetch
warm-up pod runs on each target node. The pod mounts the artifact’s volume, verifies every file reads back and exits, so the image needs only a shell, find and sha256sum — it never runs a model. It is a setting because the one thing it must guarantee is that the node can pull it, which only the cluster’s own registry arrangement decides; an air-gapped cluster points it at its mirror. |
model-artifact-huggingface-endpoint |
GPUSTACK_MODEL_ARTIFACT_HUGGINGFACE_ENDPOINT |
https://huggingface.co |
The Hugging Face Hub a ModelArtifact
resolves and revalidates against, and the HF_ENDPOINT an engine downloading the weights is given; the node plugin downloads from it too. What changing it does is under Engine delivery
. |
model-artifact-modelscope-endpoint |
GPUSTACK_MODEL_ARTIFACT_MODELSCOPE_ENDPOINT |
https://www.modelscope.cn |
The ModelScope hub a ModelArtifact
resolves and revalidates against and the node plugin downloads from; an engine downloading the weights gets its host as MODELSCOPE_DOMAIN — the runners’ SDK takes a bare host. What changing it does is under Engine delivery
. |
model-artifact-https-proxy |
GPUSTACK_MODEL_ARTIFACT_HTTPS_PROXY |
(blank) | HTTPS proxy for resolution, revalidation and the node plugin’s downloads, and the HTTPS_PROXY given to an engine downloading the weights, where a role’s own value wins. Blank keeps the worker’s own environment and renders nothing. Only an http or https URL without credentials is accepted, because the value reaches tenant Pods; a value from the environment that is not refuses the worker’s start. |
model-artifact-no-proxy |
GPUSTACK_MODEL_ARTIFACT_NO_PROXY |
(blank) | Comma-separated hosts that bypass model-artifact-https-proxy, for the node plugin too, given to engines as NO_PROXY on the same terms. |
model-artifact-ca-bundle |
GPUSTACK_MODEL_ARTIFACT_CA_BUNDLE |
(blank) | Name of a ConfigMap in the worker’s namespace whose ca.crt the resolution and the node plugin trust beside the system pool. It is not given to engine Pods: they run in tenant namespaces, which cannot mount it. |
model-artifact-revalidate-interval |
GPUSTACK_MODEL_ARTIFACT_REVALIDATE_INTERVAL |
24h |
How often a resolved hub artifact’s access is checked again, at least 1m; a Secret change checks it at once. What a refusal does is under Resolution and revalidation
. |
model-artifact-delivery-mode |
GPUSTACK_MODEL_ARTIFACT_DELIVERY_MODE |
Engine |
How a hub artifact’s weights reach a ModelDeployment: Engine, the engine downloads them, or Node, the node’s model-manager plugin
mounts a verified copy. The chart seeds Node when it deploys the plugin, and a seed never overrides a stored value. Node is refused while the CSIDriver model.csi.gpustack.ai does not exist. Changing it rolls every deployment using a hub artifact once; see Switch delivery
. |
model-store-high-watermark |
GPUSTACK_MODEL_STORE_HIGH_WATERMARK |
80 |
The node cache filesystem’s usage percent above which the plugin removes content no Pod mounts; above the low watermark, at most 95. On a filesystem shared with kubelet the plugin caps it below kubelet’s thresholds (the capacity rule
). |
model-store-low-watermark |
GPUSTACK_MODEL_STORE_LOW_WATERMARK |
70 |
The usage percent a collection removes down to; at least 1, below the high watermark. |
model-store-download-concurrency |
GPUSTACK_MODEL_STORE_DOWNLOAD_CONCURRENCY |
8 |
Concurrent HTTP requests the plugin makes per node, across every download; 1 to 64. |
model-store-peer-sync |
GPUSTACK_MODEL_STORE_PEER_SYNC |
true |
Whether a node’s plugin may pull a model’s content from the plugins of other nodes that hold it published, instead of the hub. |
model-store-download-bandwidth |
GPUSTACK_MODEL_STORE_DOWNLOAD_BANDWIDTH |
0 |
Per-node download rate limit, a quantity of bytes per second (200Mi); 0 is unlimited. |
instance-access-static-address |
GPUSTACK_INSTANCE_ACCESS_STATIC_ADDRESS |
(blank) | Static access address for all Instances; when unset, the access address is generated from host IPs. |
instance-access-wildcard-dns |
GPUSTACK_INSTANCE_ACCESS_WILDCARD_DNS |
(blank) | Wildcard DNS for all Instances (e.g. traefik.me), used to build a per-Instance domain <instance-host-ip>.<wildcard-dns>. Only effective when instance-access-static-address is not set. |
instance-privileged-allowed |
GPUSTACK_INSTANCE_PRIVILEGED_ALLOWED |
false |
Whether an Instance may request privileged mode (spec.privileged), which escapes the container boundary and exposes the node’s devices and kernel surface. Enforced by the Instance admission webhook whenever an Instance takes privileged mode — at creation, or through a later change, including one made while it is stopped. An Instance that already runs privileged keeps it: with this off it stays updatable, editable while stopped, and restartable. Like every setting here, a change takes up to 30s to reach the webhook, so an Instance created in that window is judged against the previous value. |
instance-host-path-volume-allowed |
GPUSTACK_INSTANCE_HOST_PATH_VOLUME_ALLOWED |
false |
Whether an Instance may mount a hostPath volume (spec.additionalVolumes[*].hostPath), which reaches the node’s filesystem. Kept separate from instance-privileged-allowed because it grants strictly less — the filesystem, but not the node’s devices or kernel — so an admin can allow node-path mounts without allowing a container escape. Enforced whenever an Instance takes a hostPath mount, on the same terms; a mount it already has is matched by value, so reordering the list is not a new grant while repointing an entry at a different node path is. |
instance-persistent-volume-placement |
GPUSTACK_INSTANCE_PERSISTENT_VOLUME_PLACEMENT |
true |
Whether an Instance’s persistent claims (spec.volume.persistent, spec.additionalVolumes[*].persistent) decide where its Pod runs, by the claim placement rules
: a bound claim adds its PV’s required node affinity, and a claim nothing binds holds the Pod back with the reason in status.phaseMessage. false is the escape path: the Pod is built from the claim names alone, as before, and Kueue’s topology-aware scheduling may then assign a node the volume cannot reach, leaving the Pod Pending while its Workload holds quota. Read when a Pod is created, so a change never moves a running Instance; stop and start one to apply it. |
workload-fit-affinity |
GPUSTACK_WORKLOAD_FIT_AFFINITY |
true |
Whether the Workload webhook pins a logical slice, or a shared request of two or more accelerators, on an operator queue to nodes whose fit labels
show that one accelerator can still host it. When true, Kueue’s topology-aware scheduling skips a fragmented node; when false, Workloads pass unchanged while the fit labels are still published, so the pin can be switched off on its own. Read on every webhook call, so a change takes up to 30s to reach it, and it applies to Workloads created or updated after that. |
node-management-manual |
GPUSTACK_NODE_MANAGEMENT_MANUAL |
false |
Skip auto-managing nodes. When false, the operator auto-onboards discovered nodes by injecting the gpustack.ai/managed=true label; when true, an administrator must opt nodes in manually. Read per-reconcile. |
instance-type-mixed-on-node |
GPUSTACK_INSTANCE_TYPE_MIXED_ON_NODE |
true |
Whether one node may surface both a GPU and a CPU-only InstanceType. When true, a node is summarized into every type it can serve; when false, a node with accelerators yields only a GPU InstanceType and a CPU-only node only a general one. Read per-reconcile. |
instance-type-derived-from-node |
GPUSTACK_INSTANCE_TYPE_DERIVED_FROM_NODE |
true |
Whether the operator auto-derives the InstanceType (and its backing ClusterQueue) from node hardware. When true, the NodeFlavorReconciler authors the derived InstanceType (create-only); when false, it only aligns the ResourceFlavor and the administrator defines the InstanceType via the API, and owns feasibility with it — see Authoring the InstanceType
. Read per-reconcile. |
instance-type-drain-when-no-flavors |
GPUSTACK_INSTANCE_TYPE_DRAIN_WHEN_NO_FLAVORS |
true |
Whether a ClusterQueue whose pool has lost all its ResourceFlavors is drained (HoldAndDrain, so Kueue evicts admitted workloads) before its resource groups are emptied. When true, the queue is drained first; when false, the operator waits for the reservations to clear on their own, then empties. Either way the groups are emptied only once every reservation is zero, so Kueue’s counters never go negative. Read per-reconcile. |
instance-type-aware-cpu-manufacturer |
GPUSTACK_INSTANCE_TYPE_AWARE_CPU_MANUFACTURER |
false |
Whether the derived ClusterQueue/InstanceType/InstanceTypeFlavor aggregation splits by CPU manufacturer. When false, non-accelerated flavors collapse into one generic pool per os/arch and accelerated flavors pool per accelerator (CPU ignored); when true, every pool splits by the CPU key (gpustack--${gKey}-… / gpustack--${gKey}--${aKey}-…) and the InstanceType records the raw CPU detail. The ResourceFlavors themselves are unaffected — they always carry the CPU key, so a flip only re-groups the aggregation layer. Read per-reconcile. |
These settings (node-management-manual, instance-type-mixed-on-node, instance-type-derived-from-node,
instance-type-drain-when-no-flavors, instance-type-aware-cpu-manufacturer) are read per-reconcile
(ShouldValueBool(ctx)): flipping one re-converges the scheduling chain on the next reconcile, with no
restart.
Authoring the InstanceType
With instance-type-derived-from-node=false the operator authors no InstanceType, so the
administrator owns both halves of a pool: which InstanceTypes exist, and whether what their queue
admits can actually be placed. Flipping the setting off does not retire the types the operator
already authored; the derived marker is provenance, and nothing auto-removes a type.
No ClusterQueue gains a reference to the node-devices feasibility gate in this mode. That
AdmissionCheck
has one
writer, NodeQueueReconciler
,
which reads this setting as a cluster-wide switch, not as “was this queue derived”.
So with the switch on, every accelerated queue carries the reference once the check reports Active,
whoever authored its InstanceType; with it off none is added, and a queue that already carries one
drops it the next time its resource groups are refilled, not at the moment of the flip.
Why it is not attached in this mode anyway — the gate changes what a queue admits, and in the mode where the administrator authors every InstanceType that is theirs to decide.
The failure shape to expect: without that gate, a Workload is admitted on quota alone and its Pods
then sit Pending, which reads as a full cluster while the usual truth is a request no node can
place, because quota is a scalar total and cannot see per-accelerator fragmentation. When the Pods
of an admitted Workload stay Pending, compare the request against the per-accelerator ledger
(kubectl get devices <node> -o yaml) before adding capacity.
The joint-admission check is attached in both modes. Every queue backing an InstanceType
references gpustack-model-deployment-joint once it reports Active, whoever authored the type and
whatever this setting says.
It answers Ready at once for every Workload that is not a replica of a multi-role ModelDeployment,
so the only thing it changes in this mode is that such a deployment is admitted as a set; see
Prefill and decode
, and
Status
for a set that stays short of quota.
Write the InstanceType against the InstanceTypeFlavor catalog. It is the read-only,
os/arch-agnostic view of the pools that actually exist (one entry per grouping the settings above
produce), and its spec carries the identity fields an InstanceType is built from:
acceleratorGroup, generalGroup, acceleratable, manufacturer, product, family, memory,
cores.
kubectl get instancetypeflavors # short name: instypeflavor
kubectl get instancetypeflavor gpustack--nvidia-a10g -o yamlAn InstanceType whose identity matches no entry there backs a ClusterQueue that selects no
ResourceFlavor, so the queue is left with empty resource groups and admits nothing.
SGLang TIME-WAIT port reuse
model-deployment-tcp-tw-reuse renders net.ipv4.tcp_tw_reuse=1 into the Pod
securityContext.sysctls of the prefill half of every SGLang prefill/decode pair, and of no other
Pod. The sysctl lets the kernel reuse ports held by TIME-WAIT sockets for new outgoing
connections, in that Pod’s own network namespace only.
The prefill half opens a new TCP connection for every transfer, so over TCP, with or without a
store, its TIME-WAIT sockets use up its local
ports
under sustained load.
It does not reach the decode half, which accepts transfers rather than opening them; a vLLM role,
which keeps its transfer connections open; or the members of a KV cache backend
.
A role that replaced its command gets nothing either, since the operator did not build what runs
there. Whether an SGLang server with a store runs out of ports is not measured, and the setting
does not reach it.
- Allow the sysctl on the kubelet of every node that can run such a Pod. It is not on the Kubernetes safe list , so a kubelet refuses it until told otherwise. Add it to the kubelet configuration and restart the kubelet:
# KubeletConfiguration
allowedUnsafeSysctls:
- net.ipv4.tcp_tw_reuseThe command-line form is --allowed-unsafe-sysctls=net.ipv4.tcp_tw_reuse. It is one list, so keep
any name already on it.
On a managed cluster, set it where the node pool’s kubelet configuration comes from, so a node added or replaced later keeps it; an edit made on a running node leaves with that node. On each of these it takes effect on newly created nodes:
- EKS on Amazon Linux 2023 — the launch template’s
nodeadmNodeConfig, underspec.kubelet.configas above or as the flag underspec.kubelet.flags; see thenodeadmAPI . On Amazon Linux 2, pass the flag through the bootstrap script’s--kubelet-extra-args. - AKS —
allowedUnsafeSysctlsin the kubelet configuration a node pool is created with; see custom node configuration . - GKE —
kubeletConfig.allowedUnsafeSysctlsin the node system configuration of a node pool; see node system configuration . - A platform that exposes no kubelet configuration cannot run this setting; leave it off there.
- Turn the setting on.
kubectl -n gpustack-system patch setting model-deployment-tcp-tw-reuse --type merge -p '{"spec":{"value":"true"}}'Flipping it recreates the prefill replicas it reaches, in either direction, because the sysctl is part of the Pod spec their fingerprint covers. Each deployment picks the value up on its next reconcile, and a setting change does not wake one. To apply it at once, delete one of the deployment’s prefill Pods: the reconcile that replaces it renders the new value, and the rollout replaces the other prefill replicas. A pair that has already locked up recovers the same way.
If the kubelet was not changed, the prefill replica never starts. The node’s kubelet refuses the
Pod, which ends in phase Failed with reason SysctlForbidden; the operator deletes it and logs
removing replica that failed with that reason, and the replacement is refused the same way, over
and over. The events outlive the Pods, and their message names the sysctl:
kubectl -n <namespace> get events --field-selector reason=SysctlForbidden
kubectl -n <namespace> get pods -w \
-l app.kubernetes.io/instance=<deployment>,app.kubernetes.io/component=<prefill-role>The second command shows the prefill Pods cycling through SysctlForbidden. Allow the sysctl on
those nodes or turn the setting off; either ends the loop at the next replacement.
A namespace enforcing the baseline or restricted Pod Security Standard refuses the Pod at
creation instead, since the sysctl is not on that standard’s allowed list either: no Pod appears,
and the operator logs the refusal.
Deploy-time environment variables
Read once at process startup, so changing one means restarting or redeploying the affected component.
The Worker (WK) copies every GPUSTACK_-prefixed variable from its own Pod spec onto the Device Manager
(DM) DaemonSets, so setting one on the Worker Deployment reaches the DMs automatically.
General variables
| Variable | Default | Component | Effect |
|---|---|---|---|
GPUSTACK_DATA_DIR |
/var/lib/gpustack |
all | Root directory for data storage. |
GPUSTACK_CONF_DIR |
/etc/gpustack |
all | Root directory for configuration and metadata, e.g. bundled Helm charts. |
GPUSTACK_PCI_CLASS_PREFIXES |
02,03,0b,12 |
DM | Comma-separated PCI class prefixes treated as display/accelerator devices (see the PCI class registry
). Applied to the DM’s local sysfs PCI scan, and to nothing else. The same list appears twice more, and neither reads this variable: the chart value node-feature-discovery.worker.config.sources.pci.deviceClassWhitelist decides which classes NFD labels, and pkg/nodefeature is what the gpustack-cpu-info NodeFeatureRule matches (a Go test holds those two equal). Change one and change all three. |
GPUSTACK_DEVICES_GROUP_ID_WITH_MEMORY |
false |
DM | When true, the devices group ID gains a memory-size suffix (e.g. nvidia-tesla-t4-16g instead of nvidia-tesla-t4), so same-model devices with different VRAM sizes form distinct groups. |
GPUSTACK_DEVICE_PLUGIN_SLICED_ALLOCATE_GATE |
true |
DM | Whether Allocate refuses a logical slice its accelerator cannot hold: one with no free slot, or without the units the slice needs. The Pod fails with UnexpectedAdmissionError and its controller recreates it. false allocates it anyway and clamps the ledger at zero, as before the refusal existed; see Disabling an allocator safeguard
. |
GPUSTACK_DEVICE_PLUGIN_IDENTIFY_BY_KUBELET |
true |
DM | Whether Allocate asks the kubelet’s pod-resources API which container it serves. false identifies it by the pending-Pod heuristic alone, as before the lookup existed; a node whose kubelet cannot be reached falls back to that heuristic by itself. See Disabling an allocator safeguard
. |
GPUSTACK_LEDGER_RELEASE_TERMINATED_PODS |
true |
WK, DM | Whether the accelerator ledger stops charging a Pod in phase Succeeded or Failed for the exclusive cards, shared cards and logical slices kubelet has taken back from it, so a finished Job that nobody deletes no longer keeps its card from the fit labels, the InstanceType views and the node-devices check. A hardware partition, and a MetaX or Cambricon slice, stay charged until the Pod object is gone, because only then are they destroyed. false charges every finished Pod until its object is gone, as before, and makes the sliced Allocate gate charge it too, which it did not do before; set it through the chart value ledgerReleaseTerminatedPods, which renders it on the worker and every device-manager so the two sides agree. |
Disabling an allocator safeguard
Two device-plugin safeguards default on, and each has a switch for the case where it misjudges a
node. Set it through the chart, which renders deviceManager.env onto every device-manager DaemonSet:
helm upgrade gpustack-operator <chart> --namespace gpustack-system --reset-then-reuse-values \
--set-string deviceManager.env.GPUSTACK_DEVICE_PLUGIN_SLICED_ALLOCATE_GATE=falseGPUSTACK_DEVICE_PLUGIN_SLICED_ALLOCATE_GATE=falseis for a slice refused although its accelerator has room: the Pod keeps coming backUnexpectedAdmissionErrornaming that accelerator. With it off, the slice is allocated and the ledger clamps that accelerator’sRemainingat zero.GPUSTACK_DEVICE_PLUGIN_IDENTIFY_BY_KUBELET=falseis for allocations recorded on the wrong Pod while the kubelet is reachable. With it off, the device manager picks the oldest pending Pod the request could be for, which can record two Pods bound together on each other.
The change rolls the device-manager Pods, one node at a time. Running workloads are not touched: a device-manager restart does not reach a started container, and its allocation records live on the Pods, where the restarted process reads them. A Pod admitted while its node’s device manager is restarting can fail admission, as across any device-manager restart, and its controller recreates it.
To switch the safeguard back on, remove the value. Set it to null with --set, not --set-string,
which would store the string "null":
helm upgrade gpustack-operator <chart> --namespace gpustack-system --reset-then-reuse-values \
--set deviceManager.env.GPUSTACK_DEVICE_PLUGIN_SLICED_ALLOCATE_GATE=nullPer-manufacturer overrides
Override patterns expand for every known manufacturer (amd, ascend, cambricon, hygon,
iluvatar, metax, mthreads, nvidia, thead); both the WK and the DM read them, so the propagation
above keeps the two sides consistent.
Installed by the chart, set the global.manufacturers row, not these variables. The overrides
here are fields of that row (pciVendorID, resourceName, runtimeName, partitionKind), which the
chart fans out as the matching variable to the worker and the device-managers, along with what a variable
cannot reach: the DaemonSet node selectors, the RuntimeClasses it creates, Kueue’s credits mapping. The
variable alone leaves those stale; the next render overwrites it.
GPUSTACK_${MANUFACTURER}_PCI_VENDOR_ID— the PCI vendor ID used for NFD node selection and device scanning. Accepts${vendor}or${class}_${vendor}.GPUSTACK_${MANUFACTURER}_ACCELERATABLE_RESOURCE_NAME— the extended resource name the scheduling chain allocates against.GPUSTACK_${MANUFACTURER}_ACCELERATABLE_RUNTIME_NAME— the container runtime class name used for accelerated workloads. A RuntimeClass of that name is attached only when one exists in the cluster; the chart creates one only where theglobal.manufacturersrow sets aruntimeInjectsDriverorruntimeInjectsDevicesfact (seeglobal.manufacturers) — never for a manufacturer setting neither, whose default below holds only if something else created the class.
Defaults:
| Manufacturer | PCI vendor ID | Resource name | Runtime name |
|---|---|---|---|
amd |
1002 |
amd.com/gpu |
amd |
ascend |
19e5 |
huawei.com/npu |
ascend |
cambricon |
cabc |
cambricon.com/mlu |
cambricon |
hygon |
1d94 |
hygon.com/dcu |
hygon |
iluvatar |
1e3e |
iluvatar.com/gpu |
iluvatar |
metax |
9999 |
metax-tech.com/gpu |
metax |
mthreads |
1ed5 |
mthreads.com/gpu |
mthreads |
nvidia |
10de |
nvidia.com/gpu |
nvidia |
thead |
1ded |
alibabacloud.com/ppu |
— (none) |
T-Head has no default runtime name, but
GPUSTACK_THEAD_ACCELERATABLE_RUNTIME_NAMEis still honored and can supply one.
A fourth override applies only to a manufacturer that has hardware partitioning at all:
GPUSTACK_${MANUFACTURER}_PARTITION_KIND— the manufacturer’s own name for hardware partitioning, which becomes the segment prefix of its per-profile resource key.
nvidia is the only one with a default: mig, giving nvidia.com/gpu.partitioned.mig-${profile}.
Elsewhere the variable does nothing: without hardware partitioning there is no .partitioned family, so
no key segment to rename.
A fifth override is per manufacturer, and today only NVIDIA reads it:
GPUSTACK_${MANUFACTURER}_DEVICE_INJECTION_STRATEGY— which channel that manufacturer’s allocator uses to make a granted accelerator reach a container:envvar(default, today’s*_VISIBLE_DEVICESbehavior, byte-identical),cdi-annotations(the granted accelerator is requested as acdi.k8s.io/*device-plugin annotation), orauto(opt-in detection, falling back toenvvarwith the reason logged whenever it cannot confirm CDI is safe on that node).
A manufacturer whose allocator injects /dev nodes itself, or that ships no CDI generator, has nothing
for the CDI channel to resolve, so the variable does nothing there. The partitioned (MIG) family and
partition-backed visibility always use envvar, whatever the strategy names, since a MIG instance is
materialized at Allocate time and no pre-generated specification names it.
Prefer auto over naming the CDI channel yourself. The CDI channel needs a container engine that
resolves CDI requests (containerd 2.x does; containerd 1.7 only with enable_cdi = true). Ask for it on
an engine that does not and the request is simply ignored: no variable is set either, and the container
starts with no accelerator and no error, which is the failure this setting exists to remove. auto reads
the engine first and keeps the variable when the answer is no.
Two independent questions decide which value a runtime wants: does the engine resolve CDI
requests at all, and can auto tell that it does? auto reads both answers out of containerd’s
config.toml, so a runtime that keeps them anywhere else is invisible to it. That file is where the
detection looks; it is not something the channel itself needs.
| Container runtime | Resolves CDI | Detectable by auto |
Recommended |
|---|---|---|---|
| containerd 2.x (configuration version 3) | Yes, always | Yes, from the version | auto |
containerd 1.7 with enable_cdi = true |
Yes | Yes, from the key | auto |
containerd 1.7, enable_cdi unset |
No | Yes | auto, which keeps the variable. Naming cdi-annotations here is the silent no-accelerator case above |
| CRI-O | Yes, natively | No — it has no config.toml |
cdi-annotations |
cdi-annotations is the answer on exactly one row: a runtime that does resolve CDI requests but that
auto has no way to see. Everywhere else auto reaches the same channel by itself, or correctly
declines to; naming the channel where the engine ignores it is how a container ends up running
with nothing.
Where to set it: the name embeds the manufacturer in upper case, so it is not a literal you will
find spelled out anywhere. Pass it through the chart’s deviceManager.env, which lands on every
device-manager DaemonSet:
helm upgrade ... --set deviceManager.env.GPUSTACK_NVIDIA_DEVICE_INJECTION_STRATEGY=autoIt is read once, when that manufacturer’s allocator is constructed, so a change takes effect only when
the DaemonSet restarts. A value that is not one of the supported choices is reported and the node keeps envvar:
refusing to start the allocator over it would take the node’s accelerators with it.
auto also keeps the variable when the engine’s own default runtime is already the vendor runtime.
That is the usual shape on a distribution that ships the GPU toolkit for you. There every Pod runs under
that runtime whether it asks to or not, so the variable works and a CDI request would only add a second
injection path. auto doing nothing on such a node is the correct answer, not a misconfiguration.
Which channel a node settled on, and why, is logged once per answer, but at the allocator’s own verbosity, above what the DaemonSet ships with. Runtime log verbosity raises it on a running Pod.
One node-level variable feeds the same decision:
GPUSTACK_CONTAINERD_CONFIG_DIR— the directory holding the container engine’sconfig.toml, from the chart’sdeviceManager.containerdConfigDir(default/etc/containerd).autoreads it to learn whether the engine resolves CDI requests and which runtime handler a Pod naming noruntimeClassNameruns under; unreadable keeps the node onenvvar.
Manufacturer toolkit paths
The DM device bindings locate manufacturer libraries through conventional toolkit-home variables, each falling back to the listed default when unset.
| Variable | Default | Manufacturer | Effect |
|---|---|---|---|
ROCM_HOME, then ROCM_PATH |
/opt/rocm |
AMD | ROCm root, searched for librocm_smi64.so / libamd_smi.so / libhsa-runtime64.so. |
ROCM_SMI_LIB_PATH |
— | AMD | Extra directory searched for librocm_smi64.so before the ROCm root. |
AMD_SMI_LIB_PATH |
— | AMD | Extra directory searched for libamd_smi.so before the ROCm root. |
CANN_HOME |
/usr/local/Ascend |
Ascend | Driver root, searched for libdcmi.so. |
ASCEND_TOOLKIT_HOME |
/usr/local/Ascend/cann, falling back to /usr/local/Ascend/ascend-toolkit/latest/runtime |
Ascend | CANN toolkit root used by the Ascend detector. |
NEUWARE_HOME |
/usr/local/neuware |
Cambricon | Neuware root, searched for libcndev.so. |
PPU_HOME |
/usr/local/PPU_SDK |
T-Head | PPU SDK root, searched for libhgml.so. |
COREX_HOME |
/usr/local/corex |
Iluvatar | CoreX root, searched for libixml.so. |
MACA_HOME |
/opt/maca |
MetaX | MACA root, searched for libmxsml.so. |
LD_LIBRARY_PATH |
— | all | Standard library search path, consulted as an additional source of candidate library directories. |
Kubernetes-injected variables
Populated by the Pod specs GPUStack Operator renders (Downward API or Service environment); not user-facing knobs, listed for completeness.
| Variable | Default | Effect |
|---|---|---|
KUBERNETES_NODE_NAME |
— (required) | Name of the node the Pod runs on; the DM uses it to name its NodeFeature/Devices objects. |
KUBERNETES_POD_NAME |
— | The WK’s own Pod name, used to read back its container spec (image, pull policy, GPUSTACK_* env) for rendering the DM DaemonSets. |
KUBERNETES_POD_NAMESPACE |
gpustack-system |
System namespace where managed resources live. |
KUBERNETES_POD_IP |
— | Overrides the auto-detected primary host IP in topology discovery. |
KUBERNETES_SERVICE_NAME |
gpustack-operator-worker |
Service name used for system routing. |
KUBERNETES_SERVICE_HOST |
— | Standard in-cluster marker; its presence tells the embedded runner it is inside a cluster. |
Proxy and internal flags
| Variable | Default | Effect |
|---|---|---|
ALL_PROXY / HTTP_PROXY / HTTPS_PROXY / NO_PROXY |
— | Standard proxy settings, passed through to the embedded Kubernetes installer. |
NO_PROXY / no_proxy |
— | Also parsed (hosts, IPs, CIDRs) to bypass the proxy on direct HTTP calls. |
_RUNNING_INSIDE_CONTAINER_ |
false |
Internal marker baked into the container image; switches data/conf paths to their absolute in-container locations. Not intended to be set by users. |