Accelerator Requests
The Pod webhooks check these requests at creation and reject invalid ones before a container starts.
They select Pods with the kueue.x-k8s.io/queue-name label. Without that label, a Pod can still
request device-plugin resources, but these checks and resource-unit calculations do not run.
Contents
- Two families, two accelerator populations
- The resource keys
- Worked example per family
- The request rules
- The RDMA keys, beside the accelerator families
- Co-locating an accelerator and an RDMA interface
- Requesting through the Instance API
- Pre-release breaks
- Limitations
Two families, two accelerator populations
An accelerator is shared in one of two physically incompatible ways, and GPUStack names them apart:
| Term | Definition | Isolation | Example |
|---|---|---|---|
Logical slicing (.sliced*) |
software slicing of a whole accelerator | the manufacturer’s own sharing facility caps compute and VRAM per container — which facility, per manufacturer | 50 % of an A10G |
Physical partitioning (.partitioned*) |
hardware partitioning of an accelerator put into a partitioning mode | the hardware itself; the operator materializes the instance | an NVIDIA MIG 3g.40gb, or a partition of a T-Head PPU or a Hygon DCU |
The two never apply to the same accelerator: one in a partitioning mode advertises only the partition
family, an unpartitioned one only the whole-accelerator, shared and logical-slice families. Hence the
four separate InstanceType views (EX / SH / SL / PT) — each accelerator feeds exactly one.
<kind> is the manufacturer’s own word for hardware partitioning, and manufacturers that offer it call it
mig: nvidia.com/gpu.partitioned.mig-3g.40gb, alibabacloud.com/ppu.partitioned.mig-<profile>,
hygon.com/dcu.partitioned.mig-2g.15gb. A manufacturer with no hardware partitioning has no kind, and
no .partitioned* keys at all.
The resource keys
<base> is the manufacturer’s device resource (nvidia.com/gpu, huawei.com/npu, … — see
Accelerator support
).
| Key | Served by | Eligible accelerators | Request value | Node value |
|---|---|---|---|---|
<base> |
device plugin (Exclusive) | unpartitioned only | accelerator count | Σ healthy tokens |
<base>.shared |
device plugin (Shared) | unpartitioned only | distinct accelerators on one node, one ownership share on each (10 per accelerator) | Σ healthy tokens |
<base>.sliced |
device plugin (Sliced) | logically sliceable only | always 1 |
Σ healthy tokens |
<base>.sliced.units |
node capacity | logically sliceable | webhook-derived, per accelerator | Σ accelerators × 1,600,000 |
<base>.sliced.cores-percentage |
node capacity | logically sliceable | per accelerator, (0,100] |
Σ per-accelerator budget |
<base>.sliced.memory-percentage |
node capacity | logically sliceable | per accelerator, (0,100] |
Σ per-accelerator budget |
<base>.sliced.memory-mib |
node capacity | logically sliceable | per accelerator, ≤ accelerator VRAM | Σ per-accelerator budget |
<base>.partitioned |
device plugin (Partitioned) | partitioned only | always 1 |
Σ healthy tokens |
<base>.partitioned.units |
node capacity | partitioned | webhook-derived | Σ accelerators × 1,600,000 |
<base>.partitioned.<kind>-<profile> |
node capacity | partitioned | always 1 |
Σ (allocated + remaining) |
device.gpustack.ai/<manufacturer>.visibility |
device plugin (Visibility) | every accelerator | sidecar’s accelerator count | Σ tokens |
Two things the table does not say:
- Never write a
.unitskey by hand. The Pod webhook recomputes it from the memory budget (logical) or the profile’s VRAM (partition), overwriting any client value. It feeds Kueue’screditstransformation, so a partition and a logical slice of the same VRAM cost the same credits. - The two token shapes differ.
<base>,.sharedand.slicedtokens are accelerator-bound: the token the kubelet picks is the accelerator..partitionedand visibility tokens are a fungible count: the plugin picks the accelerator itself against the live partition geometry and records the one it used. So a partition request never lands where its profile does not fit, and a rejection fromAllocatemeans the whole node has no room.
Worked example per family
All four are Pods submitted on a pool’s entrance LocalQueue (kueue.x-k8s.io/queue-name).
Exclusive — two whole accelerators:
resources:
limits:
nvidia.com/gpu: "2"Shared — two accelerators on one node, one of each accelerator’s 10 ownership shares:
resources:
limits:
nvidia.com/gpu.shared: "2"The value counts accelerators, never shares of one: "2" needs a node with two accelerators that each
still have a free share. It is per container, and each container is one holder: two Pods, or two
containers of one Pod, may hold a share of the same accelerator. A share carries no memory or compute
cap: up to 10 holders use the whole accelerator side by side; isolating them is a logical slice’s job.
The Pod webhook pins a request of two or more (its largest container’s) to nodes with that many
accelerators. It adds
acceleratable.feature.gpustack.ai/<group>.count Gt N-1 to every required node-affinity term, reading
<group> from the InstanceType fronting the queue, and rejects the Pod when that cannot be read.
Ten tokens per accelerator hide a node’s accelerator count from Kueue and the scheduler; the label does not.
Logical slice — half of one accelerator’s VRAM, capped at 40 % of its compute:
resources:
limits:
nvidia.com/gpu.sliced: "1" # always exactly 1 accelerator
nvidia.com/gpu.sliced.memory-percentage: "50" # per accelerator; or .sliced.memory-mib, never both
nvidia.com/gpu.sliced.cores-percentage: "40" # per accelerator; defaults to 100 when omitted
# nvidia.com/gpu.sliced.units is folded by the webhook — do not set itPhysical partition — one MIG 3g.40gb instance:
resources:
limits:
nvidia.com/gpu.partitioned: "1" # always exactly 1 accelerator
nvidia.com/gpu.partitioned.mig-3g.40gb: "1" # always exactly 1 instance
# nvidia.com/gpu.partitioned.units is folded by the webhook — do not set itThe SSH sidecar sits outside the four families, so no family rule counts it. It requests the internal visibility resource for its workload container’s accelerator count:
resources:
limits:
device.gpustack.ai/nvidia.visibility: "1"The request rules
The rules are scoped by a container’s lifetime group, not by the field it sits in:
- the init group is
spec.initContainerswithoutrestartPolicy: Always; - the running group is
spec.containers; - a native sidecar —
spec.initContainerswithrestartPolicy: Always— belongs to neither. It starts during the init phase and keeps running, so it overlaps every later init container as well as every app container.
Rule 1 — one family, in exactly one container group
All containers of a Pod request the same accelerator family, and the accelerator claims sit in exactly one container group.
Accepted — both app containers claim the same family:
spec:
containers:
- name: trainer
resources:
limits:
nvidia.com/gpu: "1"
- name: sidecar-metrics
resources:
limits:
nvidia.com/gpu: "1"Rejected — two families in one Pod:
spec:
containers:
- name: a
resources:
limits:
nvidia.com/gpu: "1"
- name: b
resources:
limits:
nvidia.com/gpu.sliced: "1"
nvidia.com/gpu.sliced.memory-percentage: "50"
spec: Forbidden: a Pod may request only one accelerator family, found [exclusive sliced]
Rejected — the same family claimed in both container groups:
spec:
initContainers:
- name: warmup
resources:
limits:
nvidia.com/gpu: "1"
containers:
- name: main
resources:
limits:
nvidia.com/gpu: "1"
spec.initContainers: Forbidden: a Pod's accelerator requests must all sit in one container group; spec.initContainers must give up its request, because its devices are held for the Pod's whole life while the scheduler charges the Pod only once
Needs a device in two phases? Keep the claim on the app container: its device is the same hardware the init container would have held.
Why — two independent reasons, and one group makes charge and consumption agree by construction.
- Nothing releases the earlier claim. The kubelet holds a finished init container’s devices for the Pod’s whole life, and GPUStack’s reclaimer destroys a partition on Pod deletion, not container termination. The second claim coexists with the first.
- The scheduler charges once. A Pod’s demand for a key is
max(Σ init, Σ app), so two same-family claims cost one unit of quota while consuming two accelerators. The node over-advertises by one slot per such Pod, and the next tenant fails terminally.
Rule 2 — <base>.sliced is exactly 1
Accepted: nvidia.com/gpu.sliced: "1". Rejected:
resources:
limits:
nvidia.com/gpu.sliced: "2"
nvidia.com/gpu.sliced.memory-percentage: "50"
spec.containers[0].resources.limits[nvidia.com/gpu.sliced]: Invalid value: "2": a logical slice request is always a single accelerator; multi-accelerator logical slicing is not supported yet
A multi-accelerator workload asks for several Pods, or for whole accelerators.
Why — a deferral, not a manufacturer limit: NVIDIA can isolate more than one accelerator. No node-level key expresses “N distinct accelerators” — an additive scalar cannot tell “two at half each” from “one in full”, so a one-accelerator node would accept the request and fail it at
Allocate. Lifting the cap needs a node-level accelerator-count dimension.
A logical slice must also name exactly one memory budget:
- neither
.sliced.memory-percentagenor.sliced.memory-mib→Required value: a nvidia.com/gpu.sliced request must set nvidia.com/gpu.sliced.memory-percentage or nvidia.com/gpu.sliced.memory-mib; - both →
Forbidden: cannot set both …memory-percentage and …memory-mib; - a percentage outside
(0,100], a non-positive MiB value, or a MiB value above the accelerator’s VRAM → rejected.
Rule 3 — <base>.partitioned is exactly 1
Accepted: nvidia.com/gpu.partitioned: "1". Rejected:
resources:
limits:
nvidia.com/gpu.partitioned: "2"
nvidia.com/gpu.partitioned.mig-1g.10gb: "1"
spec.containers[0].resources.limits[nvidia.com/gpu.partitioned]: Invalid value: "2": a partition request is always a single accelerator; request one Pod per instance
Unlike rule 2 this is a scope decision: the plugin picks the accelerator itself, so N > 1 would
be implementable. No workload needs it yet, so a multi-partition workload asks for several Pods.
A bare accelerator key with no profile is also rejected; there is no hardware shape to actuate.
Required value: a nvidia.com/gpu.partitioned request must name a profile, e.g. nvidia.com/gpu.partitioned.<kind>-<profile>
Rule 4 — a partition request names exactly one profile shape
Accepted: one per-profile key. Rejected:
resources:
limits:
nvidia.com/gpu.partitioned: "1"
nvidia.com/gpu.partitioned.mig-1g.10gb: "1"
nvidia.com/gpu.partitioned.mig-3g.40gb: "1"
spec: Forbidden: a Pod may request only one partition profile, found [1g.10gb 3g.40gb]
The rule is Pod-wide as well as per-container: two containers naming different profiles are rejected the same way.
Rule 5 — each per-profile key is exactly 1
Accepted: nvidia.com/gpu.partitioned.mig-3g.40gb: "1". Rejected:
resources:
limits:
nvidia.com/gpu.partitioned: "1"
nvidia.com/gpu.partitioned.mig-3g.40gb: "2"
spec.containers[0].resources.limits[nvidia.com/gpu.partitioned.mig-3g.40gb]: Invalid value: "2": a partition profile request must be exactly 1 instance
Rule 6 — at most one container may request a slicing family
Logical or physical, one claiming container per Pod. The SSH sidecar is unaffected: the visibility resource sits deliberately outside the accelerator families.
Accepted:
spec:
containers:
- name: main
resources:
limits:
nvidia.com/gpu.sliced: "1"
nvidia.com/gpu.sliced.memory-percentage: "50"
- name: sshd
resources: # visibility, not a slicing family
limits:
device.gpustack.ai/nvidia.visibility: "1"Rejected — two containers each holding a slice:
spec:
containers:
- name: a
resources:
limits:
nvidia.com/gpu.sliced: "1"
nvidia.com/gpu.sliced.memory-percentage: "50"
- name: b
resources:
limits:
nvidia.com/gpu.sliced: "1"
nvidia.com/gpu.sliced.memory-percentage: "50"
spec: Forbidden: at most one container may request a slicing family, found 2
Rule 7 — a restartable init container may not request an accelerator
Accepted — the claim sits on an app container while the native sidecar carries none; rejected — the native sidecar carries the claim:
spec: # accepted
initContainers:
- name: log-shipper
restartPolicy: Always
resources:
limits:
cpu: "100m"
containers:
- name: main
resources:
limits:
nvidia.com/gpu: "1"
---
spec: # rejected
initContainers:
- name: log-shipper
restartPolicy: Always
resources:
limits:
nvidia.com/gpu: "1"
spec.initContainers[0].resources.limits: Forbidden: a restartable init container (a native sidecar) may not request an accelerator; move the request to an app container
A native sidecar belongs to neither container group, so its claim would overlap every later init container and every app container — the exact double-consumption rule 1 exists to prevent, with no group to move it out of.
The RDMA keys, beside the accelerator families
An RDMA interface is requested through three node-level keys that sit beside the accelerator
families, not inside them: a network interface belongs to the node rather than to a manufacturer,
so no <base> applies. The mechanism (which interface serves which key, and what a grant hands
the container) is described in
Network Topology
; the request rules are here.
RDMA Operations covers how many to ask for and how to configure a node so the request lands well.
All three are served by the device plugin, and a request names what it wants one of:
| Key | Quantity meaning |
|---|---|
device.gpustack.ai/rdma |
whole interfaces |
device.gpustack.ai/rdma.shared |
concurrent uses of an interface |
device.gpustack.ai/rdma.partitioned |
virtual functions |
What a quantity of each key means, how many tokens an endpoint carries, and which allocation mode
each key belongs to, are stated once with the mechanism:
the RDMA resource keys
.
There are no .units keys on this side: nothing is webhook-derived, and the value you set is the
value that schedules.
The keys are not an accelerator family. The family classifier returns none for them, so the accelerator rules above neither apply to them nor can be violated by them: an accelerator family and an RDMA key in one Pod is legal.
The same blindness reaches Kueue, which therefore meters these keys not at all. See RDMA Operations for what that costs a fleet and what to do instead of a quota.
Neither a node without an RDMA-capable interface nor an endpoint whose link the node judged
failed carries allocatable tokens (what each one advertises is on
the mechanism page
),
so a request for these keys never schedules onto them.
A grant injects the endpoint’s own verbs character device, the node-level connection-manager device
where the host has one, and NCCL_IB_HCA naming the granted devices. That set is evidenced for
RoCE and for nothing else — see
what an allocation hands over
.
Shared accelerator and RDMA interface in one container, the shape a co-located workload uses:
resources:
limits:
nvidia.com/gpu.shared: "1"
device.gpustack.ai/rdma.shared: "1" # one concurrent use of one RDMA interfaceCo-locating an accelerator and an RDMA interface
A container that requests an accelerator and an RDMA interface together is granted both on one NUMA node (the kubelet aligns resources that publish NUMA hints, and both sides publish them) but only when every condition below holds. Outside these conditions nothing aligns the two, and no error says so.
- Both requests sit in the same container. The kubelet aligns per container by default, so an accelerator in one container and an RDMA interface in another are aligned by nothing.
- Both resources publish a hint. The accelerator’s whole-device modes do; an accelerator
partition token does not (it names no accelerator), so a partition profile paired with
rdma.partitionedis aligned to the RDMA side alone and is not covered, even though both are allocatable in the same container. An RDMA endpoint whose affinity the kernel did not report publishes none either. - The node’s kubelet runs
restrictedorsingle-numa-node. It is node-level kubelet configuration this operator cannot change, and what each of the four policies does with a hint is in the preflight runbook , which is also where you read the policy a given node runs. No release sets one for you — how to set it, and how to confirm it took . - The alignment’s unit is the NUMA node. See the RDMA resource keys for why a shared PCIe switch is finer than a hint can express.
What is not guaranteed is the node itself: the scheduler selects by quantities and cannot see inside a node’s topology, so a node whose totals suffice but whose devices sit on two NUMA nodes is selected and the container is then refused at admission — and a Pod refused that way is not rescheduled onto another node.
device-manager preflight reports the node’s TopologyManager policy as a topology section, read
off the running kubelet’s own command line and the configuration it names; unknown means no
readable source named one, never a guess of the default. See
Preflight Operations
.
Requesting through the Instance API
An Instance expresses the same request families through spec.resources, and the controller shapes them
into the keys above:
| Field | Effect |
|---|---|
accelerator: "N" |
whole accelerators (exclusive) — may span accelerators |
accelerator: "1"acceleratorSlicedMemoryPercentageacceleratorSlicedCoresPercentage (optional) |
one logical slice |
accelerator: "1"acceleratorPartitionedProfile: "3g.40gb" |
one hardware partition of that profile |
kind: Instance
spec:
type: gpustack--nvidia-h100-80gb-hbm3-linux-amd64
resources:
accelerator: "1"
acceleratorPartitionedProfile: 3g.40gbThe webhook rejects these at admission rather than leaving the Instance Pending:
- a profile the pool does not offer — the message lists the offered set;
- a profile on a manufacturer with no hardware partitioning;
- a profile and a slice percentage together (
a hardware partition and a logical slice percentage are mutually exclusive); - an
acceleratorcount other than1for a slice or a partition request; - a whole-accelerator count over the pool’s whole-accelerator capacity (
status.accelerator.capacity), never over what is free right now: an Instance submitted while every accelerator is held is admitted and waits in its queue; - a slice percentage against a pool that offers no logical slicing — a tightening; such a request
used to be silently reshaped into a whole-accelerator one and served. On an all-partitioned pool the
message points at
spec.resources.acceleratorPartitionedProfileinstead.
A partition’s host CPU and RAM are sized by the profile’s share of the accelerator’s VRAM, so a 1g
instance does not ask for a whole accelerator’s worth.
Pre-release breaks
No released version is marked stable, so each of the following is a clean break with no translation layer.
The old MIG key is gone. A MIG profile used to be requested through the logical family: a
per-profile key built from <base>.sliced with a mig-<profile> segment appended, alongside
<base>.sliced: 1. It is replaced by <base>.partitioned.<kind>-<profile> alongside
<base>.partitioned: 1.
Nothing recognizes the old one — not the builders, not the parsers, not the node-capacity reconciler, and deliberately not even a rejection by name. Rewrite such a manifest; untouched, it meets two ordinary failures, neither naming the replacement:
- its per-profile key is an extended resource no node advertises, so it never schedules;
- the
<base>.sliced: 1beside it is a logical slice with no memory budget, which rule 2 rejects.
Why — legibility would cost a legacy branch in every key path, to serve a request no documentation, no
InstanceTypeand no example produces.
A development node an earlier build wrote those legacy capacities onto keeps them, because nothing owns them. List what is left and patch it off — one JSON-patch removal per stale key:
$ kubectl get node <node> -o json | jq -r '.status.capacity | keys[] | select(contains("mig-"))'
$ kubectl patch node <node> --subresource=status --type=json \
-p '[{"op":"remove","path":"/status/capacity/<escaped-key>"}]' # escape "/" in the key as "~1"
The allocation annotation’s value changed too, so drain a node before upgrading its device manager.
The Pod annotation device.gpustack.ai/accelerator.allocated is now a per-container map; the old flat
shape is not read. Drain the node (or delete its accelerator Pods) before rolling the device-manager
DaemonSet, then let the workloads reschedule.
Why — a Pod carrying the old shape on a restarted device manager drops out of the ledger: its accelerators read free while its containers hold them, and the next opposite-mode Pod can land on an occupied one. The rebuild cannot recover — the occupancy is exactly what became unreadable — so it logs loudly, naming the Pod.
<base>.shared: N counts accelerators now. It used to be read as N shares of one accelerator in
some places and N accelerators in others; it is N distinct accelerators on one node, one share on each.
A manifest asking "3" for a heavier share of one accelerator now needs a node with three.
InstanceType.status.acceleratorShared.onceMaxRequest moved with it: it counts a node’s accelerators
with a free share, where it used to sum their shares. Output recorded before the change, in the
Walkthrough
and NVIDIA MIG Operations
, still shows the
old reading under SH — 10/10 on one free accelerator now reads 1/10.
Limitations
-
Media-engine and graphics profile variants are not exposed. A profile whose name is not a valid Kubernetes resource-name segment (the
+me,+me.alland+gfxMIG variants) is excluded from an accelerator’s inventory rather than rewritten to something key-safe, so a key always maps back to its profile by a plain prefix strip. Those variants cannot be requested. -
One accelerator per slice or partition request (rules 2 and 3), for the different reasons given above.
-
The whole-accelerator bound is the pool’s total, not one node’s. A count above the largest node’s accelerators but within the pool’s total is admitted and then stays queued, because one Pod’s accelerators come from one node and
InstanceType.statuscarries no per-node capacity to refuse it with. It applies to anInstanceand aModelDeploymentrole alike. -
Hand-carving a partition outside GPUStack is unsupported on a managed node. Every node-level key — per-profile capacity, partition token health, the admission check — comes from the Pod annotations the device plugin writes, and an instance made with
nvidia-smi mig -cgiproduces none.The node keeps advertising room it does not have, and unlike a transient over-advertisement this never converges: it persists until the instance is removed. Placement reads the live hardware and will not double-book it, but the accounting above stays wrong. Let GPUStack materialize the instances; it reuses any already on an accelerator it manages.
-
Flipping an accelerator’s partitioning mode is an operational procedure, not a live switch. An accelerator advertises a family’s tokens only while its reported capability backs that family, so a flip removes the old tokens rather than marking them unhealthy: there is no continuity across it.
Drain the accelerator (the hardware refuses the toggle under load anyway), flip the mode with the manufacturer’s tool, then restart that node’s Device Manager DaemonSet: the re-detect trigger watches the device set and health, not the partitioning mode. Deleting the node’s
Devicesobject is not required, since an existing group’s capability is rewritten in place.Full procedure: NVIDIA MIG Operations · T-Head MIG Operations · Hygon MIG Operations .
-
A shared request on a node with fewer free accelerators than it names fails for good. Each accelerator advertises 10 shared tokens, so the scheduler fits
.shared: "2"on a one-accelerator node; the device plugin refuses the two tokens the kubelet picked there rather than grant one short. The Pod endsFailedwithUnexpectedAdmissionErrorand must be recreated.A pool’s queue holds such a request first (Admission ), so a Pod reaches this only by skipping the queue or racing another allocation.
-
The shared card-count pin has two gaps. Where it misses, a request placed on too few accelerators retries for good (Admission ).
- Workloads built from a template. Kueue builds a batch Job’s, a JobSet’s or a RayJob’s Workload from its Pod template before any Pod exists, and the webhook only sees Pods, so none is pinned.
- Partitioned accelerators count.
.countis the node’s total, so a node whose accelerators are mostly in a partitioning mode still passes the pin with too few left to share.
-
A non-default
TopologyManagerpolicy can mis-align a partition. The Partitioned resource reports no NUMA topology (the plugin may not use the accelerator the kubelet aligned to), so undersingle-numa-nodethe CPU and memory providers can settle on one socket while the only accelerator with room is on the other. The defaultnoneis unaffected.