Admission
A workload passes five admission gates before a container receives a device. Kueue accounts for pool capacity, while GPUStack checks whether individual accelerators can satisfy the request.
Contents
- The five gates
- The
Devicesledger beneath the gates - Four-view status
- Capability versus availability
- The InstanceType and Instance webhooks
- The KV cache injection webhook
- Update validation while an object is deleted
- Running-instance stop
- Known behavior: the deployed Kueue Configuration
The five gates
A layered five-gate admission model: Kueue is a coarse gate, and the per-accelerator ledger is the
fine one. Each gate produces or consumes the .sliced.* / .partitioned.* values at a distinct
point on the path.
| # | Gate | Sees | Cannot see |
|---|---|---|---|
| 1 | Pod webhook (Worker) | the request’s shape; folds memory into credits | cluster-wide capacity |
| 2 | Kueue credits |
the pool’s aggregate total | per-accelerator fragmentation |
| 3 | NodeDevicesAdmission AdmissionCheck |
the accelerators of the node Kueue TAS assigned, via the ledger | how to move a Workload off that node: for an unpinned Workload, TAS re-places a Retry from the same per-node totals |
| 4 | Default scheduler / kubelet | per-node remaining capacity keys | which accelerator, for the partitioned family |
| 5 | Device-plugin allocator | the live accelerator state, under a per-node mutex | anything upstream of its own node |
Gate 1 — the Pod webhook
A pods CREATE webhook (objectSelector kueue.x-k8s.io/queue-name, failurePolicy: Fail) enforces the
normative request rules
and folds each family’s credit
input:
- logical slice —
.sliced.memory-percentage/.sliced.memory-mibinto.sliced.units: the memory demand over the pool InstanceType’s per-acceleratorMemory; - hardware partition — the profile’s VRAM into
.partitioned.unitsby that same VRAM-anchored fold, so a3g.40gbpartition costs what a same-VRAM slice costs.
Memory thus reaches the credits fold-down before Kueue scores it. The webhook defaults
.sliced.cores-percentage to 100, prefers .sliced.memory-percentage over .sliced.memory-mib, and
always recomputes the fold, since no trusted path sets it.
It then validates the seven request rules , which that page states with an accepted and a rejected example each.
The divisor is the operator-owned InstanceType’s spec.memory, found via the
schedule.gpustack.ai/queue-entrance label — never the user-writable LocalQueue. Its
MutatingWebhookConfiguration sorts before kueue-mutating-webhook-configuration, so the fold precedes
Kueue’s resource hash.
Gate 2 — Kueue credits
Coarse total admission by fractional scoring (capacity × M); a scalar total cannot see per-accelerator
fragmentation.
Gate 3 — the per-accelerator AdmissionCheck
When Kueue reserves quota, the gpustack-node-devices AdmissionCheck reads the Devices ledger of
every node the assigned ResourceFlavor names and checks the fit per accelerator. The credits gate
before it sees only a scalar total, so it cannot tell that eight accelerators sliced to 50 % each
cannot serve a request for five whole accelerators. This check can, and holds
the Workload with Retry when the fit fails.
The worker applies the gpustack-node-devices AdmissionCheck at startup, retrying
until Kueue’s CRD is established — the chart cannot ship it, since Kueue templates its own CRDs and
nothing orders them ahead of a custom resource in the same render (see Install
modes
).
The worker keeps it Active; an accelerated queue references it in spec.admissionChecksStrategy
only once it is (Queue quota and draining
).
Once Kueue reserves quota, the check reads each node’s Devices ledger for per-accelerator
Remaining ≥ demand: a whole accelerator for exclusive, .sliced.units for sliced, an owner share for
shared, a free placement of the profile for a partition. It rebuilds that ledger from the Pods bound to
the node, one uncached read, with the device manager’s own aggregation.
A Workload answered Ready holds its accelerators before its Pods do. The ledger moves only when
Allocate records a Pod, after Kueue admits the Workload, the scheduler binds its Pods and the kubelet
admits them. A second Workload judged in that window would see the first one’s accelerators as free.
- So the check first fits every other Workload on the node that is admitted or holds this check’s
Ready: the Pods its topology assignment puts there, minus those the rebuilt ledger already holds. - A Pod is held once its allocation record names every container asking for an accelerator. The
kueue.x-k8s.io/workloadannotation andkueue.x-k8s.io/podsetlabel tie it to its Workload. - The ledger and the held Pods come from the same read, so no Pod is counted twice or missed.
- A finished Pod, in phase
SucceededorFailed, still counts as held, while the ledger stops charging the cards the kubelet took back from it (device discovery ). A replacement Pod its Workload creates is then counted by neither untilAllocaterecords it. - The check judges one Workload at a time, and remembers each
Readyit wrote until the cache shows it.
Known behavior: the inflight count follows the allocator’s hint. An inflight slice is fitted on the accelerator the allocator’s packing order prefers, which the kubelet normally takes. A kubelet that ignores the hint can put it elsewhere, and a slice judged after it can then find no room on that accelerator:
Allocaterefuses it and the Pod fails withUnexpectedAdmissionError. A Workload without a hostname-level assignment is not counted as inflight; every queue this operator derives assigns one.
Slices share an accelerator the way the allocator packs them: each is charged its .sliced.units and
one of the accelerator’s slice slots on the fullest accelerator that still fits it, so two 30 % slices
fit one free accelerator and two 60 % slices do not. Only a free or already-sliced accelerator takes one.
The slots are counted from the ledger’s allocatedSlices as well as this Workload’s own slices, so a
Hygon accelerator holding its four slices takes no fifth however much memory it has left.
Under TAS the check judges the node already assigned. Every queue this operator derives is TAS-only
and ends at kubernetes.io/hostname, so Kueue writes the node into the Workload’s
podSetAssignments[].topologyAssignment as it reserves quota, before any AdmissionCheck settles. Each
node named there must host its own share of the PodSet from its own accelerators.
- A PodSet without a hostname-level assignment is judged across the flavor’s whole node pool.
- A node whose
Nodeobject is gone serves no node-scoped demand, so the Workload is held.
Known behavior: a fragmented node is skipped only when its Workload is pinned. TAS sums each node’s capacity keys, so it cannot see how the free capacity is spread over the node’s accelerators:
- a node’s free
.sliced.unitsmay be spread over accelerators none of which fits the slice;- a node’s 10 shared tokens per accelerator let a node with too few accelerators take
.shared: N. The Pod webhook’s card-count pin keeps such a request off those nodes and their flavors. Where the pin misses (Limitations ) it never frees, since Kueue restarts the flavor scan at the smallest node after every eviction.Left alone, TAS prefers such a node, the check answers
Retry, and TAS places the Workload there again every 30 s. The per-card fit labels close that loop for a pinned Workload: TAS skips the node, and when no node fits, the Workload stays pending onexcluded: affinityand holds no quota.Reading the pool instead let such a request through to
Allocate, which refuses a slice its accelerator cannot hold and a shared request its node cannot spread over distinct accelerators, failing the Pod.
The Workload fit pin: a mutating webhook on Kueue Workload CREATE and UPDATE adds the pin to each
PodSet, so a logical slice of U units per accelerator gets sliced-max-free-units… Gt U-1, and a
shared request of N >= 2 accelerators gets shared-free-cards… Gt N-1. N is the largest count
any one container asks for.
- It changes only a Workload whose LocalQueue points at a ClusterQueue carrying this operator’s InstanceType mark, with a same-named InstanceType naming an accelerator group, so other tenants' Workloads in a shared Kueue are out of its reach.
- It acts on an UPDATE only while the old Workload holds no quota reservation, which is when Kueue rebuilds a suspended job’s spec and still lets PodSets change.
- It never denies. Its
failurePolicyisIgnore, and the settingworkload-fit-affinityturns it off while the labels keep being published. - The pin never reaches the Pod or its owner’s template: Kueue copies only labels, annotations, a
nodeSelector, tolerations and scheduling gates onto them. A Pod-level pin would be unsafe, because the kubelet re-admits running Pods against current node labels when it restarts.
What still retries:
- A label not yet refreshed. A Workload placed before the label follows an allocation still reaches this check and retries once; the next placement reads the new value.
- Room promised to a Workload that is still starting. The label cannot see it, so TAS may place a Workload on those accelerators. The check holds it until the first Workload’s Pods are allocated.
- Several Pods, or several accelerators, on one node. The label admits a node when one accelerator fits one Pod, so two Pods of one PodSet, two PodSets, or one template-built Pod asking for two sliced accelerators can still be placed where only one fits. That does not converge by itself.
- An unpinned Workload. One created while the webhook was unavailable or the setting was off, one
created before the upgrade, and a template-built slice whose template carries no folded units keep
the loop described above. A
Retrydoes not repair a missed pin, because Kueue requeues that same object; only a spec update before reservation re-pins it.
Each family gets one correlated (accelerators, per-accelerator demand, profile) tuple scoped to the
accelerators that can serve it, so an exclusive or shared request is never judged feasible against a
partitioned accelerator: it stays queued, not admitted into a permanent Pending. For partitions
accelerators counts instances, not accelerators — one hosts as many replicas as its remaining
geometry allows.
Feasibility is per role. A Workload composed from a pod group carries one PodSet per role, and each PodSet gets its own ResourceFlavor — so a demand is judged against the accelerators of the flavor its own PodSet was assigned, never against every accelerator in the pool.
- Two PodSets’ demands merge only when the flavor agrees as well as the family.
- An accelerator is eligible for a demand only when its own key is one that demand’s flavor covers.
- The per-accelerator budget stays global to the Workload, so two roles on one flavor cannot spend one accelerator twice.
- The verdict message names the role that is starved.
Nothing this operator renders produces a multi-PodSet Workload today. A
ModelDeploymentadmits each replica as a pod group of its own, and every member of that group carries the same role — so the Workload holds one PodSet, whose count is the role’ssize, never one per role.The reasoning above therefore describes the check’s shape rather than a state its own workloads reach. It stands because the check answers for every workload in the queues it is referenced from, not only for the ones rendered here.
The ledger seeds every accelerator at M, so an exclusive over-admit that coarse credits let
through is caught exactly, held with Retry, and self-heals once Kueue re-admits after the backoff.
The gate only checks: it never preempts and never answers Rejected.
The check never judges an evicted Workload: Kueue resets its checks and quota reservation in two separate writes, and a verdict written between them would overwrite the reset and deadlock the backoff loop instead of letting it self-heal. Re-reserving quota clears the eviction condition and re-opens evaluation.
Gate 4 — default scheduler / kubelet
Node-level counting of each family’s remaining keys — the bare .sliced / .partitioned token
plus its logical
or
partitioned
counting keys — checks the node
against those keys.
On a TAS queue it does not choose the node: TAS assigned one at quota reservation and the Pod is
pinned to it, so a node without room leaves the Pod Pending rather than moving it elsewhere.
Disjoint accelerator populations advertise the two families’ keys, so the resource name alone rules out
a node that cannot serve the kind at all — the one placement error Allocate can never repair.
Gate 5 — the device-plugin allocator
At Allocate, the Device Manager settles the accelerator, injects the container’s visibility env and
runtime isolation, and records the allocation in the Devices ledger. “Settles” differs by family:
- accelerator-bound — the kubelet chose the accelerator by choosing the token, so the allocator refuses one another mode holds and, for a logical slice, one without a free slot or the units the slice needs;
- partitioned — the tokens are a fungible count, so the allocator picks the accelerator itself and materializes the hardware instance on it.
The slice refusal is the one gate every path reaches — a Pod outside the scheduling chain, or a hint the kubelet declined — so a slice the gates above let through by mistake fails its Pod rather than oversubscribing an accelerator. Its controller recreates the Pod; the refusal and its off switch are in Container identification .
Both paths, and the per-manufacturer isolation each slice gets — it covers every sliceable manufacturer
— are in Device Discovery
. The Pod webhook caps
.sliced and .partitioned at exactly 1, so the manufacturer-specific multi-slice divergence that
cap hid can no longer be requested.
The Devices ledger beneath the gates
The Devices CR AcceleratorAllocation ledger is the single authoritative accounting, written below the
kubelet by the device-plugin Allocate for every allocation, Kueue-routed or not. It drives the
four-view and backs gate 3, but is not itself a gate.
Four-view status
InstanceType.status carries four per-accelerator bin-packing projections from the Devices ledger, not
a credits fold-down:
| View | Column | Counts |
|---|---|---|
Accelerator |
EX | free whole accelerators |
AcceleratorShared |
SH | shareable ownership slots |
AcceleratorSliced |
SL | logically sliceable VRAM-percent units |
AcceleratorPartitioned |
PT | hardware partition instances the pool’s partitioned accelerators can still host |
Each accelerator feeds exactly one group, by the capability it reports and never its scalar ledger:
EX/SH/SL count unpartitioned accelerators, PT partitioned ones. A partition-only pool thus reads
0/0 0/0 0/0 for the first three, a logical-only pool 0/0 for PT, and an exclusive tenant is never
shown capacity a partitioned accelerator could not serve.
OnceMaxRequest differs per view:
EX— the largest single node’s free accelerators; one request can span a node’s accelerators;SH— the largest single node’s accelerators that still have a free share, since a shared request of N names N distinct accelerators on one node, and a node’s spare shares on one accelerator do not add;SL— the freest single accelerator’s; a slice targets one accelerator;PT—1while any accelerator can host an instance, else0; a partition request is validated to be exactly one instance on one accelerator, so nothing larger is requestable.
kubectl get instancetypes folds them into the Accelerator(EX/SH/SL/PT) column as four
onceMaxRequest/remaining groups.
Capability versus availability
For a partitioned pool, neither PT number reports which profiles are still available.
onceMaxRequest is the 1/0 answer to whether any partition fits. remaining sums each
accelerator’s largest free count across profiles that compete for the same physical slices;
it does not give a total for each profile.
The per-profile answer is status.acceleratorPartitioned.remainingProfiles, paired with
allocatedProfiles: the pool-level Σ by profile name of the per-accelerator ledger on
Devices.status. Every profile the pool offers is listed even at zero, so one a sibling’s instance
filled reads 0 instead of vanishing, keeping “offered but full” distinct from “not offered”.
kubectl get instancetypes -o wide shows the same list as the PARTITIONS column; the worker gateway
sums it across clusters (Active members only, like every availability dimension).
Do not read status.detail.slicedDetail for this. It aggregates the static slicing capability
catalog from the Devices spec side, and that catalog deliberately does not move as instances are
carved and released. The Instance webhook uses the catalog to reject an unoffered profile while
naming the offered set, and to size a request from the profile’s MemoryMib, which the ledger does
not carry.
Repurposing those counts as availability would make a momentarily-full profile vanish from the offered
set, turning a request that should stay Retry at the AdmissionCheck into a permanent rejection.
The partition views are likewise enumerated from the capability side, scoped to the pool’s own accelerator group (a node can carry several models) and joined to the ledger. An accelerator the detector reported is never dropped for a missing ledger row, nor read as full for an empty one: it falls back to its catalog ceilings, as the node’s per-profile capacity keys do.
Why the status lives on a real CRD — the worker watches the
Devicesresource and writes into a real CRD’s.status, sokubectl get instancetype -wobserves capacity move as pods allocate and free. A read-only projection over the ClusterQueue could not.
The InstanceType API proxies the controller-managed resource and converts it to the public representation.
The InstanceType and Instance webhooks
The unit spec lives only on the InstanceType: a derived type is stamped with its per-product preset at creation, and an admin edit touches only the InstanceType, never a Node or the ClusterQueue notes.
- InstanceType validating, create — requires the complete input, read independently of any editable
setting:
acceleratorGroup(only whenacceleratable),os,arch, the unit triple (unitResources.cpu/.ram+localStorage); empty or partial is rejected. A CPU-only (acceleratable=false) type’s unit CPU must be exactly 1 core, an accelerated type any unitless positive integer. - InstanceType validating, update — freezes the spec: every field immutable except
displayName(rename) andinactive(in/out of service), so re-sizing or re-pointing a pool means delete and re-create. Only immutability is re-checked, never the create-time shape, so a legacy type stored before a tightened rule can still be renamed or deactivated. - InstanceType defaulting — an empty
generalGroupto thegenericsentinel; the pool’s schedule labels (grouped byinstance-type-aware-cpu-manufacturer) andschedule.gpustack.ai/queue-entrance(the LocalQueue name format); descriptors enriched from a matching ResourceFlavor; and when awareness is on, that flavor’scpuDetailnote folded intospec.cpu(generic) orspec.accelerator.cpu(accelerated). - Instance validating — enforces the unit spec on Create and Update: a submission’s RAM must not
exceed
unitRAM × count, its local storage not the InstanceType’sLocalStorage.- On a non-accelerated type, create and start also bound the CPU request by the pool’s CPU
capacity (
status.cpu.capacity), never by what is unrequested right now: an Instance submitted while every core is requested is admitted and waits in its queue. - The CPU bound is the pool’s total, not one node’s. A request above the largest node’s cores but
within the pool’s total is admitted and cannot run until a node that large joins, because one Pod’s
cores come from one node and
InstanceType.statuscarries no per-node capacity to refuse it with.
- On a non-accelerated type, create and start also bound the CPU request by the pool’s CPU
capacity (
The KV cache injection webhook
A second mutating webhook on Pods writes the client configuration an inference engine needs to use a KV cache pool . It sits outside the five gates:
- It admits or refuses on its own inputs, consumes no
.sliced.*or.partitioned.*value, produces none, and never touchesresources. - It cannot be a branch of gate 1, because the two select on independent criteria: gate 1 fires on
kueue.x-k8s.io/queue-name, this one onkvcache.gpustack.ai/inject, and aLabelSelectorcannot express the union. The two are independent but not disjoint: a Pod may carry both labels and be served by both entries, which is exactly what the next point is about. - Both entries live in the single
gpustack-worker-mutationconfiguration, whose name sorts before Kueue’s so that gate 1 folds a Pod’s units before Kueue hashes its resources. Their order within it is immaterial, and both run over one Pod whichever order the engine picks. - Before it adds connector arguments, the KV cache webhook removes the transparent launcher prefixes
it knows and accepts only the entry point for the declared engine:
vllmfor the vLLM family orpython3 -m sglang.launch_serverfor SGLang. An unrecognised launcher, launch form, or mismatched engine is refused, so a missing parser entry is visible at admission rather than producing a Pod marked injected whose engine never receives the connector arguments. - A Pod author may set
kvcache.gpustack.ai/launch-args-forwarded: "true"only to declare that an unrecognised launcher, script, or command line hidden in one argument forwards appended arguments to the engine. It does not exempt a shell command mode such assh -c: admission knows that the appended arguments become shell positional parameters and never reach the engine. - The injected record includes
launchProgramandlaunchArgsForwarded, so a JSONPath query can show the resolved executable and whether that author declaration admitted the launch. - A pool whose member groups offer two or more transports is refused for an engine that declares no
required transport, because such an engine is configured with one transport and reads a block from
any group. The
ModelDeploymentvalidating webhook refuses the same binding for the same reason, since its own engine resolution happens there rather than on the Pod.
Update validation while an object is deleted
An UPDATE is neither validated nor defaulted once its object carries a metadata.deletionTimestamp,
unless the handler opts out. The guard exists because an update that clears a finalizer is an UPDATE,
so a rule reading another object and refusing when it is absent could hold an object past its own
teardown. Every handler of this operator except Instance opts out where its rules are safe, and
those rules still validate during deletion.
Instance is the one that keeps the guard, so its UPDATE checks are skipped while it is being
deleted. The reason is its mutating half rather than its validation: its defaulting reads the
referenced InstanceType and is registered failurePolicy: Fail on UPDATE. An InstanceType deleted
ahead of its Instances would leave each one undeletable.
Running-instance stop
Before (re)creating an Instance’s Pod, the worker reads the backing
ClusterQueue’s StopPolicy and stops the Instance (spec.stop=true) rather than recreate a Pod
the queue can never admit. That covers a queue in HoldAndDrain (a pool drain, or a teardown evicting
admitted workloads), an InstanceType being deleted, and an InstanceType already gone.
An admin Hold (the Inactive switch) is deliberately not a stop: running Pods keep running, a new
Instance stays pending.
Why it keys on
StopPolicy— the InstanceType phase collapses bothHoldand a fully-drainedHoldAndDraintoInactive, and a fast drain clears the reservation before a durableDrainingphase is ever observed.
A ClusterQueue watch (on StopPolicy) re-enqueues the type’s Instances so the stop is prompt even when
no Pod event fires; the InstanceType watch is narrowed to the deletion signal for a prompt teardown
stop.
Known behavior: the deployed Kueue Configuration
The feature gate AssignQueueLabelsForPods is disabled in the deployed Kueue Configuration
(kueue.managerConfig.controllerManagerConfigYaml in the chart’s values.yaml), so Kueue never copies
cluster/local queue names onto Pod labels; long ClusterQueue names would not fit a label value.
TopologyAwareScheduling is enabled. Generated ResourceFlavors reference Kueue Topologies, and the
Pod integration converts a ModelDeployment role’s required-level annotation into the Workload
PodSet request described in Topology-Aware Scheduling
.
It also sets resources.quotaCheckStrategy: IgnoreUndeclared, so a single-dimension queue (only cpu,
or only the manufacturer credits) does not reject a Workload for the Pod resources it does not cover
(memory/ephemeral-storage). Its resources.transformations list is generated from the worker’s
node-feature tables by make generate chart.
The same strategy makes a queue with no resource groups admit every Workload: it declares no resource, so nothing is checked, the Workload gets no flavor, and with no flavor it carries no AdmissionCheck. The operator never leaves such a queue admitting — see Queue quota and draining .