# Accelerator Requests

The Pod webhooks check these requests at creation and reject invalid ones before a container starts.
They select Pods with the `kueue.x-k8s.io/queue-name` label. Without that label, a Pod can still
request device-plugin resources, but these checks and resource-unit calculations do not run.

## Contents

- [Two families, two accelerator populations](#two-families-two-accelerator-populations)
- [The resource keys](#the-resource-keys)
- [Worked example per family](#worked-example-per-family)
- [The request rules](#the-request-rules)
- [The RDMA keys, beside the accelerator families](#the-rdma-keys-beside-the-accelerator-families)
- [Co-locating an accelerator and an RDMA interface](#co-locating-an-accelerator-and-an-rdma-interface)
- [Requesting through the Instance API](#requesting-through-the-instance-api)
- [Pre-release breaks](#pre-release-breaks)
- [Limitations](#limitations)

## Two families, two accelerator populations

An accelerator is shared in one of two physically incompatible ways, and GPUStack names them apart:

| Term | Definition | Isolation | Example |
|---|---|---|---|
| **Logical slicing** (`.sliced*`) | software slicing of a whole accelerator | the manufacturer's own sharing facility caps compute and VRAM per container — [which facility, per manufacturer](/gpustack-operator/main/docs/modules/devices/discovery/index.md#sliced-logical-slicing) | 50 % of an A10G |
| **Physical partitioning** (`.partitioned*`) | hardware partitioning of an accelerator put into a partitioning mode | the hardware itself; the operator materializes the instance | an NVIDIA MIG `3g.40gb`, or a partition of a T-Head PPU or a Hygon DCU |

The two never apply to the same accelerator: one in a partitioning mode advertises **only** the partition
family, an unpartitioned one **only** the whole-accelerator, shared and logical-slice families. Hence the
four separate `InstanceType` views (`EX` / `SH` / `SL` / `PT`) — each accelerator feeds exactly one.

`<kind>` is the manufacturer's own word for hardware partitioning, and manufacturers that offer it call it
`mig`: `nvidia.com/gpu.partitioned.mig-3g.40gb`, `alibabacloud.com/ppu.partitioned.mig-<profile>`,
`hygon.com/dcu.partitioned.mig-2g.15gb`. A manufacturer with no hardware partitioning has no kind, and
no `.partitioned*` keys at all.

## The resource keys

`<base>` is the manufacturer's device resource (`nvidia.com/gpu`, `huawei.com/npu`, … — see
[Accelerator support](https://github.com/gpustack/gpustack-operator/blob/4b81d9b1077c38a5355e00bcb8bdd8f0da96239a/README.md#accelerator-support)).

| Key | Served by | Eligible accelerators | Request value | Node value |
|---|---|---|---|---|
| `<base>` | device plugin (Exclusive) | unpartitioned only | accelerator count | Σ healthy tokens |
| `<base>.shared` | device plugin (Shared) | unpartitioned only | distinct accelerators on one node, one ownership share on each (10 per accelerator) | Σ healthy tokens |
| `<base>.sliced` | device plugin (Sliced) | logically sliceable only | always `1` | Σ healthy tokens |
| `<base>.sliced.units` | node capacity | logically sliceable | **webhook-derived**, per accelerator | Σ accelerators × 1,600,000 |
| `<base>.sliced.cores-percentage` | node capacity | logically sliceable | per accelerator, `(0,100]` | Σ per-accelerator budget |
| `<base>.sliced.memory-percentage` | node capacity | logically sliceable | per accelerator, `(0,100]` | Σ per-accelerator budget |
| `<base>.sliced.memory-mib` | node capacity | logically sliceable | per accelerator, ≤ accelerator VRAM | Σ per-accelerator budget |
| `<base>.partitioned` | device plugin (Partitioned) | partitioned only | always `1` | Σ healthy tokens |
| `<base>.partitioned.units` | node capacity | partitioned | **webhook-derived** | Σ accelerators × 1,600,000 |
| `<base>.partitioned.<kind>-<profile>` | node capacity | partitioned | always `1` | Σ (allocated + remaining) |
| `device.gpustack.ai/<manufacturer>.visibility` | device plugin (Visibility) | every accelerator | sidecar's accelerator count | Σ tokens |

Two things the table does not say:

- **Never write a `.units` key by hand.** The Pod webhook recomputes it from the memory budget (logical)
  or the profile's VRAM (partition), overwriting any client value. It feeds Kueue's `credits`
  transformation, so a partition and a logical slice of the same VRAM cost the same credits.
- **The two token shapes differ.** `<base>`, `.shared` and `.sliced` tokens are *accelerator-bound*: the
  token the kubelet picks is the accelerator. `.partitioned` and visibility tokens are a *fungible
  count*: the plugin picks the accelerator itself against the live partition geometry and records the
  one it used. So a partition request never lands where its profile does not fit, and a rejection from
  `Allocate` means the whole node has no room.

## Worked example per family

All four are Pods submitted on a pool's entrance `LocalQueue` (`kueue.x-k8s.io/queue-name`).

**Exclusive** — two whole accelerators:

```yaml
resources:
  limits:
    nvidia.com/gpu: "2"
```

**Shared** — two accelerators on one node, one of each accelerator's 10 ownership shares:

```yaml
resources:
  limits:
    nvidia.com/gpu.shared: "2"
```

The value counts accelerators, never shares of one: `"2"` needs a node with two accelerators that each
still have a free share. It is per container, and each container is one holder: two Pods, or two
containers of one Pod, may hold a share of the same accelerator. A share carries no memory or compute
cap: up to 10 holders use the whole accelerator side by side; isolating them is a logical slice's job.

The Pod webhook pins a request of two or more (its largest container's) to nodes with that many
accelerators. It adds
`acceleratable.feature.gpustack.ai/<group>.count Gt N-1` to every required node-affinity term, reading
`<group>` from the `InstanceType` fronting the queue, and rejects the Pod when that cannot be read.
Ten tokens per accelerator hide a node's accelerator count from Kueue and the scheduler; the label does not.

**Logical slice** — half of one accelerator's VRAM, capped at 40 % of its compute:

```yaml
resources:
  limits:
    nvidia.com/gpu.sliced: "1"                       # always exactly 1 accelerator
    nvidia.com/gpu.sliced.memory-percentage: "50"    # per accelerator; or .sliced.memory-mib, never both
    nvidia.com/gpu.sliced.cores-percentage: "40"     # per accelerator; defaults to 100 when omitted
    # nvidia.com/gpu.sliced.units is folded by the webhook — do not set it
```

**Physical partition** — one MIG `3g.40gb` instance:

```yaml
resources:
  limits:
    nvidia.com/gpu.partitioned: "1"                  # always exactly 1 accelerator
    nvidia.com/gpu.partitioned.mig-3g.40gb: "1"      # always exactly 1 instance
    # nvidia.com/gpu.partitioned.units is folded by the webhook — do not set it
```

**The SSH sidecar** sits outside the four families, so no family rule counts it. It requests the
internal visibility resource for its workload container's accelerator count:

```yaml
resources:
  limits:
    device.gpustack.ai/nvidia.visibility: "1"
```

## The request rules

The rules are scoped by a container's lifetime group, not by the field it sits in:

- the **init group** is `spec.initContainers` *without* `restartPolicy: Always`;
- the **running group** is `spec.containers`;
- a **native sidecar** — `spec.initContainers` *with* `restartPolicy: Always` — belongs to neither. It
  starts during the init phase and keeps running, so it overlaps every later init container as well as
  every app container.

### Rule 1 — one family, in exactly one container group

*All containers of a Pod request the same accelerator family, and the accelerator claims sit in exactly
one container group.*

Accepted — both app containers claim the same family:

```yaml
spec:
  containers:
    - name: trainer
      resources:
        limits:
          nvidia.com/gpu: "1"
    - name: sidecar-metrics
      resources:
        limits:
          nvidia.com/gpu: "1"
```

Rejected — two families in one Pod:

```yaml
spec:
  containers:
    - name: a
      resources:
        limits:
          nvidia.com/gpu: "1"
    - name: b
      resources:
        limits:
          nvidia.com/gpu.sliced: "1"
          nvidia.com/gpu.sliced.memory-percentage: "50"
```

> `spec: Forbidden: a Pod may request only one accelerator family, found [exclusive sliced]`

Rejected — the same family claimed in both container groups:

```yaml
spec:
  initContainers:
    - name: warmup
      resources:
        limits:
          nvidia.com/gpu: "1"
  containers:
    - name: main
      resources:
        limits:
          nvidia.com/gpu: "1"
```

> `spec.initContainers: Forbidden: a Pod's accelerator requests must all sit in one container group;
> spec.initContainers must give up its request, because its devices are held for the Pod's whole life while
> the scheduler charges the Pod only once`

Needs a device in two phases? Keep the claim on the app container: its device is the same hardware the
init container would have held.

> **Why** — two independent reasons, and one group makes charge and consumption agree by construction.
>
> - *Nothing releases the earlier claim.* The kubelet holds a finished init container's devices for the
>   Pod's whole life, and GPUStack's reclaimer destroys a partition on **Pod** deletion, not container
>   termination. The second claim coexists with the first.
> - *The scheduler charges once.* A Pod's demand for a key is `max(Σ init, Σ app)`, so two same-family
>   claims cost **one** unit of quota while consuming **two** accelerators. The node over-advertises by
>   one slot per such Pod, and the *next* tenant fails terminally.

### Rule 2 — `<base>.sliced` is exactly 1

Accepted: `nvidia.com/gpu.sliced: "1"`. Rejected:

```yaml
resources:
  limits:
    nvidia.com/gpu.sliced: "2"
    nvidia.com/gpu.sliced.memory-percentage: "50"
```

> `spec.containers[0].resources.limits[nvidia.com/gpu.sliced]: Invalid value: "2": a logical slice
> request is always a single accelerator; multi-accelerator logical slicing is not supported yet`

A multi-accelerator workload asks for several Pods, or for whole accelerators.

> **Why** — a **deferral, not a manufacturer limit**: NVIDIA can isolate more than one accelerator. No
> node-level key expresses "N *distinct* accelerators" — an additive scalar cannot tell "two at half
> each" from "one in full", so a one-accelerator node would accept the request and fail it at
> `Allocate`. Lifting the cap needs a node-level accelerator-count dimension.

A logical slice must also name exactly one memory budget:

- neither `.sliced.memory-percentage` nor `.sliced.memory-mib` → `Required value: a nvidia.com/gpu.sliced
  request must set nvidia.com/gpu.sliced.memory-percentage or nvidia.com/gpu.sliced.memory-mib`;
- both → `Forbidden: cannot set both …memory-percentage and …memory-mib`;
- a percentage outside `(0,100]`, a non-positive MiB value, or a MiB value above the accelerator's
  VRAM → rejected.

### Rule 3 — `<base>.partitioned` is exactly 1

Accepted: `nvidia.com/gpu.partitioned: "1"`. Rejected:

```yaml
resources:
  limits:
    nvidia.com/gpu.partitioned: "2"
    nvidia.com/gpu.partitioned.mig-1g.10gb: "1"
```

> `spec.containers[0].resources.limits[nvidia.com/gpu.partitioned]: Invalid value: "2": a partition
> request is always a single accelerator; request one Pod per instance`

Unlike rule 2 this is a scope decision: the plugin picks the accelerator itself, so `N > 1` would
be implementable. No workload needs it yet, so a multi-partition workload asks for several Pods.

A bare accelerator key with no profile is also rejected; there is no hardware shape to actuate.

> `Required value: a nvidia.com/gpu.partitioned request must name a profile, e.g.
> nvidia.com/gpu.partitioned.<kind>-<profile>`

### Rule 4 — a partition request names exactly one profile shape

Accepted: one per-profile key. Rejected:

```yaml
resources:
  limits:
    nvidia.com/gpu.partitioned: "1"
    nvidia.com/gpu.partitioned.mig-1g.10gb: "1"
    nvidia.com/gpu.partitioned.mig-3g.40gb: "1"
```

> `spec: Forbidden: a Pod may request only one partition profile, found [1g.10gb 3g.40gb]`

The rule is Pod-wide as well as per-container: two containers naming different profiles are rejected the
same way.

### Rule 5 — each per-profile key is exactly 1

Accepted: `nvidia.com/gpu.partitioned.mig-3g.40gb: "1"`. Rejected:

```yaml
resources:
  limits:
    nvidia.com/gpu.partitioned: "1"
    nvidia.com/gpu.partitioned.mig-3g.40gb: "2"
```

> `spec.containers[0].resources.limits[nvidia.com/gpu.partitioned.mig-3g.40gb]: Invalid value: "2": a
> partition profile request must be exactly 1 instance`

### Rule 6 — at most one container may request a slicing family

Logical or physical, one claiming container per Pod. The SSH sidecar is unaffected: the visibility
resource sits deliberately outside the accelerator families.

Accepted:

```yaml
spec:
  containers:
    - name: main
      resources:
        limits:
          nvidia.com/gpu.sliced: "1"
          nvidia.com/gpu.sliced.memory-percentage: "50"
    - name: sshd
      resources: # visibility, not a slicing family
        limits:
          device.gpustack.ai/nvidia.visibility: "1"
```

Rejected — two containers each holding a slice:

```yaml
spec:
  containers:
    - name: a
      resources:
        limits:
          nvidia.com/gpu.sliced: "1"
          nvidia.com/gpu.sliced.memory-percentage: "50"
    - name: b
      resources:
        limits:
          nvidia.com/gpu.sliced: "1"
          nvidia.com/gpu.sliced.memory-percentage: "50"
```

> `spec: Forbidden: at most one container may request a slicing family, found 2`

### Rule 7 — a restartable init container may not request an accelerator

Accepted — the claim sits on an app container while the native sidecar carries none; rejected — the
native sidecar carries the claim:

```yaml
spec:                                                       # accepted
  initContainers:
    - name: log-shipper
      restartPolicy: Always
      resources:
        limits:
          cpu: "100m"
  containers:
    - name: main
      resources:
        limits:
          nvidia.com/gpu: "1"
---
spec:                                                       # rejected
  initContainers:
    - name: log-shipper
      restartPolicy: Always
      resources:
        limits:
          nvidia.com/gpu: "1"
```

> `spec.initContainers[0].resources.limits: Forbidden: a restartable init container (a native sidecar) may
> not request an accelerator; move the request to an app container`

A native sidecar belongs to neither container group, so its claim would overlap every later init
container *and* every app container — the exact double-consumption rule 1 exists to prevent, with no
group to move it out of.

## The RDMA keys, beside the accelerator families

An RDMA interface is requested through three node-level keys that sit beside the accelerator
families, not inside them: a network interface belongs to the node rather than to a manufacturer,
so no `<base>` applies. The mechanism (which interface serves which key, and what a grant hands
the container) is described in
[Network Topology](/gpustack-operator/main/docs/modules/rdma/network-topology/index.md#the-rdma-resource-keys); the request rules are here.

[RDMA Operations](/gpustack-operator/main/docs/modules/rdma/operations/index.md) covers how many to ask for and how to configure a node so
the request lands well.

All three are served by the device plugin, and a request names what it wants one of:

| Key | Quantity meaning |
|---|---|
| `device.gpustack.ai/rdma` | whole interfaces |
| `device.gpustack.ai/rdma.shared` | concurrent uses of an interface |
| `device.gpustack.ai/rdma.partitioned` | virtual functions |

What a quantity of each key means, how many tokens an endpoint carries, and which allocation mode
each key belongs to, are stated once with the mechanism:
[the RDMA resource keys](/gpustack-operator/main/docs/modules/rdma/network-topology/index.md#the-rdma-resource-keys).
There are no `.units` keys on this side: nothing is webhook-derived, and the value you set is the
value that schedules.

**The keys are not an accelerator family.** The family classifier returns none for them, so the
accelerator rules above neither apply to them nor can be violated by them: an accelerator family and an
RDMA key in one Pod is legal.

The same blindness reaches Kueue, which therefore meters these keys not at all. See
[RDMA Operations](/gpustack-operator/main/docs/modules/rdma/operations/index.md#kueue-and-the-rdma-keys) for what that costs a
fleet and what to do instead of a quota.

Neither a node without an RDMA-capable interface nor an endpoint whose link the node judged
`failed` carries allocatable tokens (what each one advertises is on
[the mechanism page](/gpustack-operator/main/docs/modules/rdma/network-topology/index.md#the-rdma-resource-keys)),
so a request for these keys never schedules onto them.

A grant injects the endpoint's own verbs character device, the node-level connection-manager device
where the host has one, and `NCCL_IB_HCA` naming the granted devices. That set is evidenced for
RoCE and for nothing else — see
[what an allocation hands over](/gpustack-operator/main/docs/modules/rdma/network-topology/index.md#the-allocation-response).

**Shared accelerator and RDMA interface in one container**, the shape a co-located workload uses:

```yaml
resources:
  limits:
    nvidia.com/gpu.shared: "1"
    device.gpustack.ai/rdma.shared: "1"     # one concurrent use of one RDMA interface
```

## Co-locating an accelerator and an RDMA interface

A container that requests an accelerator and an RDMA interface together is granted both on one NUMA
node (the kubelet aligns resources that publish NUMA hints, and both sides publish them) but only
when every condition below holds. Outside these conditions nothing aligns the two, and no error
says so.

- **Both requests sit in the same container.** The kubelet aligns per container by default, so an
  accelerator in one container and an RDMA interface in another are aligned by nothing.
- **Both resources publish a hint.** The accelerator's whole-device modes do; an accelerator
  partition token does not (it names no accelerator), so a partition profile paired with
  `rdma.partitioned` is aligned to the RDMA side alone and is not covered, even though both are
  allocatable in the same container. An RDMA endpoint whose affinity the kernel did not report
  publishes none either.
- **The node's kubelet runs `restricted` or `single-numa-node`.** It is node-level kubelet
  configuration this operator cannot change, and what each of the four policies does with a hint is
  in [the preflight runbook](/gpustack-operator/main/docs/modules/devices/preflight/index.md#reading-the-result), which is also where you
  read the policy a given node runs. No release sets one for you — [how to set it, and how to
  confirm it took](../rdma/operations.md#enabling-numa-alignment-on-the-kubelet).
- **The alignment's unit is the NUMA node.** See
  [the RDMA resource keys](/gpustack-operator/main/docs/modules/rdma/network-topology/index.md#the-rdma-resource-keys)
  for why a shared PCIe switch is finer than a hint can express.

What is not guaranteed is the node itself: the scheduler selects by quantities and cannot see
inside a node's topology, so a node whose totals suffice but whose devices sit on two NUMA nodes is
selected and the container is then refused at admission — and a Pod refused that way is not
rescheduled onto another node.

`device-manager preflight` reports the node's TopologyManager policy as a `topology` section, read
off the running kubelet's own command line and the configuration it names; `unknown` means no
readable source named one, never a guess of the default. See
[Preflight Operations](/gpustack-operator/main/docs/modules/devices/preflight/index.md).

## Requesting through the `Instance` API

An `Instance` expresses the same request families through `spec.resources`, and the controller shapes them
into the keys above:

| Field | Effect |
|---|---|
| `accelerator: "N"` | whole accelerators (exclusive) — may span accelerators |
| `accelerator: "1"`<br/>`acceleratorSlicedMemoryPercentage`<br/>`acceleratorSlicedCoresPercentage` (optional) | one logical slice |
| `accelerator: "1"`<br/>`acceleratorPartitionedProfile: "3g.40gb"` | one hardware partition of that profile |

```yaml
kind: Instance
spec:
  type: gpustack--nvidia-h100-80gb-hbm3-linux-amd64
  resources:
    accelerator: "1"
    acceleratorPartitionedProfile: 3g.40gb
```

The webhook rejects these at admission rather than leaving the Instance Pending:

- a profile the pool does not offer — the message lists the offered set;
- a profile on a manufacturer with no hardware partitioning;
- a profile **and** a slice percentage together (`a hardware partition and a logical slice percentage are
  mutually exclusive`);
- an `accelerator` count other than `1` for a slice or a partition request;
- a whole-accelerator count over the pool's whole-accelerator **capacity** (`status.accelerator.capacity`),
  never over what is free right now: an Instance submitted while every accelerator is held is admitted
  and waits in its queue;
- **a slice percentage against a pool that offers no logical slicing** — a tightening; such a request
  used to be silently reshaped into a whole-accelerator one and served. On an all-partitioned pool the
  message points at `spec.resources.acceleratorPartitionedProfile` instead.

A partition's host CPU and RAM are sized by the profile's share of the accelerator's VRAM, so a `1g`
instance does not ask for a whole accelerator's worth.

## Pre-release breaks

No released version is marked stable, so each of the following is a clean break with **no** translation
layer.

**The old MIG key is gone.** A MIG profile used to be requested through the *logical* family: a
per-profile key built from `<base>.sliced` with a `mig-<profile>` segment appended, alongside
`<base>.sliced: 1`. It is replaced by `<base>.partitioned.<kind>-<profile>` alongside
`<base>.partitioned: 1`.

Nothing recognizes the old one — not the builders, not the parsers, not the node-capacity reconciler,
and deliberately **not even a rejection by name**. Rewrite such a manifest; untouched, it meets two
ordinary failures, neither naming the replacement:

- its per-profile key is an extended resource no node advertises, so it never schedules;
- the `<base>.sliced: 1` beside it is a logical slice with no memory budget, which rule 2 rejects.

> **Why** — legibility would cost a legacy branch in every key path, to serve a request no
> documentation, no `InstanceType` and no example produces.

A development node an earlier build wrote those legacy capacities onto keeps them, because nothing owns
them. List what is left and patch it off — one JSON-patch removal per stale key:

```console
$ kubectl get node <node> -o json | jq -r '.status.capacity | keys[] | select(contains("mig-"))'
$ kubectl patch node <node> --subresource=status --type=json \
    -p '[{"op":"remove","path":"/status/capacity/<escaped-key>"}]'   # escape "/" in the key as "~1"
```

**The allocation annotation's value changed too, so drain a node before upgrading its device manager.**
The Pod annotation `device.gpustack.ai/accelerator.allocated` is now a per-container map; the old flat
shape is not read. Drain the node (or delete its accelerator Pods) before rolling the device-manager
DaemonSet, then let the workloads reschedule.

> **Why** — a Pod carrying the old shape on a restarted device manager **drops out of the ledger**: its
> accelerators read *free* while its containers hold them, and the next opposite-mode Pod can land on an
> occupied one. The rebuild cannot recover — the occupancy is exactly what became unreadable — so it
> logs loudly, naming the Pod.

**`<base>.shared: N` counts accelerators now.** It used to be read as N shares of one accelerator in
some places and N accelerators in others; it is N distinct accelerators on one node, one share on each.
A manifest asking `"3"` for a heavier share of one accelerator now needs a node with three.

`InstanceType.status.acceleratorShared.onceMaxRequest` moved with it: it counts a node's accelerators
with a free share, where it used to sum their shares. Output recorded before the change, in the
[Walkthrough](/gpustack-operator/main/docs/getting-started/walkthrough/index.md) and [NVIDIA MIG Operations](/gpustack-operator/main/docs/modules/devices/nvidia-mig/index.md), still shows the
old reading under `SH` — `10/10` on one free accelerator now reads `1/10`.

## Limitations

- **Media-engine and graphics profile variants are not exposed.** A profile whose name is not a valid
  Kubernetes resource-name segment (the `+me`, `+me.all` and `+gfx` MIG variants) is **excluded** from
  an accelerator's inventory rather than rewritten to something key-safe, so a key always maps back to
  its profile by a plain prefix strip. Those variants cannot be requested.
- **One accelerator per slice or partition request** (rules 2 and 3), for the different reasons given
  above.
- **The whole-accelerator bound is the pool's total, not one node's.** A count above the largest node's
  accelerators but within the pool's total is admitted and then stays queued, because one Pod's
  accelerators come from one node and `InstanceType.status` carries no per-node capacity to refuse it
  with. It applies to an `Instance` and a `ModelDeployment` role alike.
- **Hand-carving a partition outside GPUStack is unsupported on a managed node.** Every node-level key —
  per-profile capacity, partition token health, the admission check — comes from the Pod annotations the
  device plugin writes, and an instance made with `nvidia-smi mig -cgi` produces none.

  The node keeps advertising room it does not have, and unlike a transient over-advertisement this
  **never converges**: it persists until the instance is removed. Placement reads the live hardware and
  will not double-book it, but the accounting above stays wrong. Let GPUStack materialize the instances;
  it reuses any already on an accelerator it manages.

- **Flipping an accelerator's partitioning mode is an operational procedure, not a live switch.** An
  accelerator advertises a family's tokens only while its reported capability backs that family, so a
  flip removes the old tokens rather than marking them unhealthy: there is no continuity across it.

  Drain the accelerator (the hardware refuses the toggle under load anyway), flip the mode with the
  manufacturer's tool, then **restart that node's Device Manager DaemonSet**: the re-detect trigger
  watches the device set and health, not the partitioning mode. Deleting the node's `Devices` object is
  **not** required, since an existing group's capability is rewritten in place.

  Full procedure: [NVIDIA MIG Operations](/gpustack-operator/main/docs/modules/devices/nvidia-mig/index.md#enabling-partitioning-on-a-node) ·
  [T-Head MIG Operations](/gpustack-operator/main/docs/modules/devices/thead-mig/index.md#enabling-partitioning-on-a-node) ·
  [Hygon MIG Operations](/gpustack-operator/main/docs/modules/devices/hygon-mig/index.md#enabling-partitioning-on-a-node).
- **A shared request on a node with fewer free accelerators than it names fails for good.** Each
  accelerator advertises 10 shared tokens, so the scheduler fits `.shared: "2"` on a one-accelerator
  node; the device plugin refuses the two tokens the kubelet picked there rather than grant one short.
  The Pod ends `Failed` with `UnexpectedAdmissionError` and must be recreated.

  A pool's queue holds such a request first
  ([Admission](/gpustack-operator/main/docs/modules/devices/admission/index.md#gate-3--the-per-accelerator-admissioncheck)), so a Pod
  reaches this only by skipping the queue or racing another allocation.
- **The shared card-count pin has two gaps.** Where it misses, a request placed on too few
  accelerators retries for good ([Admission](/gpustack-operator/main/docs/modules/devices/admission/index.md#gate-3--the-per-accelerator-admissioncheck)).
  - *Workloads built from a template.* Kueue builds a batch Job's, a JobSet's or a RayJob's Workload
    from its Pod template before any Pod exists, and the webhook only sees Pods, so none is pinned.
  - *Partitioned accelerators count.* `.count` is the node's total, so a node whose accelerators are
    mostly in a partitioning mode still passes the pin with too few left to share.
- **A non-default `TopologyManager` policy can mis-align a partition.** The Partitioned resource reports
  no NUMA topology (the plugin may not use the accelerator the kubelet aligned to), so under
  `single-numa-node` the CPU and memory providers can settle on one socket while the only accelerator
  with room is on the other. The default `none` is unaffected.

---

**See also** — [NVIDIA MIG Operations](/gpustack-operator/main/docs/modules/devices/nvidia-mig/index.md) (the administrator runbook for an
accelerator's partitioning mode, plus a recorded enable → request → reclaim → disable walkthrough) ·
[Hygon MIG Operations](/gpustack-operator/main/docs/modules/devices/hygon-mig/index.md) (the same runbook for Hygon, whose mode is
node-wide and whose partitioned nodes serve no whole-card or sliced request at all) ·
[T-Head MIG Operations](/gpustack-operator/main/docs/modules/devices/thead-mig/index.md) (the same runbook for T-Head's own
partitioning) · [Admission](/gpustack-operator/main/docs/modules/devices/admission/index.md) (where these keys are checked) ·
[Device Discovery](/gpustack-operator/main/docs/modules/devices/discovery/index.md#the-device-plugin-allocator) (where they are
served) · [RDMA Operations](/gpustack-operator/main/docs/modules/rdma/operations/index.md) (how many RDMA endpoints to ask for beside N
accelerators, and the kubelet policy that aligns the two sides)

**Next** → [Walkthrough](/gpustack-operator/main/docs/getting-started/walkthrough/index.md) — the same requests on a live cluster.
