GPUStack Operator

Architecture

GPUStack Operator manages accelerator-backed workloads on Kubernetes. It discovers devices, turns node capacity into schedulable pools, and runs GPU Instances and model deployments against those pools. Optional model delivery, topology placement and shared KV cache support inference workloads that need more than a single GPU.

Contents

Components

The gpustack-operator binary provides these subcommands:

Component Deployment Role
worker Control-plane Deployment Serves the public APIs and manages workloads, queues and supporting resources.
device-manager One DaemonSet per manufacturer Discovers accelerators and network interfaces, reports available capacity and allocates devices to containers.
model-manager Node DaemonSet Caches model weights and mounts them into workloads using node delivery.
worker-gateway Installed separately Combines InstanceTypes and capacity from multiple clusters into a fleet view.

The chart also includes Node Feature Discovery (NFD) , Kueue and two CSI drivers. Topograph is optional and remains disabled until an administrator selects a provider. See Installation Modes for the deployment choices and Internals for contributor details.

Workload flow

GPU Instances provide containers with optional SSH access. Model deployments run inference engines and, where configured, routers or prefill/decode groups. Both use the same device capacity and admission chain. Model delivery and KV cache are optional services alongside that chain.

Rendering diagram…

Mermaid source
mermaid
flowchart TB
    REQUEST["GPU Instances<br/>ModelDeployment / Pods"]
    WORKER["Worker"]
    DEVICES["Devices Manager<br/>Accelerators / RDMA"]
    TOPOLOGY["Topology Aware"]
    QUEUES["Kueue<br/>Pools and queues"]
    ARTIFACT["Model Delivery<br/>ModelArtifact"]
    MODELS["Model Manager<br/>Node cache"]
    CACHE["KV Cache<br/>Backends and pools"]
    STORE["Cache service"]
    PODS["Workload Pods<br/>SSH / inference"]

    REQUEST --> WORKER
    WORKER --> QUEUES
    DEVICES -- "capacity" --> QUEUES
    TOPOLOGY -- "placement" --> QUEUES
    QUEUES -- "admission" --> PODS
    DEVICES -- "device access" --> PODS
    ARTIFACT -- "node delivery" --> MODELS
    MODELS -- "weights" --> PODS
    ARTIFACT -- "Pod / claim delivery" --> PODS
    CACHE --> STORE
    STORE -. "optional KV reuse" .-> PODS

The diagram shows the main dependencies. It does not require every workload to use a model artifact, RDMA, a topology source or a shared cache. An ordinary Pod can request device resources directly.

Device scheduling

Device scheduling follows this chain:

  1. NFD identifies node hardware and labels the nodes.
  2. The Device Manager discovers accelerators and network interfaces and maintains the Devices inventory and allocation record.
  3. The Worker derives schedulable CPU and accelerator capacity from that inventory.
  4. The Worker creates InstanceTypes and Kueue resources. Each pool has an isolated queue; topology profiles constrain placement, and admission checks verify whether individual accelerators fit.

Device Discovery covers node labeling and device discovery. Scheduling Chain covers capacity and queues. RDMA Operations explains network resource requests, while Topology-Aware Scheduling explains placement domains.

Model delivery and KV cache

A ModelArtifact describes model weights from a model repository, an image or a persistent-volume claim. A workload can download weights in its own Pod, use an existing claim, or mount weights from an eligible node’s verified cache. Node delivery is handled by the Model Manager; it is separate from device admission. See Model Delivery .

A shared KV cache stores reusable inference state. A backend runs the cache service or points to an external one; pools and namespace bindings provide access and quota grants. Compatible workloads receive the cache configuration when their Pods are created. This service is optional, and its quota is separate from accelerator quota. See KV Cache .

Model Deployment brings these services together with engine, router and replica configuration. GPU Instances support interactive containers and SSH access using the same accelerator allocation.

Request admission

The Pod webhook validates an accelerator request before Kueue reserves quota. Admission checks confirm that devices can satisfy it, and the scheduler, kubelet and allocator complete the allocation. Admission owns the gate sequence and its failure behavior.

InstanceType status reports the remaining capacity as workloads allocate and release devices. Use kubectl get instancetype -w to watch those changes. The Walkthrough shows the request and its resulting Kubernetes resources.

Vocabulary

Term Meaning
Pool Nodes with compatible CPU, accelerator, operating system and architecture identities, served by an isolated queue.
InstanceType The resource configuration and available capacity offered by a pool.
Allocation mode Whole accelerators, shared accelerators, logical slices or hardware partitions.
Credits Quota accounting units .
Four capacity views Available allocation modes reported by an InstanceType.
Devices ledger The per-node inventory and record of accelerator allocations.
LocalQueue The namespace queue through which a workload enters its pool.

The hardware terms Device, Accelerator and Resource are defined in Device Discovery .