GPUStack Operator

GPUStack Operator

GPUStack Operator discovers hardware, shapes accelerator capacity, and helps Kubernetes place workloads where that capacity is available.

Start here

Architecture

See how discovery, capacity, and admission fit together.

Walkthrough

Follow an accelerator request through a working cluster.

Installation modes

Choose how the operator and its dependencies are deployed.

Capabilities

Heterogeneous Devices

Discover, slice, and allocate accelerators.

RDMA Networking

Give workloads network interfaces alongside accelerators.

Topology Management

Place workloads within the requested network or location domain.

KV Cache

Share inference cache across workloads.

Model Delivery

Fetch and cache model weights before workloads need them.

Model Deployment

Run serving replicas with routing and prefill/decode roles.

Accelerated Instances

Work in an accelerator-backed container over SSH.