Documentation
GPUStack Operator
GPUStack Operator discovers hardware, shapes accelerator capacity, and helps Kubernetes place workloads where that capacity is available.
Start here
Architecture
See how discovery, capacity, and admission fit together.
Walkthrough
Follow an accelerator request through a working cluster.
Installation modes
Choose how the operator and its dependencies are deployed.
Capabilities
Heterogeneous Devices
Discover, slice, and allocate accelerators.
RDMA Networking
Give workloads network interfaces alongside accelerators.
Topology Management
Place workloads within the requested network or location domain.
KV Cache
Share inference cache across workloads.
Model Delivery
Fetch and cache model weights before workloads need them.
Model Deployment
Run serving replicas with routing and prefill/decode roles.
Accelerated Instances
Work in an accelerator-backed container over SSH.