Model Artifact API
GPUStack server and consoles use worker.gpustack.ai/v1 to manage model artifacts and inspect
node caches. A node’s download progress is stored on
thresholds
; progress answers between them, on request, and
stores nothing.
Contents
- Resources
- The progress subresource
- Authorization
- Capability map for GPUStack server
- Requirements and limits
Resources
| Resource | Scope | Verbs | Printer columns |
|---|---|---|---|
modelartifacts.v1.worker.gpustack.ai |
namespaced | create, get, list, watch, update, patch, delete; subresource progress |
Name, Source, Revision (12 characters), Size, Ready, Downloading, Resolved; Age with -o wide |
nodemodelstores.v1.worker.gpustack.ai |
cluster | get, list, watch, delete |
Name, Ready, Used, Models, Downloading; Age with -o wide |
kubectl get modelartifacts.v1.worker.gpustack.ai -n team-a
kubectl get nodemodelstores.v1.worker.gpustack.aiworker.gpustack.ai/v1 is the public API. A bare kubectl get nms or
kubectl get modelartifacts.worker.gpustack.ai uses it.
- ModelArtifact exposes no writable
statussubresource. The worker owns its status, which is used to authorize model mounts. Updating the main resource preserves the stored status. - NodeModelStore supports reads and deletion. The worker manages its
spec, and each node’s plugin reportsstatus. Deletion supports Kubernetes garbage collection; the worker recreates the object while its node runs the plugin. - Resource fields are described in Model Artifact and Node Model Store .
The progress subresource
kubectl get --raw "/apis/worker.gpustack.ai/v1/namespaces/<ns>/modelartifacts/<name>/progress"kind: ModelArtifactProgress
apiVersion: worker.gpustack.ai/v1
metadata:
name: qwen-72b
namespace: team-a
timestamp: "2026-09-26T00:00:00Z" # when it was computed
manifestDigest: sha256:...
sizeBytes: 145424101604
ready: 1 # nodes holding the content published
downloading: 2 # nodes downloading it
failed: 1 # nodes waiting to retry a failed attempt
downloadingPercent: 42 # the downloading nodes' mean, whole percent
downloadingBytes: 123482472448 # what the downloading nodes hold, summed
live: 2 # downloading nodes read from their plugin for this answer
failureReasons:
- reason: SourceUnavailable
count: 1- It is computed on each request and never written. The worker reads the artifact and the
NodeModelStores listing its digest; for each node downloading it, it reads that node’s plugin for its running downloads. The request is bounded at 13 seconds. Live reads cover at most 64 downloading nodes; a node past that cap, or one that does not answer within the request’s budget, contributes the bytes it last wrote and is not counted inlive. - It names no node. It is namespaced and tenants read it; node names would give every tenant the cluster’s topology. Read NodeModelStore for the node-level details.
- The mean covers the downloading nodes only, each downloading a whole copy; a node starting a
download does not pull the ready ones down.
downloadingPercentis absent while none downloads. - Nothing to count answers zeros and a
reason: the artifact is not resolved yet, or its claim source is mounted from its volume and never downloaded. - The counts describe the content. Artifacts with the same digest, in any namespace, see the same nodes; only numbers cross.
Authorization
The aggregated API server authorizes each request with Kubernetes RBAC through delegated
authorization: get on modelartifacts/progress in the artifact’s namespace. get modelartifacts
alone does not grant it. The chart grants tenants nothing in this group; an administrator gives a
tenant’s subjects a Role such as:
apiVersion: rbac.authorization.k8s.io/v1
kind: Role
metadata:
name: model-artifact-reader
namespace: team-a
rules:
- apiGroups:
- worker.gpustack.ai
resources:
- modelartifacts
- modelartifacts/progress
verbs:
- get
- list
- watchA subject with that Role in team-a reads team-a’s progress and is refused another namespace’s.
nodemodelstores is cluster-scoped: grant it with a ClusterRole to administrators and GPUStack
server only.
Capability map for GPUStack server
The table matches each GPUStack server capability to the API surface that covers it; the gap column names what is missing and who owns it.
| GPUStack server | This API | Gap and owner |
|---|---|---|
| A source: Hugging Face repository, ModelScope model, local path | ModelArtifact.spec.source: huggingFace or modelScope; a local path is a persistentVolumeClaim source |
none |
One file of a repository (huggingface_filename, a GGUF) |
allowPatterns with a single <file> entry |
none |
A model file per worker, its state and state_message |
the node’s status.models[] entry: state, reason, message |
none |
download_progress and size |
the entry’s downloadedBytes and sizeBytes; live through progress; the artifact’s status.nodes |
none |
resolved_paths on the worker |
the fixed mount path in the consumer’s container; host paths are never exposed | by design |
local_dir |
none: the plugin owns the cache layout | by design |
| Listing a worker’s model files | NodeModelStore get, list, watch | none |
| Download to a worker ahead of use | a ModelPrefetch naming nodes | none |
Delete with cleanup_on_delete |
delete the prefetch; collection after the last reference and its grace | none |
reset (retry now) |
automatic backoff to retryTime |
not planned |
| One row per source per worker | one entry per digest per node, shared by artifacts with the same content | none |
| Prefer workers holding the files | placement preference , compute first | none |
| LoRA and draft-model files | not in this batch | later |
| Tenant scope | the artifact’s namespace; the node store names no tenant | none |
Requirements and limits
- The aggregated API,
apiregistration.k8s.io/v1, and delegated authentication and authorization (TokenReview,SubjectAccessReview): available on every version the chart installs on. - Live bytes need the plugin’s Pod reachable from the worker on its HTTPS port. A NetworkPolicy
that blocks it leaves the stored values, and
livesays how many nodes answered. - No watch on
progress: it answers one request. Watch NodeModelStores for changes.