Model Image Source
A ModelArtifact can deliver weights from an OCI image pinned by digest. Kubernetes mounts the
image read-only for the workload; packaging and pulling it adds disk and registry costs.
Contents
- The source and its digest contract
- Building the image
- Delivery
- Versions and prerequisites
- Disk cost
- Image GC
- Registry mirrors and private registries
The source and its digest contract
spec:
source:
image:
reference: registry.example.com/team/qwen@sha256:669ed7b1...48 # digest-pinned, requiredThe reference must pin a digest. A tag is mutable, so one artifact could deliver different weights on different pulls, and the identity a frozen reference pins would be nothing; admission refuses a tag with that reason. Patterns are refused too: an image is mounted whole, so there is no listing to select from.
Why — the digest pins the image’s manifest bytes. It is not the manifest digest a Hugging Face source resolves to, and
status.resolvedstays empty for an image source, as for a claim: the operator never reads the registry, so the reference in the immutable spec is the only record of what the artifact delivers, and the KV reuse identity falls back to the artifact’s UID. Two artifacts naming one image never share KV blocks; that is safe, only not deduplicated.
Creation is refused on an apiserver older than 1.35, where the image-volume field would be
dropped from the Pod silently; the refusal names the floor and the alternatives. On a supported
apiserver the artifact resolves on its first pass with no network: Resolved=True, resolved
carrying only resolvedTime, no nodes aggregation, no revalidation.
Building the image
The build’s contract: the weights live at the image’s root, because delivery mounts the root
whole and offers no sub-path (a sub-path would raise the containerd floor; see
Versions
). A recipe measured end to end, from
Qwen/Qwen2.5-7B-Instruct at commit a09a3545…, one safetensors shard per layer:
for f in *.safetensors tokenizer.json tokenizer_config.json config.json generation_config.json; do
oras push "$IMAGE" --artifact-type application/vnd.gpustack.weights "$f:application/octet-stream"
doneThe mount is a read-only overlay of the layer snapshots, and every file the engine reads matches the Hub copy byte for byte, measured sha256-equal to the LFS oids on Kubernetes 1.35.7 / containerd 2.2.6. That equality is a property of the build, which this recipe produces; the operator does not verify it.
Delivery
An image source always delivers Image, whatever model-artifact-delivery-mode says; the Setting
governs hub sources (Hugging Face and ModelScope). A ModelDeployment renders one image volume per role, mounted
read-only at /var/lib/gpustack/model, and nothing else: no cache emptyDir, no HF_* or proxy
environment, no --revision, no ephemeral-storage raise.
The served path and the --served-model-name rule are the claim’s, and the engine-argument
refusals apply unchanged. A take-over role gets the mount like a claim’s.
An Instance model volume mounts the image read-only and whole, and needs no CSIDriver; the
node’s plugin plays no part. An Instance pinned to a node checks that node before rendering: kubelet
below 1.35, containerd below 2.1, a runtime that is not containerd, or a version that cannot be read
holds the Pod with a phase message naming the node and the floor. An unpinned Instance cannot know,
and waits like a ModelDeployment does.
WeightsReady follows the claim path: WeightsNotMounted until every Pod has started. A failed or
slow pull keeps that reason; kubelet’s own event on the Pod names the pull error, and that event is
the diagnostic path.
Versions and prerequisites
| Component | Floor | Why |
|---|---|---|
| apiserver | 1.35 | the ImageVolume feature gate is beta on by default from 1.35, GA in 1.36; below that the gate must be opened on the apiserver, which this version does not opt into, and a 1.32 or older apiserver refuses the source at admission |
| kubelet | 1.35 | it mounts the volume; the same gate line applies, checked per pinned node |
| containerd | 2.1 | the runtime that mounts image volumes; sub-path mounts would need 2.2, which is why none are offered |
PSA is not an obstacle: on an apiserver ≥ 1.33, Restricted admits image volumes
(kubernetes#130394, fixed in 1.33, not backported); on ≤ 1.32 Restricted refuses them, and
enforce-version makes no difference either way. Baseline admits wherever the apiserver knows the
field.
A private registry is reached with the role’s or Instance’s existing imagePullSecrets; the image
volume pull assembles credentials the same way a container image pull does.
Disk cost
A pulled image holds its bytes twice while discard_unpacked_layers is false (the containerd
default): the compressed blobs in the content store and the unpacked snapshots. Measured on
containerd 2.2.6 with a 15,242,807,270-byte model:
| Image shape | Content store | Snapshots | Total | vs raw weights |
|---|---|---|---|---|
| gzip-compressed layers | +12,074,816,089 B | +15,242,807,270 B | ~27.3 GB | 1.79× |
| uncompressed layers | +15,242,807,270 B | +15,242,807,270 B | ~30.5 GB | 2.00× |
Uncompressed layers save decompression time, not disk. Kubelet’s own accounting does not see the
doubling: the CRI image size is the compressed size, ~12.07 GB where the disk really holds ~27.3
GB (~2.26× apart). Size a node’s disk from the table, never from kubectl’s image sizes.
A node image that ships discard_unpacked_layers = true (the local kind images do) keeps only the
blobs and does not pay the snapshot copy; do not take that as the fleet default.
Image GC
Three rules, verified on 1.35.7 / 2.2.6:
-
In use, not collected. While any running container mounts the image, kubelet’s image GC passes over it (the protection needs kubelet ≥ 1.31 and containerd ≥ 2.1). A completed Pod counts for nothing.
-
Released, collected in the next round. After the last referencing Pod goes, the image can be collected on the next GC cycle (5 minutes; measured release→collection 4m2s).
imageMinimumGCAgedoes not protect it: kubelet compares against the image’s first-detected time, not its release.Scaling a deployment to zero and back can re-pull a dozen GB (a first pull measured 12m43s on a 2-vCPU node at ~20 MB/s).
-
The 85% cascade. A model image that pushes the disk past
imageGCHighThresholdPercent(default 85) drags every other unused image on the node into the same rounds (measured: 18 unrelated images collected alongside).
The only retention is another reference: a resident Pod mounting the image keeps it. A deployment’s replicas are themselves that reference while they run.
Registry mirrors and private registries
The pull is a normal containerd pull, so it rides whatever the node’s registry configuration resolves: a Harbor pull-through cache, a Spegel peer-to-peer mirror, or Dragonfly as a configured containerd mirror. None of these need operator support, and none are measured from here; Dragonfly as an operator-integrated prefetcher is future work.
Two images sharing layer digests share snapshots and the download: measured, a second image whose layers overlapped pulled in 112 ms and added 11.8 MB where a cold node paid 2m10s and 7.89 GB.