A staging run of the studio image build died with ENOSPC during the
Studio venv install: the hosted runners' default free space does not
fit the base image plus buildkit state plus the Studio layer. Drop all
unused preinstalled toolchains and the runner's preloaded docker
images in both build jobs.
- Dockerfile: lift numba past vllm's 0.61.2 pin after the numpy>=2.4
re-upgrade; 0.61.2 refuses numpy 2.3+ at import time and the stack
cannot move numpy down. Verified numba 0.65 + numpy 2.4.6 + vllm
import cleanly together.
- docker-publish.yml: resolve UNSLOTH_ZOO_REF in a step that mirrors
the pushed tag only when the tag exists in unsloth-zoo (the zoo
currently cuts no tags, so blind mirroring broke every tag publish);
falls back to main.
- Dockerfile.studio: Studio venv stays on cu128 for arm64 too, matching
the base venv (cu130 wheels would lift the driver floor to 580+), and
gets the same NVRTC cu13 swap for DGX Spark / GB10 sm_121 support.
- docker_confirm.sh: do not drop to CPU mode when docker info lacks a
nvidia runtime entry; CDI installs and Docker Desktop WSL2 expose
GPUs without one. The phase 3 --gpus probe is now the authority.
- docker_confirm.ps1: GPU selector built as an args array; comma device
lists get version-aware CSV quoting (native arg passing changed in
PowerShell 7.3).
- studio_launch.sh: no fixed Jupyter default password; generate a
random one and print it when JUPYTER_PASSWORD is unset. Env snapshot
for SSH sessions now written via shlex.quote instead of sed so
values with quotes or command substitution cannot break or inject
into /etc/profile.d.
- install.ps1: honour UNSLOTH_TORCH_INDEX_FAMILY like install.sh does.
entrypoint.sh: a container started without a GPU request has no
nvidia-smi at all (the toolkit injects it), so the old check 1 reported
'CUDA runtime in this image is broken, re-pull' for the most common user
error. Fold the missing-binary case into the actionable 'No GPU visible'
message and document the CPU-only option (UNSLOTH_ALLOW_CPU=1).
run.sh / test_locally.sh: guard empty-array expansions with the
${arr[@]+...} form; bash 3.2 (macOS /bin/bash) treats "${empty[@]}"
as unbound under set -u, which broke the documented macOS CPU path.
studio_launch.sh: exclude *_TOKEN, *_API_KEY, *_PASSWORD, *_SECRET,
*_LICENSE from the env snapshot written for SSH sessions; secrets stay
in process env only, never on disk.
supervisord.conf / Dockerfile.studio: pin HOME=/root for the studio and
jupyter programs (jupyter would silently fall back to token auth if HOME
were unset), default JUPYTER_PORT and UNSLOTH_ENABLE_SSHD at the image
level so a direct supervisord invocation cannot hit a bad %(ENV_*)s
expansion, and document the root-services decision (non-root parity with
the previous production image is a tracked follow-up).
docker_confirm.ps1: mirror the bash script's GPU selector translation so
GPUS=0 / 0,1 select devices instead of silently using all GPUs.
docker-publish.yml: studio cache scope moves to mode=min; a mode=max
cache of a ~24GB image would evict everything else in the 10GB GHA
quota for no hit-rate gain.
The first bake attempt reused studio/install_llama_prebuilt.py, but that
resolver selects a bundle for the CURRENT host: on a GPU build host
/proc/driver/nvidia leaks into docker build and the resolver goes down the
CUDA path with no readable driver runtime (chosen_asset=none, exit 2),
while on a GPU-less CI runner it would resolve a CPU bundle instead. Both
violate the image's build-host-independence rule.
fetch_llama_prebuilt.py pins by build target only: amd64 takes the
linux-x64-cuda12-portable bundle, arm64 the linux-arm64-cuda13-portable
bundle (DGX Spark / Grace), both sha256-verified against the release's
llama-prebuilt-sha256.json. convert_hf_to_gguf.py plus gguf-py/ are
hydrated from the same release's source tarball so the converter's tensor
mappings match the binaries, mirroring unsloth_zoo's
_hydrate_converter_sources layout. LLAMA_PREBUILT_TAG build-arg overrides
the pinned release.
docker_confirm.sh: one-command confirmation script for any machine
(Linux / WSL2 / macOS) following the staging confirm-script conventions:
host + docker + GPU detection with CPU-mode auto-fallback, image pulls,
in-container torch.cuda check, 5-step LoRA training smoke, baked llama.cpp
verification, full-image boot probing Studio /api/health and JupyterLab
/api, PASS/WARN/FAIL summary with RESULT line.
Base image (docker/Dockerfile):
- Install JupyterLab + notebook + ipywidgets in a separate pure-Python uv
pass so the cu128 pin set cannot move; EXPOSE 8888.
- Bake the prebuilt llama.cpp bundle into /opt/unsloth/llama.cpp at the
runtime stage using studio/install_llama_prebuilt.py from the same
UNSLOTH_REF (sha256-verified, portable CUDA bundle since the build host
has no GPU; arm64 resolves the linux-arm64-cuda13 bundle). Export
UNSLOTH_LLAMA_CPP_PATH so unsloth_zoo's save_pretrained_gguf finds it
and never reaches the interactive install prompt or a source build.
- Optional github_token BuildKit secret for the resolver's API calls on
shared CI runner IPs.
Entrypoint: UNSLOTH_ALLOW_CPU=1 degrades a missing GPU to a warning so
Docker Desktop on macOS / Windows-without-WSL2-GPU and plain CPU hosts can
run Jupyter, GGUF tooling and Studio chat; with a GPU visible the normal
pre-flight still runs.
Full image (docker/Dockerfile.studio): now mirrors the production service
set under supervisord - Studio on 8000, JupyterLab on 8888, key-only sshd
on 22 (enabled only when PUBLIC_KEY/SSH_KEY is set). Points Studio's
llama.cpp dir at the baked bundle to skip a duplicate download, accepts
any git ref via fetch+checkout (CI passes commit SHAs), and FROMs a
digest-pinned BASE_IMAGE.
Publish workflow: base image moves to the base-* tag namespace; new
build-studio/merge-studio jobs publish the full image as :latest (hub
parity with the previous production image, which shipped Studio + Jupyter
+ SSH). Studio builds FROM the exact base manifest digest published by the
same run. GPU smoke job now also boots the full image and probes Studio
/api/health and Jupyter /api.
run.sh: UNSLOTH_GPUS=none, UNSLOTH_ALLOW_CPU forwarding, UNSLOTH_PORTS
publish flags, CPU-mode and Jupyter usage examples.
Previously, UNSLOTH_REF was pinned to the triggering tag (e.g. v2026.5.8)
but UNSLOTH_ZOO_REF was hardcoded to main. That made release-tag images
ship a zoo from whatever was on main at build time rather than the zoo
release cut alongside that unsloth tag, so a 2026.5.8 tag image could
install a zoo from days later. Mirror the tag branch of UNSLOTH_REF.
SHA-based branch pushes still fall through to main because the unsloth
SHA does not exist in the unsloth-zoo repo. workflow_dispatch still
honours the unsloth_zoo_ref input.
- docker-publish.yml: add `concurrency: docker-publish-${{ github.ref }}`
(cancel-in-progress: false) so two pushes to main never race the
`:latest` retag. Don't cancel in-progress runs -- the build is
expensive and a half-built image left around is worse than a stale
:latest for a few minutes.
- Dockerfile: soften the requirements.lock.txt comment. `pip freeze`
captures versions but not wheel hashes, and several deps resolve
from VCS / nightly indexes that float, so the file is not actually
byte-reproducible. Reword as an "informational pin record".
1. Stop leaking secrets via docker run -e VAR=VALUE argv (run.sh, test_locally.sh)
`docker run ... -e HF_TOKEN=hf_xxx ...` puts the literal token in
the docker CLI's argv, which is visible to any user on the host
via `ps auxe` / `/proc/<pid>/cmdline` for the lifetime of the
process. Switch to the dash-only form `-e HF_TOKEN`, which tells
docker to read the value from the parent shell's env and never
appears in argv. Same fix for WANDB_API_KEY and UNSLOTH_LICENSE in
run.sh and HF_TOKEN in test_locally.sh.
2. Stop stripping numpy/tests/ in the runtime layer (Dockerfile)
The Dockerfile explicitly upgrades numpy >= 2.4 because numpy 2.2.6
shipped a stripped wheel where `from numpy._core.tests._natype
import pd_NA` fails. Numpy 2.4 restores `numpy/_core/tests/`, then
the existing `find ${VENV} -name tests -exec rm -rf {} +` deleted
it again -- re-introducing the same broken-import state on the
deployed image (the build-time verification at line 220 runs
BEFORE the strip so it passed). Whitelist numpy's tests directories
from the strip; keep stripping the rest.
3. Align :latest tag gate between merge and smoke-test jobs
(.github/workflows/docker-publish.yml)
merge job: enable = is-default-branch AND unsloth_ref == ''
smoke-test job: enable = is_default_branch only
On `workflow_dispatch`, `github.event.inputs.unsloth_ref` defaults to
"main" (not ""), so the merge step skipped `:latest` but the smoke
step still emitted `:latest` as tags[0]. The smoke step then
`docker pull`-ed a prior `:latest` from Docker Hub instead of the
image just merged -- so the smoke test verified the OLD image, not
the new one. Copy the merge step's exact `enable=` expression into
the smoke-test step so the two stay byte-identical and a workflow_
dispatch run validates whatever was actually merged.
Round-2 of the 12-persona reviewer.py pass found 17 issues. Address the
P1s + the regression-class P2s in this commit; the remaining nits are
left for a follow-up cleanup pass.
1. unsloth/_gpu_init.py: the `NVIDIA_VISIBLE_DEVICES in os.environ` check
triggered for every NVIDIA-runtime container including `--gpus all`
(NVIDIA_VISIBLE_DEVICES=all is the default). Gate strictly on a
non-special device list. Also drop the precondition that the env var
was absent: if the user already pinned TORCHINDUCTOR_COMPILE_THREADS=1
we should still plant the UNSLOTH_FORCE_SINGLE_COMPILE_WORKER sentinel
so the zoo-side patch knows to preserve the forcing.
2. unsloth/_gpu_init.py: after the post-`import unsloth_zoo` reassertion,
monkey-patch `unsloth_zoo.temporary_patches.common.determine_compile_threads`
to return 1, so any later `torch.compile` call that rebuilds the
options dict still sees the single-worker forcing even if a downstream
patch_torch_compile pops the env var again.
3. docker/Dockerfile: torchaudio==2.11.0 mismatched the torch==2.10.0
release pairing; pin to 2.10.0 so the ABI is correct and the audio
stack matches torch/cu128.
4. docker/Dockerfile: drop `12.1+PTX` from TORCH_CUDA_ARCH_LIST. The
cu128 toolkit compiler does not know about compute_121; the trailing
PTX entry forced nvcc to emit a `sm_121` gencode that breaks any
in-container source builds.
5. docker/smoke_test.py: the device-capability floor said `cap[0] < 8`,
rejecting Turing (sm_75) while the Dockerfile + entrypoint advertise
sm_75 as supported. Lower the smoke floor to sm_75 and print a hint
that bf16 is not available on Turing.
6. docker/run.sh: `-it` is unconditional; CI / non-TTY invocations died
with "the input device is not a TTY". Probe `[ -t 0 ] && [ -t 1 ]`
first. Also remove `set -x` which echoed the forwarded HF_TOKEN /
WANDB_API_KEY / UNSLOTH_LICENSE values to stdout.
7. docker/test_locally.sh: `-e HF_TOKEN="${HF_TOKEN:-}"` either pasted
the secret verbatim into the process arg list or shadowed any
in-container value with an empty string. Forward conditionally.
8. .github/workflows/docker-publish.yml: gate `latest` on default branch
AND on `unsloth_ref` not being overridden via workflow_dispatch.
Otherwise a maintainer testing a feature SHA from main could overwrite
`:latest` with non-main source.
9. docker/Dockerfile.studio: add an `UNSLOTH_STUDIO_REF` build-arg so
the Studio companion image is pinned to a known unsloth ref instead
of cloning `main` whenever it builds.
Round-trip with the reviewer.py 12-persona pass surfaced four real
issues. Fix all four in this PR so the new Docker release path is
self-consistent.
1. docker/smoke_test.py used `import xformers` unconditionally, which
guarantees a failure on arm64 (built with `[huggingface]` extras to
skip xformers since it has no aarch64 cu128 wheel). Wrap the import
in try/except so the same smoke script validates both arches.
2. unsloth/_gpu_init.py forced `TORCHINDUCTOR_COMPILE_THREADS=1` before
`import unsloth_zoo`, but `patch_torch_compile` in unsloth_zoo main
pops that env var in non-debug mode. After unsloth_zoo init the
guard was effectively undone, so cgroup-pinned `docker --gpus
'"device=N"'` containers still spawned the Inductor subprocess pool
that cannot enumerate the GPU. Set `torch._inductor.config.
compile_threads = 1` directly post-import-torch and re-populate the
env var so `determine_compile_threads()` in the zoo options dict
also returns 1, regardless of whether the zoo-side fix from PR #694
has shipped yet.
3. docker-publish.yml UNSLOTH_REF build-arg defaulted to `'main'` for
tag pushes and scheduled runs, so a `v1.2.3` release image would
contain whatever `main` happened to be at build time, not v1.2.3.
Pick the tag's `github.ref_name` for tag events and `github.sha`
for branch/schedule events.
4. The smoke-test job pulled `:latest` regardless of which tag the
merge job had just published, so tag/schedule/sha publishes were
never actually validated. Re-run docker/metadata-action with the
same config the merge job used, then smoke-test the first tag from
its output.
All four changes are gated and backwards-compatible.
GitHub announced free linux/arm64 hosted runners for public repos (GA Aug
2025) under labels `ubuntu-24.04-arm` / `ubuntu-22.04-arm`. Switching the
arm64 leg from QEMU-on-amd64 to a native arm64 matrix runner is ~3x
faster and avoids QEMU's occasional flakiness on long cu128 installs.
The workflow now:
* builds amd64 and arm64 in parallel on their native runners,
pushing each as a single-arch image *by digest* (no tag)
* stitches both digests into one multi-platform manifest in a
follow-up `merge` job, using `docker buildx imagetools create`
* keeps a separate buildx cache scope per platform to avoid
cross-arch cache collisions
Smoke-test job now needs `merge` (was `build`) so it only runs once the
final manifest is published.
Dockerfile header: replace the speculative aarch64 SASS list with the
verified one from pytorch/pytorch v2.10.0 .ci/manywheel/build_cuda.sh
(8.0;9.0;10.0;12.0 on aarch64), and note that sm_120 is forward-compatible
to sm_121 per PyTorch maintainers -- which is what makes DGX Spark work
without an explicit sm_121 SASS section in the wheel.
setup_qemu.sh / test_locally.sh --platform stay in place: they're for
the local-dev path on x86_64 boxes that don't have arm64 hardware.
Make the docker image multi-arch so DGX Spark (GB10, sm_121, aarch64) and
the Grace-Hopper / Grace-Blackwell SoCs (GH200 arm64, GB200 arm64) pull a
natively-built arm64 child from the same manifest. Runtime emulation is
NOT involved -- QEMU is used only for the cross-compile step on x86_64
CI runners; consumers on aarch64 hosts get a normal arm64 image and CUDA
works as on any other host.
Dockerfile:
* ARG TARGETARCH; switch unsloth extras between cu128-ampere-torch2100
(amd64, with xformers) and huggingface (arm64, no xformers -- there
is no cu128 aarch64 xformers wheel as of 0.0.34, so we fall back to
Unsloth's native SDPA path; ~5-10% slowdown but functionally complete).
* Build-time torch._C._cuda_getArchFlags() assertion: amd64 still
requires sm_120, arm64 accepts sm_120 or sm_121.
* Same TORCH_CUDA_ARCH_LIST on both arches; nvcc emits whatever's listed.
docker/setup_qemu.sh (new):
One-time host setup -- registers binfmt_misc handlers via
tonistiigi/binfmt and creates a 'unsloth-multiarch' docker-container
buildx builder. Required only on x86_64 build hosts targeting arm64.
docker/test_locally.sh:
--platform amd64|arm64 flag. Cross-builds verify QEMU is registered,
then build through the in-image arch-flags assertion. Smoke + notebook
blocks auto-skip when image arch != host arch (CUDA cannot run under
user-space QEMU + nvidia-container-toolkit cannot bridge a QEMU guest
to a real GPU).
.github/workflows/docker-publish.yml:
platforms: linux/amd64,linux/arm64 (single manifest, two children).
Timeout bumped 60 -> 150 min for the slower arm64-under-QEMU leg.
docker/setup-qemu-action@v3 with platforms: arm64 (was implicit before).
Adds a multi-stage Dockerfile producing an image that works on Ampere through
Blackwell (sm_80 through sm_120: A100, RTX 30/40, H100, B100/B200, RTX 50-series,
RTX 6000 Pro Blackwell). The build itself requires no GPU at all and runs on a
free GitHub-hosted ubuntu-latest runner.
How the GPU-less build works:
1. cu128 PyTorch wheels are fat binaries. torch._C._cuda_getArchFlags() returns
'sm_70 sm_75 sm_80 sm_86 sm_90 sm_100 sm_120' regardless of which GPU
compiled the image, because the wheels are cross-compiled upstream by the
PyTorch team.
2. All deps resolve in a single uv pip install pass with explicit pins
(torch==2.10.0, --extra-index-url cu128, no --torch-backend=auto, no
install.sh). This prevents the silent cu cascade where bitsandbytes'
transitive cuda-toolkit==13 dep upgrades torch to 2.12+cu130 in a later
resolver pass, leaving xformers and other cu128 wheels stranded.
3. Build-time verification uses package metadata (importlib.metadata.version)
and the raw torch._C._cuda_getArchFlags() accessor. We deliberately avoid
import unsloth at build time because unsloth.__init__ calls
torch.cuda.get_device_properties(0), which requires an actual CUDA device
and is not bypassable. Import-time correctness is exercised at deploy time
by smoke_test.py with --gpus all.
4. UNSLOTH_COMPILE_DISABLE=1 and CUDA_VISIBLE_DEVICES="" during the build stage
prevent any code path from JIT-compiling kernels for the build host's
compute capability and baking the resulting cache into the image. The
deploy GPU produces its own cache on first use.
Other notes:
- --index-strategy unsafe-best-match is needed because the PyTorch wheel index
serves an old requests==2.28.1 that conflicts with datasets>=2.32.2, which
the default first-index-wins strategy rejects.
- Extra is cu128-ampere-torch2100 (ampere precedes the torch version in the
pyproject ordering).
- No flash-attn in the base image. FA3 is hard-refused on Blackwell upstream
and unsloth gracefully falls back to xformers + SDPA. Users on Ampere /
Ada / Hopper who want FA2 can pip install flash-attn on top.
- Two stages: nvidia/cuda:12.8.1-cudnn-devel-ubuntu24.04 for the build,
-cudnn-runtime for the deploy image. No nvcc in the published image.
- A lockfile is emitted at /opt/unsloth-venv/requirements.lock.txt inside
the image and can be extracted with docker/freeze.sh for byte-identical
rebuilds even after PyPI moves on.
CI workflow .github/workflows/docker-publish.yml:
- Builds on ubuntu-latest on every push to main, every tag, weekly via cron,
and manually via workflow_dispatch. Pushes to docker.io/unsloth/unsloth
with cache via type=gha.
- Optional smoke-test job runs on a self-hosted GPU runner if vars.HAS_GPU_RUNNER
is set; skipped otherwise. End-to-end verification on sm_120 hardware is a
nice-to-have, not a publish blocker.
Validation:
- Install path validated on a B200 host with CUDA_VISIBLE_DEVICES="" set
(simulating the GPU-less CI runner): torch 2.10.0+cu128 holds, xformers
0.0.34, bitsandbytes 0.49.2, triton 3.6.0, transformers 5.5.0, trl 0.24.0,
peft 0.19.1, accelerate 1.13.0. Arch flags include sm_100 and sm_120.
- Runtime path validated end-to-end on B200: smoke_test.py imports unsloth,
loads Llama-3.2-1B-Instruct-bnb-4bit in 4-bit, completes 5 LoRA steps with
loss decreasing 4.11 -> 3.75. xformers fallback active as designed.
Files:
- docker/Dockerfile multi-stage cu128 build
- docker/build.sh local build wrapper
- docker/freeze.sh extract lockfile from a built image
- docker/smoke_test.py runtime verification, run with --gpus all
- docker/.dockerignore
- .github/workflows/docker-publish.yml