- Dockerfile: lift numba past vllm's 0.61.2 pin after the numpy>=2.4 re-upgrade; 0.61.2 refuses numpy 2.3+ at import time and the stack cannot move numpy down. Verified numba 0.65 + numpy 2.4.6 + vllm import cleanly together. - docker-publish.yml: resolve UNSLOTH_ZOO_REF in a step that mirrors the pushed tag only when the tag exists in unsloth-zoo (the zoo currently cuts no tags, so blind mirroring broke every tag publish); falls back to main. - Dockerfile.studio: Studio venv stays on cu128 for arm64 too, matching the base venv (cu130 wheels would lift the driver floor to 580+), and gets the same NVRTC cu13 swap for DGX Spark / GB10 sm_121 support. - docker_confirm.sh: do not drop to CPU mode when docker info lacks a nvidia runtime entry; CDI installs and Docker Desktop WSL2 expose GPUs without one. The phase 3 --gpus probe is now the authority. - docker_confirm.ps1: GPU selector built as an args array; comma device lists get version-aware CSV quoting (native arg passing changed in PowerShell 7.3). - studio_launch.sh: no fixed Jupyter default password; generate a random one and print it when JUPYTER_PASSWORD is unset. Env snapshot for SSH sessions now written via shlex.quote instead of sed so values with quotes or command substitution cannot break or inject into /etc/profile.d. - install.ps1: honour UNSLOTH_TORCH_INDEX_FAMILY like install.sh does.
124 lines
5.9 KiB
Text
124 lines
5.9 KiB
Text
# Full Unsloth image: base training stack + Studio + JupyterLab + sshd.
|
|
#
|
|
# This is the image published as docker.io/unsloth/unsloth:latest. It layers
|
|
# Unsloth Studio on top of the lean base image (Dockerfile, published under
|
|
# the `base` tags) and runs the same service trio as the previous production
|
|
# image: Studio on 8000, JupyterLab on 8888, key-only sshd on 22.
|
|
#
|
|
# Build (local):
|
|
# docker buildx build \
|
|
# --build-arg BASE_IMAGE=unsloth-blackwell:test \
|
|
# -f docker/Dockerfile.studio \
|
|
# -t unsloth-blackwell:studio docker/
|
|
#
|
|
# Run:
|
|
# docker run --rm --gpus all -p 8000:8000 -p 8888:8888 \
|
|
# -v $HOME/.cache/huggingface:/workspace/.cache/huggingface \
|
|
# unsloth-blackwell:studio
|
|
#
|
|
# Open http://localhost:8000 for Studio (first-boot admin password is printed
|
|
# in the container logs and persisted under /opt/unsloth-studio/auth/) and
|
|
# http://localhost:8888 for JupyterLab (password: JUPYTER_PASSWORD env; when
|
|
# unset a random one is generated and printed in the container logs). On
|
|
# hosts without GPU passthrough (Docker Desktop on macOS, Windows without
|
|
# WSL2 GPU) add -e UNSLOTH_ALLOW_CPU=1: training is unavailable but Studio
|
|
# chat / Data Recipes / GGUF tooling / Jupyter work.
|
|
#
|
|
# CI pins BASE_IMAGE to the just-published multi-arch base digest so the two
|
|
# images always ship the same stack.
|
|
|
|
ARG BASE_IMAGE=unsloth-blackwell:test
|
|
FROM ${BASE_IMAGE}
|
|
|
|
# Studio source ref to clone. Defaults to `main`, but a CI publish pipeline
|
|
# that pins BASE_IMAGE to a digest should pin this too (same UNSLOTH_REF as
|
|
# the base) so the published image is reproducible against a known ref.
|
|
ARG UNSLOTH_STUDIO_REF=main
|
|
ARG TARGETARCH
|
|
|
|
# Services run as root in this revision (the base image is root-only by
|
|
# design); the previous production image ran them as a dedicated uid-1001
|
|
# user. Non-root parity is a tracked follow-up. sshd is key-only and stays
|
|
# disabled unless a PUBLIC_KEY/SSH_KEY is provided, and no secrets are
|
|
# persisted to disk (see studio_launch.sh).
|
|
#
|
|
# The JUPYTER_PORT / UNSLOTH_ENABLE_SSHD defaults exist so supervisord's
|
|
# %(ENV_*)s expansions still resolve when someone bypasses the launcher
|
|
# and runs supervisord directly.
|
|
USER root
|
|
ENV UNSLOTH_STUDIO_HOME=/opt/unsloth-studio \
|
|
JUPYTER_PORT=8888 \
|
|
UNSLOTH_ENABLE_SSHD=false \
|
|
DEBIAN_FRONTEND=noninteractive
|
|
|
|
# install.sh needs curl + git; supervisor + openssh-server run the service
|
|
# trio. The base image already has python + uv + pip.
|
|
RUN apt-get update \
|
|
&& apt-get install -y --no-install-recommends \
|
|
curl git ca-certificates supervisor openssh-server \
|
|
&& rm -rf /var/lib/apt/lists/*
|
|
|
|
# Clone + install Studio into a dedicated venv under $UNSLOTH_STUDIO_HOME.
|
|
# --local makes install.sh use the just-cloned source tree (editable
|
|
# install), so the source dir MUST persist for the venv's `unsloth_cli`
|
|
# entrypoint to keep resolving. Move it under $UNSLOTH_STUDIO_HOME/src
|
|
# (already inside the persistent layer) instead of deleting it. Strip
|
|
# .git to save ~120MB.
|
|
#
|
|
# The llama.cpp symlink BEFORE install.sh points Studio's prebuilt dir at
|
|
# the bundle already baked into the base image (validated, sha256-checked,
|
|
# UNSLOTH_PREBUILT_INFO.json present), so the installer's prebuilt step
|
|
# recognises it and skips a second ~400MB download. The
|
|
# .unsloth-studio-owned marker satisfies setup.sh's ownership assertion for
|
|
# custom STUDIO_HOMEs -- the dir IS provisioned exclusively for Studio.
|
|
#
|
|
# UNSLOTH_TORCH_INDEX_FAMILY pins the torch wheel index for the Studio
|
|
# venv: at build time there is no GPU and no nvidia-smi, so install.sh's
|
|
# probing would land on cpu or cu126 wheels depending on which host built
|
|
# the image. cu128 on BOTH arches, mirroring the base venv: cu130 wheels
|
|
# would silently lift the arm64 driver floor to 580+ while the base venv
|
|
# keeps the documented 570+ floor. DGX Spark / GB10 (sm_121) support comes
|
|
# from the same NVRTC cu13 swap the base image applies to its venv --
|
|
# repeated below for the Studio venv's own bundled libnvrtc (the base's
|
|
# arm64 layer already installed cuda-nvrtc-13-0, so the cu13 .so exists).
|
|
#
|
|
# fetch+checkout FETCH_HEAD instead of `clone --branch` because the CI
|
|
# pipeline passes a commit SHA as the ref (clone --branch only accepts
|
|
# branch/tag names).
|
|
RUN set -eux \
|
|
&& case "${TARGETARCH:-amd64}" in \
|
|
amd64|arm64) TORCH_FAMILY="cu128" ;; \
|
|
*) echo "ERROR: unsupported TARGETARCH=${TARGETARCH}" >&2; exit 1 ;; \
|
|
esac \
|
|
&& mkdir -p "${UNSLOTH_STUDIO_HOME}" \
|
|
&& ln -s /opt/unsloth/llama.cpp "${UNSLOTH_STUDIO_HOME}/llama.cpp" \
|
|
&& touch /opt/unsloth/llama.cpp/.unsloth-studio-owned \
|
|
&& git init -q "${UNSLOTH_STUDIO_HOME}/src" \
|
|
&& cd "${UNSLOTH_STUDIO_HOME}/src" \
|
|
&& git remote add origin https://github.com/unslothai/unsloth \
|
|
&& git fetch -q --depth 1 origin "${UNSLOTH_STUDIO_REF}" \
|
|
&& git checkout -q FETCH_HEAD \
|
|
&& UNSLOTH_STUDIO_HOME="${UNSLOTH_STUDIO_HOME}" \
|
|
UNSLOTH_TORCH_INDEX_FAMILY="${TORCH_FAMILY}" \
|
|
bash install.sh --local \
|
|
&& rm -rf "${UNSLOTH_STUDIO_HOME}/src/.git" /root/.cache \
|
|
&& if [ "${TARGETARCH:-amd64}" = "arm64" ]; then \
|
|
for NVRTC_DIR in "${UNSLOTH_STUDIO_HOME}"/unsloth_studio/lib/python*/site-packages/nvidia/cuda_nvrtc/lib; do \
|
|
if [ -f "${NVRTC_DIR}/libnvrtc.so.12" ]; then \
|
|
mv "${NVRTC_DIR}/libnvrtc.so.12" "${NVRTC_DIR}/libnvrtc.so.12.cu128.orig"; \
|
|
ln -s /usr/local/cuda-13.0/lib64/libnvrtc.so.13 "${NVRTC_DIR}/libnvrtc.so.12"; \
|
|
fi; \
|
|
done; \
|
|
fi
|
|
|
|
COPY supervisord.conf /etc/supervisor/supervisord.conf
|
|
COPY studio_launch.sh /usr/local/bin/unsloth-studio-launch
|
|
RUN chmod +x /usr/local/bin/unsloth-studio-launch
|
|
|
|
# Studio web UI, JupyterLab, sshd. All bind 0.0.0.0 inside the container's
|
|
# network namespace; the operator publishes them explicitly with -p.
|
|
EXPOSE 8000 8888 22
|
|
|
|
# The base ENTRYPOINT (unsloth-entrypoint) still runs its GPU pre-flight
|
|
# first, then hands off to the service launcher.
|
|
CMD ["/usr/local/bin/unsloth-studio-launch"]
|