Two defects found by running the built image rather than reading it.
1. Every unsloth_cli subcommand that touches the studio backend died on
import. `unsloth list-checkpoints` on the published image:
ModuleNotFoundError: No module named 'structlog'
and the same for train / export / chat, since all four import
studio.backend.core.*. structlog is a studio backend requirement, not
an unsloth[huggingface] one, so nothing in the base install pulled it
in. Added it to the base venv, and added a build-time
`from studio.backend.core.export import ExportBackend` so a future
missing dependency in that closure fails the build instead of the
user's first CLI invocation. That guard has to live in the LAST
builder verification block: the closure also needs starlette, which
only arrives with vLLM two stages later.
2. flashinfer-jit-cache was pinned to a literal 0.6.6 while vLLM 0.26.0
resolves flashinfer-python 0.6.14. flashinfer raises at import when
the two disagree, and that exception is thrown inside the vLLM
EngineCore, so Unsloth's GRPO fast_inference path fails at engine
start with no earlier warning. A literal pin drifts again on the next
vLLM bump, so the version is now read back from the resolved
flashinfer-python, and the build proves `import flashinfer` works.
Verified on the rebuilt image: flashinfer-python 0.6.14 with
flashinfer-jit-cache 0.6.14+cu128, structlog 26.1.0, the export backend
importable, and `unsloth list-checkpoints` exiting 0.
tests/python/test_docker_llama_cuda_backend.py gains two static cases
pinning both: the jit-cache version must be derived rather than literal
and the build must import flashinfer, and the base venv must ask for
structlog with the CLI reachability guard present.