The lazy loader retries `import unsloth_zoo.hf_xet_fallback` under
UNSLOTH_ZOO_DISABLE_GPU_INIT=1 whenever the first attempt raises. That flag makes
unsloth_zoo take its MLX/CPU path, which injects triton and bitsandbytes STUBS into
sys.modules for the rest of the process.
On a working GPU box whose first import failed for an unrelated reason (a
bitsandbytes/CUDA mismatch, say) the retry succeeds, so Studio boots looking healthy
and then dies at the first CUDA-only kernel: a GGUF or compiled diffusion generation
hits the stub and returns
NotImplementedError: Unsloth: 'triton.tools.experimental_descriptor.enable_in_pytorch'
was called on Apple Silicon / MLX, where triton is stubbed out.
so every image generation 500s with an Apple-Silicon message on a Linux CUDA host,
while the load reports success. Found by loading Z-Image-Turbo GGUF through the API on
a box where bitsandbytes could not initialise.
Gate the retry on the host genuinely having no accelerator. The Xet stall watchdog is
optional and already degrades with a warning; a process whose triton is stubbed out is
not recoverable. The warning now says why it did not retry.