Commit graph

1 commit

Author SHA1 Message Date
Daniel Han
7b208bc35c
Fix misleading 'only for image models' error for Qwen3-VL when torchvision is missing (#6525)
* Fix misleading 'only for image models' error for Qwen3-VL when torchvision is missing

transformers >= 5.4 hard-requires torchvision for VLM image/video processors and
no longer falls back to a slow processor. Without torchvision the processor load
raises ImportError, unsloth degrades to a text-only tokenizer, and the vision data
collator later fails with 'UnslothVisionDataCollator is only for image models!'.

Detect this case at load time and raise a clear, actionable error pointing at the
missing torchvision dependency instead.

Fixes unslothai/unsloth#4202

* Apply kwarg-spacing format hook to vision torchvision guard (pre-commit)

* Make torchvision-missing detection precise: check availability first, match specific error text

* Tighten code comments (no logic change)

* Make missing-torchvision VLM error version-agnostic

The raise also fires on transformers 4.57.x for VLMs with a video processor
(Qwen2.5-VL, Qwen3-VL), where AutoVideoProcessor requires torchvision. The old
message claimed 'transformers >= 5.4 requires torchvision', which is inaccurate
on 4.57.x. Reword to state torchvision is required for this model's vision
processors without a version-specific claim.

---------

Co-authored-by: danielhanchen <michaelhan2050@gmail.com>
2026-06-23 01:28:09 -07:00