- Changed absolute import to relative: from ._utils import prepare_model_for_kbit_training
- Added SUPPORTS_BFLOAT16 import for proper dtype detection
- Handle devices that don't support bfloat16 by falling back to float16
Changed from peft.prepare_model_for_kbit_training to
unsloth.models._utils.prepare_model_for_kbit_training.
Unsloth's version provides:
- Float32 mixed precision upcasting for LoRA layers
- Better numerical stability
- Consistency with rest of Unsloth codebase
The original implementation was 31% slower than naive SentenceTransformer due to
conflicting decorators from Unsloth's auto-compiler (@torch.compile on attention
modules but @torch.compiler.disable on sub-modules).
Changes:
- Add fast encoder path that bypasses Unsloth patching for encoder models
- Use native torch.compile with mode="reduce-overhead" for 6x speedup
- Auto-detect and enable SDPA for models that support it (BERT, RoBERTa, etc.)
- Change defaults: load_in_16bit=True, load_in_4bit=False (16-bit is optimal)
- Change default: use_gradient_checkpointing=False (conflicts with torch.compile)
- Add UNSLOTH_COMPILE_DISABLE=1 env var to fall back to old path if needed
Supported encoder types: mpnet, bert, distilbert, roberta, xlm-roberta, albert, electra
Benchmark results (BS=32, seq_len=128):
- Naive 16-bit LoRA: 13-50ms per iter
- Unsloth 16-bit LoRA: 2-9ms per iter (5.4x-6.7x faster)
- Memory usage: 61MB-1.3GB (even largest model fits easily)
Note: 4-bit + torch.compile has a PyTorch bug (pytorch/pytorch#90665).
4-bit is also 1.7-1.9x slower than 16-bit due to dequantization overhead,
so 16-bit is recommended for these small encoder models anyway.
When users load a model with fast_inference=False but then try to use
vLLM-style arguments with fast_generate, they previously got confusing
errors. This adds a wrapper that detects common mistakes and provides
helpful guidance:
- Using sampling_params: explains to use HF generate args instead
- Using lora_request: explains LoRA weights are already merged
- Passing text strings: shows how to tokenize input first
Changes:
- Add make_fast_generate_wrapper to _utils.py
- Apply wrapper in llama.py when fast_inference=False
- Apply wrapper in vision.py when fast_inference=False
Gemma3 models have a large vocabulary (262144 tokens) which causes
training loss to explode when using int8 embedding quantization.
This fix auto-detects Gemma3 models and switches from int8-int4
(phone-deployment) to int4 weight-only QAT for stable training.
1. cohere.py:347-348 - Fixed wrong variable names in QK normalization.
Used `Q`/`K` but variables were named `Qn`/`Kn`. This caused NameError
when `use_qk_norm=True` (e.g., c4ai-command-r-plus models).
2. cohere.py:482 - Fixed wrong object reference in inference loop.
Used `self.mlp` but should be `decoder_layer.mlp` since we're
iterating through decoder layers. Caused AttributeError during inference.
3. falcon_h1.py:459,461 - Fixed wrong attribute names in inference path.
Used `post_attention_layernorm` and `mlp` but Falcon H1 uses
`pre_ff_layernorm` and `feed_forward`. Caused AttributeError during generation.
4. qwen3_moe.py:210 - Fixed wrong module path with incorrect capitalization.
Used `transformers.models.Qwen3Moe` but should be `transformers.models.qwen3_moe`.
Caused AttributeError when patching rotary embeddings.
5. qwen3_moe.py:239 - Fixed wrong model_patcher class.
Used `FastQwen3Model` but should be `FastQwen3MoeModel` for MoE models.
Caused incorrect patching for Qwen3 MoE models.
6. hf_hub.py:21-22 - Fixed floor division and missing return for billion values.
Used `//` instead of `/` for millions, and had no return for values >= 1B.
Caused incorrect formatting and None return for large numbers.
7. save.py:550 - Fixed self-assignment that did nothing.
`sharded_ram_usage = sharded_ram_usage` should be `= max_shard_size`.
Caused integer shard sizes to be ignored.
8. rl.py:562-567 - Fixed orphan string not included in length_check.
The elif branch for max_seq_length validation was a standalone string
expression, not concatenated to length_check. Caused silent skip of
the max_seq_length > model_max_seq_length warning.
9. granite.py:49-52 - Fixed wrong model name and version in error message.
Said "Gemma2" and "4.42.3" but should be "Granite" and "4.45.0".
* Fix correctness bugs in rl.py, rl_replacements.py, and vision.py
1. rl_replacements.py (lines 864, 870): Fixed undefined `nanmin`/`nanmax`
functions by using `.nan_to_num(nan=inf/-inf).min()/.max()` pattern.
PyTorch doesn't have torch.nanmin/nanmax, so we replace NaN values
before computing min/max.
2. vision.py (line 150): Fixed bug where code checked for "input" key
but then accessed kwargs["input_ids"] instead of kwargs["input"].
3. vision.py (line 159): Fixed bug where literal string "key" was used
instead of the variable `key` when accessing kwargs.
4. rl.py (lines 903, 905): Fixed non-existent `MathError` exception
by replacing with `ValueError`.
* [pre-commit.ci] auto fixes from pre-commit.com hooks
for more information, see https://pre-commit.ci
---------
Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>