* Update llama.py * offload * Update llama.py * Update llama.py * Update llama.py * Update llama.py * Update llama.py * Update llama.py * Update llama.py * continued pretraining trainer * Update trainer.py * Update trainer.py * Update trainer.py * Update trainer.py * is_bfloat16_supported * Update __init__.py * Update README.md * Update llama.py * is_bfloat16_supported * Update __init__.py * Mistral v3 * Phi 3 medium * Update chat_templates.py * Update chat_templates.py * Phi-3 * Update save.py * Update README.md Mistral v3 to Mistral v0.3 * Untrained tokens * Update tokenizer_utils.py * Update tokenizer_utils.py * Update tokenizer_utils.py * Update tokenizer_utils.py * Update tokenizer_utils.py * Update tokenizer_utils.py * Update tokenizer_utils.py * Update tokenizer_utils.py * Update tokenizer_utils.py * Update tokenizer_utils.py * Update tokenizer_utils.py * Update tokenizer_utils.py * Update tokenizer_utils.py * Update tokenizer_utils.py * Update tokenizer_utils.py * Update tokenizer_utils.py * Update tokenizer_utils.py * Update tokenizer_utils.py * Update tokenizer_utils.py * Update llama.py * Update tokenizer_utils.py * Update tokenizer_utils.py * Update tokenizer_utils.py * Update tokenizer_utils.py * Update save.py * Update save.py * Update save.py * checkpoint * Update _utils.py * Update tokenizer_utils.py * Update tokenizer_utils.py * Update tokenizer_utils.py * Update llama.py * accelerate * Update _utils.py * Update _utils.py * Update _utils.py * Update _utils.py * Update _utils.py * Update _utils.py * Update _utils.py * Update tokenizer_utils.py * train_dataloader * Update llama.py * Update llama.py * Update llama.py * use_fast_convert * Update save.py * Update save.py * Update save.py * Update save.py * remove_special_tokens * Ollama * Update chat_templates.py * Update chat_templates.py * Update chat_templates.py * Update llama.py * Update chat_templates.py * Support bfloat16 GGUF * Update save.py * Update llama.py * fast_forward_inference * Update mapper.py * Update loader.py * Update llama.py * Update tokenizer_utils.py * info * edits * Create chat template * Fix tokenizer * Update tokenizer_utils.py * fix case where gguf saving fails due to first_conversion dtype (#630) * Support revision parameter in FastLanguageModel.from_pretrained (#629) * support `revision` parameter * match unsloth formatting of named parameters * clears any selected_adapters before calling internal_model.save_pretrained (#609) * Update __init__.py (#602) Check for incompatible modules before importing unsloth * Fixed unsloth/tokenizer_utils.py for chat training (#604) * Add GGML saving option to Unsloth for easier Ollama model creation and testing. (#345) * Add save to llama.cpp GGML to save.py. * Fix conversion command and path of convert to GGML function. * Add autosaving lora to the GGML function * Create lora save function for conversion to GGML * Test fix #2 for saving lora * Test fix #3 to save the lora adapters to convert to GGML * Remove unwated tokenizer saving for conversion to ggml and added a few print statements. * Needed tokenizer for saving, added it back, also made it more unslothy style by having positional arguments, and added a few messages. * Positional arguments didn't work out, so reverted to older version of the code, and added a few comments. * Test fix 1 for arch * Test fix 2 new Mistral error. * Test fix 3 * Revert to old version for testing. * Upload issue test fix 1 * Fix 2 uploading ggml * Positional ags added. * Temporray remove positional args * Fix upload again!!! * Add print statements and fix link * Make the calling name better * Create local saving for GGML * Add choosing directory to save local GGML. * Fix lil variable error in the save_to_custom_dir func * docs: Add LoraConfig parameters documentation (#619) * llama.cpp failing (#371) llama.cpp is failing to generate quantize versions for the trained models. Error: ```bash You might have to compile llama.cpp yourself, then run this again. You do not need to close this Python program. Run the following commands in a new terminal: You must run this in the same folder as you're saving your model. git clone https://github.com/ggerganov/llama.cpp cd llama.cpp && make clean && LLAMA_CUDA=1 make all -j Once that's done, redo the quantization. ``` But when i do clone this with recursive it works. Co-authored-by: Daniel Han <danielhanchen@gmail.com> * fix libcuda_dirs import for triton 3.0 (#227) * fix libcuda_dirs import for triton 3.0 * Update __init__.py * Update __init__.py --------- Co-authored-by: Daniel Han <danielhanchen@gmail.com> * Update save.py * Update __init__.py * Update fast_lora.py * Update save.py * Update save.py * Update save.py * Update loader.py * Update save.py * Update save.py * quantize now llama-quantize * Update chat_templates.py * Update loader.py * Update mapper.py * Update __init__.py * embedding size * Update qwen2.py * docs * Update README.md * Update qwen2.py * README: Fix minor typo. (#559) * README: Fix minor typo. One-character typo fix while reading. * Update README.md --------- Co-authored-by: Daniel Han <danielhanchen@gmail.com> * Update mistral.py * Update qwen2.py * Update qwen2.py * Update qwen2.py * Update llama.py * Update llama.py * Update llama.py * Update README.md * FastMistralModel * Update mistral.py * Update mistral.py * Update mistral.py * Update mistral.py * Update mistral.py * Auto check rope scaling * Update llama.py * Update llama.py * Update llama.py * GPU support * Typo * Update gemma.py * gpu * Multiple GGUF saving * Update save.py * Update save.py --------- Co-authored-by: Michael Han <107991372+shimmyshimmer@users.noreply.github.com> Co-authored-by: Eliot Hall <60240707+chrehall68@users.noreply.github.com> Co-authored-by: Rickard Edén <rickardeden@gmail.com> Co-authored-by: XiaoYang <xyangk@gmail.com> Co-authored-by: Oseltamivir <58582368+Oseltamivir@users.noreply.github.com> Co-authored-by: mahiatlinux <110882203+mahiatlinux@users.noreply.github.com> Co-authored-by: Sébastien De Greef <sebdg@binarycompany.com> Co-authored-by: Alberto Ferrer <albertof@barrahome.org> Co-authored-by: Thomas Viehmann <tv.github-private@beamnet.de> Co-authored-by: Walter Korman <lemurware@gmail.com>
323 lines
12 KiB
Python
323 lines
12 KiB
Python
# Copyright 2023-present Daniel Han-Chen & the Unsloth team. All rights reserved.
|
|
#
|
|
# Licensed under the Apache License, Version 2.0 (the "License");
|
|
# you may not use this file except in compliance with the License.
|
|
# You may obtain a copy of the License at
|
|
#
|
|
# http://www.apache.org/licenses/LICENSE-2.0
|
|
#
|
|
# Unless required by applicable law or agreed to in writing, software
|
|
# distributed under the License is distributed on an "AS IS" BASIS,
|
|
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
# See the License for the specific language governing permissions and
|
|
# limitations under the License.
|
|
|
|
from .llama import *
|
|
import os
|
|
from ._utils import __version__
|
|
|
|
from transformers.models.mistral.modeling_mistral import (
|
|
MistralAttention,
|
|
MistralDecoderLayer,
|
|
MistralModel,
|
|
MistralForCausalLM,
|
|
)
|
|
# For Pytorch 2.1.1
|
|
try:
|
|
from transformers.models.mistral.modeling_mistral import (
|
|
MistralSdpaAttention,
|
|
MistralFlashAttention2,
|
|
)
|
|
except:
|
|
MistralSdpaAttention = MistralAttention
|
|
MistralFlashAttention2 = MistralAttention
|
|
pass
|
|
|
|
|
|
def MistralAttention_fast_forward(
|
|
self,
|
|
hidden_states: torch.Tensor,
|
|
causal_mask: Optional[xformers.attn_bias.BlockDiagonalCausalMask] = None,
|
|
attention_mask: Optional[torch.Tensor] = None,
|
|
position_ids: Optional[torch.LongTensor] = None,
|
|
past_key_value: Optional[Tuple[torch.Tensor]] = None,
|
|
output_attentions: bool = False,
|
|
use_cache: bool = False,
|
|
padding_mask: Optional[torch.LongTensor] = None,
|
|
*args, **kwargs,
|
|
) -> Tuple[torch.Tensor, Optional[torch.Tensor], Optional[Tuple[torch.Tensor]]]:
|
|
|
|
# Clear inference
|
|
if hasattr(self, "paged_attention"):
|
|
del self.paged_attention_K
|
|
del self.paged_attention_V
|
|
del self.paged_attention
|
|
del self.temp_QA
|
|
del self.temp_KV
|
|
del self.RH_Q
|
|
del self.attention
|
|
pass
|
|
|
|
bsz, q_len, _ = hidden_states.size()
|
|
|
|
n_heads = self.num_heads
|
|
n_groups = self.num_key_value_groups
|
|
n_kv_heads = self.num_key_value_heads
|
|
head_dim = self.head_dim
|
|
assert(n_kv_heads * n_groups == n_heads)
|
|
|
|
Q, K, V = self.apply_qkv(self, hidden_states)
|
|
Q = Q.view(bsz, q_len, n_heads, head_dim).transpose(1, 2)
|
|
K = K.view(bsz, q_len, n_kv_heads, head_dim).transpose(1, 2)
|
|
V = V.view(bsz, q_len, n_kv_heads, head_dim).transpose(1, 2)
|
|
|
|
kv_seq_len = K.shape[-2]
|
|
if past_key_value is not None:
|
|
kv_seq_len += past_key_value[0].shape[-2]
|
|
|
|
if position_ids is None:
|
|
cos = self.rotary_emb.cos_cached
|
|
sin = self.rotary_emb.sin_cached
|
|
Q, K = fast_rope_embedding(Q, K, cos, sin)
|
|
else:
|
|
cos, sin = self.rotary_emb(V, seq_len = kv_seq_len)
|
|
Q, K = inplace_rope_embedding(Q, K, cos, sin, position_ids)
|
|
pass
|
|
|
|
if past_key_value is not None:
|
|
K = torch.cat([past_key_value[0], K], dim = 2)
|
|
V = torch.cat([past_key_value[1], V], dim = 2)
|
|
pass
|
|
past_key_value = (K, V) if use_cache else None
|
|
|
|
# Attention module
|
|
if (not HAS_FLASH_ATTENTION and attention_mask is None):
|
|
# Xformers memory efficient attention
|
|
Q = Q.transpose(1, 2)
|
|
K = K.transpose(1, 2)
|
|
V = V.transpose(1, 2)
|
|
K_M = V_M = bsz * kv_seq_len
|
|
Q_M = bsz * q_len
|
|
|
|
has_swa = isinstance(causal_mask, xformers.attn_bias.BlockDiagonalCausalMask)
|
|
|
|
# Group query attention
|
|
K = K .view(bsz, kv_seq_len, n_kv_heads, 1, head_dim)
|
|
V = V .view(bsz, kv_seq_len, n_kv_heads, 1, head_dim)
|
|
K = K.expand(bsz, kv_seq_len, n_kv_heads, n_groups, head_dim)
|
|
V = V.expand(bsz, kv_seq_len, n_kv_heads, n_groups, head_dim)
|
|
if hidden_states.requires_grad:
|
|
K = K.reshape(bsz, kv_seq_len, n_heads, head_dim)
|
|
V = V.reshape(bsz, kv_seq_len, n_heads, head_dim)
|
|
|
|
if has_swa:
|
|
Q = Q.view(1, Q_M, n_heads, head_dim)
|
|
K = K.view(1, K_M, n_heads, head_dim)
|
|
V = V.view(1, V_M, n_heads, head_dim)
|
|
pass
|
|
else:
|
|
# Xformers does support the forward pass though
|
|
Q = Q.view(bsz, q_len, n_kv_heads, n_groups, head_dim)
|
|
|
|
if has_swa:
|
|
Q = Q.view(1, Q_M, n_kv_heads, n_groups, head_dim)
|
|
K = K.view(1, K_M, n_kv_heads, n_groups, head_dim)
|
|
V = V.view(1, V_M, n_kv_heads, n_groups, head_dim)
|
|
pass
|
|
pass
|
|
|
|
A = xformers_attention(Q, K, V, attn_bias = causal_mask)
|
|
A = A.view(bsz, q_len, n_heads, head_dim)
|
|
|
|
elif HAS_FLASH_ATTENTION and attention_mask is None:
|
|
Q = Q.transpose(1, 2)
|
|
K = K.transpose(1, 2)
|
|
V = V.transpose(1, 2)
|
|
sw = getattr(self.config, "sliding_window", None)
|
|
sw = kv_seq_len if (sw is None or sw == "null") else sw
|
|
window = (-1, -1) if (kv_seq_len <= sw) else (sw, sw)
|
|
A = flash_attn_func(Q, K, V, causal = True, window_size = window)
|
|
else:
|
|
# Grouped query attention
|
|
# if n_groups != 1:
|
|
K = K[:, :, None, :, :].expand(bsz, n_kv_heads, n_groups, kv_seq_len, head_dim)
|
|
V = V[:, :, None, :, :].expand(bsz, n_kv_heads, n_groups, kv_seq_len, head_dim)
|
|
K = K.reshape(bsz, n_heads, kv_seq_len, head_dim)
|
|
V = V.reshape(bsz, n_heads, kv_seq_len, head_dim)
|
|
# pass
|
|
# Must be contiguous or else results are False!
|
|
# https://github.com/pytorch/pytorch/issues/112577
|
|
Q, K, V = Q.contiguous(), K.contiguous(), V.contiguous()
|
|
# Needs (batch_size, n_heads, seq_len, head_dim)
|
|
# is_casual and attention_mask must not be both set!
|
|
A = scaled_dot_product_attention(Q, K, V, attn_mask = attention_mask, is_causal = False)
|
|
# Go back to (batch_size, seq_len, n_heads, head_dim)
|
|
A = A.transpose(1, 2).contiguous()
|
|
pass
|
|
|
|
attn_output = A.reshape(bsz, q_len, self.hidden_size)
|
|
attn_output = self.apply_o(self, attn_output)
|
|
attn_weights = None
|
|
return attn_output, attn_weights, past_key_value
|
|
pass
|
|
|
|
|
|
def MistralForCausalLM_fast_forward(
|
|
self,
|
|
input_ids: torch.LongTensor = None,
|
|
causal_mask: Optional[xformers.attn_bias.BlockDiagonalCausalMask] = None,
|
|
attention_mask: Optional[torch.Tensor] = None,
|
|
position_ids: Optional[torch.LongTensor] = None,
|
|
past_key_values: Optional[List[torch.FloatTensor]] = None,
|
|
inputs_embeds: Optional[torch.FloatTensor] = None,
|
|
labels: Optional[torch.LongTensor] = None,
|
|
use_cache: Optional[bool] = None,
|
|
output_attentions: Optional[bool] = None,
|
|
output_hidden_states: Optional[bool] = None,
|
|
return_dict: Optional[bool] = None,
|
|
*args, **kwargs,
|
|
) -> Union[Tuple, CausalLMOutputWithPast]:
|
|
|
|
if causal_mask is None and past_key_values is None:
|
|
bsz, q_len = input_ids.shape
|
|
sliding_window = getattr(self.config, "sliding_window", None)
|
|
if sliding_window is None or sliding_window == "null" or sliding_window <= 0:
|
|
causal_mask = xformers.attn_bias.LowerTriangularMask()
|
|
elif q_len <= sliding_window:
|
|
causal_mask = xformers.attn_bias.LowerTriangularMask()
|
|
else:
|
|
# Fix from https://github.com/Rypo
|
|
causal_mask = xformers.attn_bias.BlockDiagonalCausalMask\
|
|
.from_seqlens([q_len]*bsz)\
|
|
.make_local_attention(window_size = sliding_window)
|
|
pass
|
|
|
|
output_attentions = output_attentions if output_attentions is not None else self.config.output_attentions
|
|
output_hidden_states = (
|
|
output_hidden_states if output_hidden_states is not None else self.config.output_hidden_states
|
|
)
|
|
return_dict = return_dict if return_dict is not None else self.config.use_return_dict
|
|
|
|
# decoder outputs consists of (dec_features, layer_state, dec_hidden, dec_attn)
|
|
self.model._has_no_labels = labels is None
|
|
|
|
if past_key_values is not None:
|
|
outputs = LlamaModel_fast_forward_inference(
|
|
self,
|
|
input_ids,
|
|
past_key_values,
|
|
position_ids = position_ids,
|
|
attention_mask = attention_mask,
|
|
)
|
|
else:
|
|
outputs = self.model(
|
|
input_ids=input_ids,
|
|
causal_mask=causal_mask,
|
|
attention_mask=attention_mask,
|
|
position_ids=position_ids,
|
|
past_key_values=past_key_values,
|
|
inputs_embeds=inputs_embeds,
|
|
use_cache=use_cache,
|
|
output_attentions=output_attentions,
|
|
output_hidden_states=output_hidden_states,
|
|
return_dict=return_dict,
|
|
)
|
|
pass
|
|
|
|
hidden_states = outputs[0]
|
|
bsz, q_len, hd = hidden_states.shape
|
|
lm_head = self.lm_head.weight
|
|
if bsz == 1 and q_len == 1:
|
|
logits = torch.mv(lm_head, hidden_states.ravel().to(lm_head.dtype))
|
|
logits = logits.unsqueeze(0).unsqueeze(0)
|
|
else:
|
|
logits = self.lm_head(hidden_states.to(lm_head.dtype))
|
|
pass
|
|
logits = logits.to(self.config.torch_dtype)
|
|
|
|
loss = None
|
|
if labels is not None:
|
|
shift_logits = logits
|
|
if not hasattr(self, "extra_ignored_labels"):
|
|
device_ids = os.environ.get("CUDA_VISIBLE_DEVICES", "0") + ","
|
|
device = device_ids[:device_ids.find(',')] # Unsloth only works on NVIDIA GPUs for now
|
|
device = f"cuda:{device if device.isdigit() else '0'}"
|
|
# Fixes https://github.com/unslothai/unsloth/issues/10
|
|
self.extra_ignored_labels = torch.full((self.max_seq_length, 1), -100, device = device)
|
|
pass
|
|
|
|
shift_labels = torch.hstack((labels[..., 1:], self.extra_ignored_labels[:labels.shape[0]]))
|
|
loss = fast_cross_entropy_loss(
|
|
logits = shift_logits,
|
|
labels = shift_labels,
|
|
)
|
|
pass
|
|
|
|
if not return_dict:
|
|
output = (logits,) + outputs[1:]
|
|
return (loss,) + output if loss is not None else output
|
|
|
|
return CausalLMOutputWithPast(
|
|
loss=loss,
|
|
logits=logits,
|
|
past_key_values=outputs.past_key_values,
|
|
hidden_states=outputs.hidden_states,
|
|
attentions=outputs.attentions,
|
|
)
|
|
pass
|
|
|
|
|
|
class FastMistralModel(FastLlamaModel):
|
|
|
|
@staticmethod
|
|
def pre_patch():
|
|
MistralAttention .forward = MistralAttention_fast_forward
|
|
MistralSdpaAttention .forward = MistralAttention_fast_forward
|
|
MistralFlashAttention2.forward = MistralAttention_fast_forward
|
|
MistralDecoderLayer .forward = LlamaDecoderLayer_fast_forward
|
|
MistralModel .forward = LlamaModel_fast_forward
|
|
MistralForCausalLM .forward = MistralForCausalLM_fast_forward
|
|
PeftModelForCausalLM .forward = PeftModelForCausalLM_fast_forward
|
|
|
|
# Solves https://github.com/unslothai/unsloth/issues/168
|
|
# Static KV Cache was introduced in 4.38.0, causing training to be much slower.
|
|
# Inferene can now be CUDAGraphed, but we shall retain the old rotary embeddings.
|
|
# https://github.com/huggingface/transformers/pull/27931
|
|
# https://github.com/huggingface/transformers/blob/v4.37.2/src/transformers/models/llama/modeling_llama.py
|
|
import transformers.models.mistral.modeling_mistral
|
|
transformers.models.mistral.modeling_mistral.MistralRotaryEmbedding = LlamaRotaryEmbedding
|
|
return
|
|
pass
|
|
|
|
|
|
@staticmethod
|
|
def from_pretrained(
|
|
model_name = "unsloth/mistral-7b-bnb-4bit",
|
|
max_seq_length = None,
|
|
dtype = None,
|
|
load_in_4bit = True,
|
|
token = None,
|
|
device_map = "sequential",
|
|
rope_scaling = None, # Mistral does not support RoPE scaling
|
|
fix_tokenizer = True,
|
|
model_patcher = None,
|
|
tokenizer_name = None,
|
|
trust_remote_code = False,
|
|
**kwargs,
|
|
):
|
|
return FastLlamaModel.from_pretrained(
|
|
model_name = model_name,
|
|
max_seq_length = max_seq_length,
|
|
dtype = dtype,
|
|
load_in_4bit = load_in_4bit,
|
|
token = token,
|
|
device_map = device_map,
|
|
rope_scaling = rope_scaling,
|
|
fix_tokenizer = fix_tokenizer,
|
|
model_patcher = FastMistralModel,
|
|
tokenizer_name = tokenizer_name,
|
|
trust_remote_code = trust_remote_code,
|
|
**kwargs,
|
|
)
|
|
pass
|
|
pass
|