Edd
539fcea071
Fix/casting continue pretraining ( #1200 )
...
* Bring back float32 if float16 instead of bfloat16
* Refactor mixed precision handling for lm_head and embed_tokens to ensure correct dtype usage
* Fix dtype retrieval for embed_tokens and lm_head in mixed precision training
* Fix dtype retrieval for embed_tokens and lm_head to use weight dtype in mixed precision training
* Fix dtype handling for embed_tokens and lm_head to ensure correct float32 usage in mixed precision training
* Fix dtype assignment for lm_head modules to ensure correct weight dtype usage in mixed precision training
2024-10-27 15:06:45 -07:00
Daniel Han
9d5f58224d
Update pyproject.toml
2024-10-26 18:05:55 -07:00
Daniel Han
e7ede2f7db
Torch 2.5
2024-10-26 18:03:15 -07:00
Daniel Han
c58dc701c8
Bug fixes ( #1195 )
...
* Fix TRL
* Update mistral.py
* Patch processing_class
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Installation guide (#1165 )
* chore: update chat_templates.py (#1166 )
orginal -> original
* Disable Flex Attention
* Update tokenizer_utils.py
* Update _utils.py
* n_items
* Update cross_entropy_loss.py
* Fix DPO, ORPO
* Update _utils.py
* Update _utils.py
* fix/transformers-unpack (#1180 )
* Fix DPO, ORPO (#1177 )
* Fix TRL
* Update mistral.py
* Patch processing_class
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Installation guide (#1165 )
* chore: update chat_templates.py (#1166 )
orginal -> original
* Disable Flex Attention
* Update tokenizer_utils.py
* Update _utils.py
* n_items
* Update cross_entropy_loss.py
* Fix DPO, ORPO
* Update _utils.py
---------
Co-authored-by: timothelaborie <97834767+timothelaborie@users.noreply.github.com>
Co-authored-by: Ikko Eltociear Ashimine <eltociear@gmail.com>
* Add warning for missing Unpack and KwargsForCausalLM in older Transformers versions
---------
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
Co-authored-by: timothelaborie <97834767+timothelaborie@users.noreply.github.com>
Co-authored-by: Ikko Eltociear Ashimine <eltociear@gmail.com>
* Update cross_entropy_loss.py
* Update _utils.py
* Update _utils.py
* donot upcast lm_head and embeddings to float32 (#1186 )
* Cleanup upcast logs (#1188 )
* Fix/phi-longrope (#1193 )
* Enhance rotary embedding handling in LlamaAttention and LongRopeRotaryEmbedding
* Typo
* Improve rotary embedding handling in LlamaAttention to prevent errors with short KV cache
* Update llama.py
* Update llama.py
---------
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
* Update transformers
---------
Co-authored-by: timothelaborie <97834767+timothelaborie@users.noreply.github.com>
Co-authored-by: Ikko Eltociear Ashimine <eltociear@gmail.com>
Co-authored-by: Edd <68678137+Erland366@users.noreply.github.com>
Co-authored-by: Datta Nimmaturi <datta.nimmaturi@nutanix.com>
2024-10-26 01:21:24 -07:00
Daniel Han
519c0df00c
Update _utils.py
2024-10-24 12:17:48 -07:00
Daniel Han
06a5c752e3
Fix 4.47 issue ( #1182 )
...
* Fix TRL
* Update mistral.py
* Patch processing_class
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Installation guide (#1165 )
* chore: update chat_templates.py (#1166 )
orginal -> original
* Disable Flex Attention
* Update tokenizer_utils.py
* Update _utils.py
* n_items
* Update cross_entropy_loss.py
* Fix DPO, ORPO
* Update _utils.py
* Update _utils.py
* fix/transformers-unpack (#1180 )
* Fix DPO, ORPO (#1177 )
* Fix TRL
* Update mistral.py
* Patch processing_class
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Installation guide (#1165 )
* chore: update chat_templates.py (#1166 )
orginal -> original
* Disable Flex Attention
* Update tokenizer_utils.py
* Update _utils.py
* n_items
* Update cross_entropy_loss.py
* Fix DPO, ORPO
* Update _utils.py
---------
Co-authored-by: timothelaborie <97834767+timothelaborie@users.noreply.github.com>
Co-authored-by: Ikko Eltociear Ashimine <eltociear@gmail.com>
* Add warning for missing Unpack and KwargsForCausalLM in older Transformers versions
---------
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
Co-authored-by: timothelaborie <97834767+timothelaborie@users.noreply.github.com>
Co-authored-by: Ikko Eltociear Ashimine <eltociear@gmail.com>
* Update cross_entropy_loss.py
* Update _utils.py
* Update _utils.py
---------
Co-authored-by: timothelaborie <97834767+timothelaborie@users.noreply.github.com>
Co-authored-by: Ikko Eltociear Ashimine <eltociear@gmail.com>
Co-authored-by: Edd <68678137+Erland366@users.noreply.github.com>
2024-10-24 12:17:21 -07:00
Daniel Han
a6e4a8bf76
Fix DPO, ORPO ( #1177 )
...
* Fix TRL
* Update mistral.py
* Patch processing_class
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Installation guide (#1165 )
* chore: update chat_templates.py (#1166 )
orginal -> original
* Disable Flex Attention
* Update tokenizer_utils.py
* Update _utils.py
* n_items
* Update cross_entropy_loss.py
* Fix DPO, ORPO
* Update _utils.py
---------
Co-authored-by: timothelaborie <97834767+timothelaborie@users.noreply.github.com>
Co-authored-by: Ikko Eltociear Ashimine <eltociear@gmail.com>
2024-10-24 00:36:37 -07:00
Daniel Han
aa48184c41
Update _utils.py
2024-10-23 12:39:58 -07:00
Edd
da8e547678
Fix/patch tokenizer ( #1171 )
...
* fix: correct tokenizer handling in patch_sft_trainer_tokenizer
* Revert "fix: correct tokenizer handling in patch_sft_trainer_tokenizer"
This reverts commit 7a98e465cbd4f980c8b364b0396d44f2d052090f.
* fix: correct condition for test_text assignment in patch_sft_trainer_tokenizer
2024-10-23 12:32:33 -07:00
Daniel Han
4c85177719
Many bug fixes ( #1162 )
...
* Fix TRL
* Update mistral.py
* Patch processing_class
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Installation guide (#1165 )
* chore: update chat_templates.py (#1166 )
orginal -> original
* Disable Flex Attention
* Update tokenizer_utils.py
* Update _utils.py
---------
Co-authored-by: timothelaborie <97834767+timothelaborie@users.noreply.github.com>
Co-authored-by: Ikko Eltociear Ashimine <eltociear@gmail.com>
2024-10-23 03:14:57 -07:00
Daniel Han
a51a84f62d
Update save.py
2024-10-20 01:52:21 -07:00
Daniel Han
108fa8dfbe
Update _utils.py
2024-10-18 23:10:30 -07:00
Daniel Han
828ef9afc5
Fix get_token
2024-10-18 23:08:30 -07:00
vo1d-ai
f63a2a5026
fix: compute_loss bug ( #1151 )
...
Currently, Unsloth doesn't pass additional parameters to Trainer.compute_loss such as return_outputs. This leads to errors when calling trainer.evaluate(). This change fixes the bug by properly passing parameters to Trainer.compute_loss.
2024-10-18 20:46:07 -07:00
Daniel Han
cde7401259
Update _utils.py
2024-10-17 20:50:05 -07:00
Daniel Han
139c3b29b3
Update README.md
2024-10-17 20:46:11 -07:00
Daniel Han
3a33dad3c9
Update README.md
2024-10-17 20:45:40 -07:00
Daniel Han
d57dcf58a1
Gradient Accumulation Fix ( #1146 )
...
* Unsloth Zoo
* Update trainer.py
* Update trainer.py
* Update cross_entropy_loss.py
* n_items
* Update llama.py
* kwargs
* Remove extraneous f prefixes (#1133 )
Co-authored-by: Emil Sadek <esadek@users.noreply.github.com>
* Update __init__.py
* kwargs
* Update trainer.py
* Update trainer.py
* Update trainer.py
* Fix GA
* Update _utils.py
* Update llama.py
* Update tokenizer_utils.py
* Warn on old versions
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
---------
Co-authored-by: Emil Sadek <esadek@hotmail.com>
Co-authored-by: Emil Sadek <esadek@users.noreply.github.com>
2024-10-17 20:43:07 -07:00
Daniel Han
6ba202708c
Update mapper.py
2024-10-16 21:48:05 -07:00
Daniel Han
d2a032e117
Gradient Accumulation Fix ( #1134 )
...
* Unsloth Zoo
* Update trainer.py
* Update trainer.py
* Update cross_entropy_loss.py
* n_items
* Update llama.py
* kwargs
* Remove extraneous f prefixes (#1133 )
Co-authored-by: Emil Sadek <esadek@users.noreply.github.com>
* Update __init__.py
---------
Co-authored-by: Emil Sadek <esadek@hotmail.com>
Co-authored-by: Emil Sadek <esadek@users.noreply.github.com>
2024-10-14 19:17:35 -07:00
Daniel Han
5bd7d3640f
Update save.py
2024-10-11 00:00:06 -07:00
Daniel Han
9dd4462bf9
Update save.py
2024-10-10 23:22:17 -07:00
Giulia Baldini
592191b061
Only remove folder in sentenpiece check if it was created ( #1121 )
2024-10-10 23:21:27 -07:00
Giulia Baldini
5f2d5a3021
Handle absolute paths using pathlib ( #1120 )
2024-10-10 23:20:34 -07:00
Daniel Han
e130e748f0
Reload
2024-10-05 17:21:48 -07:00
Daniel Han
c89ae6b9b4
Merge branch 'nightly'
2024-10-01 00:45:16 -07:00
Daniel Han
3c47723bb2
Update README.md
2024-10-01 00:40:17 -07:00
Daniel Han
7fc9b07b94
Update tokenizer_utils.py
2024-10-01 00:35:54 -07:00
Daniel Han
3017eae097
Update tokenizer_utils.py
2024-10-01 00:20:01 -07:00
Daniel Han
dfc5cd3c80
Update tokenizer_utils.py
2024-10-01 00:14:52 -07:00
Daniel Han
ac3f564f7e
Update chat_templates.py
2024-09-30 23:08:46 -07:00
Daniel Han
248c27d205
Fix merges ( #1079 )
...
* Layernorm
* Update layernorm.py
* Update layernorm.py
* Update layernorm.py
* Update layernorm.py
* Update layernorm.py
* Update layernorm.py
* Patch layernorm
* Update layernorm.py
* RMS Layernorm
* Update rms_layernorm.py
* Causal LM
* Update cross_entropy_loss.py
* Update cross_entropy_loss.py
* Update cross_entropy_loss.py
* Update cross_entropy_loss.py
* Update layernorm.py
* Update cross_entropy_loss.py
* Update cross_entropy_loss.py
* Update cross_entropy_loss.py
* Update cross_entropy_loss.py
* Update cross_entropy_loss.py
* Update cross_entropy_loss.py
* Update cross_entropy_loss.py
* Update _utils.py
* Update _utils.py
* Llama 3.2
* Update _utils.py
* Update _utils.py
* Update _utils.py
* Update llama.py
* Update vision.py
* Update llama.py
* Update llama.py
* Update llama.py
* Update llama.py
* Update llama.py
* Update llama.py
* Update llama.py
* Update llama.py
* Update loader.py
* Update loader.py
* Update loader.py
* Dependencies
* Update pyproject.toml
* Update _utils.py
2024-09-30 03:03:01 -07:00
Daniel Han
0f1f2cf728
Update _utils.py
2024-09-30 02:51:32 -07:00
Daniel Han
906f88bb4d
Update pyproject.toml
2024-09-30 02:48:31 -07:00
Daniel Han
e31152134e
Dependencies
2024-09-30 02:09:15 -07:00
Daniel Han
45916d36cf
Update loader.py
2024-09-29 23:22:05 -07:00
Daniel Han
a529a39c81
Update loader.py
2024-09-29 23:15:22 -07:00
Daniel Han
f744b3159e
Update loader.py
2024-09-29 23:13:44 -07:00
Daniel Han
ab43e02a94
Merge branch 'main' into nightly
2024-09-29 23:13:16 -07:00
Daniel Han
afbb140a79
Update loader.py
2024-09-29 01:42:58 -07:00
Daniel Han
b314837622
Update pyproject.toml
2024-09-27 01:36:45 -07:00
Daniel Han
c0b4d640f2
Update tokenizer_utils.py
2024-09-26 01:23:40 -07:00
Daniel Han
88a542a129
Update README.md
2024-09-26 00:12:42 -07:00
Daniel Han
6bbca3aaa8
Update README.md
2024-09-26 00:05:38 -07:00
Daniel Han
4f4ef22035
Update README.md
2024-09-26 00:02:15 -07:00
Daniel Han
930d2ad1a8
Update pyproject.toml
2024-09-25 23:47:15 -07:00
Daniel Han
5b345ec757
Update pyproject.toml
2024-09-25 23:13:49 -07:00
Daniel Han
c331c886ee
Remove version checks
2024-09-25 23:00:09 -07:00
Daniel Han
63e3a85efb
Update _utils.py
2024-09-25 22:56:41 -07:00
Daniel Han
dc8bca6713
Update llama.py
2024-09-25 22:12:21 -07:00