Daniel Han
ede5d5893a
Update cross_entropy_loss.py
2024-10-28 14:30:06 -07:00
Daniel Han
a859b5d796
Update _utils.py
2024-10-28 10:41:31 -07:00
Daniel Han
e8855dfd4f
Update _utils.py
2024-10-28 01:12:33 -07:00
Daniel Han
2748eecc9d
More patching
2024-10-28 01:10:23 -07:00
Daniel Han
67c258b4fc
Revert "ignored labels"
...
This reverts commit 110ee41971 .
2024-10-27 22:18:05 -07:00
Daniel Han
110ee41971
ignored labels
2024-10-27 22:10:59 -07:00
Daniel Han
54bf89c11f
Typo
2024-10-27 19:09:27 -07:00
Daniel Han
f14d37a778
Update llama.py
2024-10-27 19:08:14 -07:00
Daniel Han
1254304ea9
Fix pad token
2024-10-27 19:06:57 -07:00
Daniel Han
0a445a417a
Update _utils.py
2024-10-27 17:34:33 -07:00
Daniel Han
94d03dba99
Unk token issues
2024-10-27 17:32:26 -07:00
Daniel Han
cdd128f2eb
Merge branch 'main' into nightly
2024-10-27 16:24:20 -07:00
Daniel Han
a9afc2f291
Merge branch 'main' of https://github.com/unslothai/unsloth
2024-10-27 15:09:42 -07:00
Daniel Han
6de3d3b858
Update _utils.py
2024-10-27 15:09:35 -07:00
Edd
1d1d36de46
Fix/casting continue pretraining ( #1200 )
...
* Bring back float32 if float16 instead of bfloat16
* Refactor mixed precision handling for lm_head and embed_tokens to ensure correct dtype usage
* Fix dtype retrieval for embed_tokens and lm_head in mixed precision training
* Fix dtype retrieval for embed_tokens and lm_head to use weight dtype in mixed precision training
* Fix dtype handling for embed_tokens and lm_head to ensure correct float32 usage in mixed precision training
* Fix dtype assignment for lm_head modules to ensure correct weight dtype usage in mixed precision training
2024-10-27 15:06:45 -07:00
Daniel Han
8faac33203
Update pyproject.toml
2024-10-26 18:05:55 -07:00
Daniel Han
9f8b4589c0
Torch 2.5
2024-10-26 18:03:15 -07:00
Daniel Han
558efd0899
Merge branch 'main' into nightly
2024-10-26 01:22:21 -07:00
Daniel Han
4b02021f68
Bug fixes ( #1195 )
...
* Fix TRL
* Update mistral.py
* Patch processing_class
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Installation guide (#1165 )
* chore: update chat_templates.py (#1166 )
orginal -> original
* Disable Flex Attention
* Update tokenizer_utils.py
* Update _utils.py
* n_items
* Update cross_entropy_loss.py
* Fix DPO, ORPO
* Update _utils.py
* Update _utils.py
* fix/transformers-unpack (#1180 )
* Fix DPO, ORPO (#1177 )
* Fix TRL
* Update mistral.py
* Patch processing_class
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Installation guide (#1165 )
* chore: update chat_templates.py (#1166 )
orginal -> original
* Disable Flex Attention
* Update tokenizer_utils.py
* Update _utils.py
* n_items
* Update cross_entropy_loss.py
* Fix DPO, ORPO
* Update _utils.py
---------
Co-authored-by: timothelaborie <97834767+timothelaborie@users.noreply.github.com>
Co-authored-by: Ikko Eltociear Ashimine <eltociear@gmail.com>
* Add warning for missing Unpack and KwargsForCausalLM in older Transformers versions
---------
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
Co-authored-by: timothelaborie <97834767+timothelaborie@users.noreply.github.com>
Co-authored-by: Ikko Eltociear Ashimine <eltociear@gmail.com>
* Update cross_entropy_loss.py
* Update _utils.py
* Update _utils.py
* donot upcast lm_head and embeddings to float32 (#1186 )
* Cleanup upcast logs (#1188 )
* Fix/phi-longrope (#1193 )
* Enhance rotary embedding handling in LlamaAttention and LongRopeRotaryEmbedding
* Typo
* Improve rotary embedding handling in LlamaAttention to prevent errors with short KV cache
* Update llama.py
* Update llama.py
---------
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
* Update transformers
---------
Co-authored-by: timothelaborie <97834767+timothelaborie@users.noreply.github.com>
Co-authored-by: Ikko Eltociear Ashimine <eltociear@gmail.com>
Co-authored-by: Edd <68678137+Erland366@users.noreply.github.com>
Co-authored-by: Datta Nimmaturi <datta.nimmaturi@nutanix.com>
2024-10-26 01:21:24 -07:00
Daniel Han
49f427ddbe
Update transformers
2024-10-26 01:20:37 -07:00
Edd
e771f2a932
Fix/phi-longrope ( #1193 )
...
* Enhance rotary embedding handling in LlamaAttention and LongRopeRotaryEmbedding
* Typo
* Improve rotary embedding handling in LlamaAttention to prevent errors with short KV cache
* Update llama.py
* Update llama.py
---------
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
2024-10-25 15:44:10 -07:00
Datta Nimmaturi
7bb0fd96f6
Cleanup upcast logs ( #1188 )
2024-10-25 12:17:54 -07:00
Datta Nimmaturi
ffe4ba1a43
donot upcast lm_head and embeddings to float32 ( #1186 )
2024-10-25 01:28:12 -07:00
Daniel Han
2c1c1016c2
Merge branch 'main' into nightly
2024-10-24 12:17:57 -07:00
Daniel Han
828ebf815a
Update _utils.py
2024-10-24 12:17:48 -07:00
Daniel Han
dfdff91258
Fix 4.47 issue ( #1182 )
...
* Fix TRL
* Update mistral.py
* Patch processing_class
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Installation guide (#1165 )
* chore: update chat_templates.py (#1166 )
orginal -> original
* Disable Flex Attention
* Update tokenizer_utils.py
* Update _utils.py
* n_items
* Update cross_entropy_loss.py
* Fix DPO, ORPO
* Update _utils.py
* Update _utils.py
* fix/transformers-unpack (#1180 )
* Fix DPO, ORPO (#1177 )
* Fix TRL
* Update mistral.py
* Patch processing_class
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Installation guide (#1165 )
* chore: update chat_templates.py (#1166 )
orginal -> original
* Disable Flex Attention
* Update tokenizer_utils.py
* Update _utils.py
* n_items
* Update cross_entropy_loss.py
* Fix DPO, ORPO
* Update _utils.py
---------
Co-authored-by: timothelaborie <97834767+timothelaborie@users.noreply.github.com>
Co-authored-by: Ikko Eltociear Ashimine <eltociear@gmail.com>
* Add warning for missing Unpack and KwargsForCausalLM in older Transformers versions
---------
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
Co-authored-by: timothelaborie <97834767+timothelaborie@users.noreply.github.com>
Co-authored-by: Ikko Eltociear Ashimine <eltociear@gmail.com>
* Update cross_entropy_loss.py
* Update _utils.py
* Update _utils.py
---------
Co-authored-by: timothelaborie <97834767+timothelaborie@users.noreply.github.com>
Co-authored-by: Ikko Eltociear Ashimine <eltociear@gmail.com>
Co-authored-by: Edd <68678137+Erland366@users.noreply.github.com>
2024-10-24 12:17:21 -07:00
Daniel Han
3ba142fd62
Update _utils.py
2024-10-24 12:17:09 -07:00
Daniel Han
10f3eedaf7
Update _utils.py
2024-10-24 12:14:14 -07:00
Daniel Han
6e0fa4ec2b
Update cross_entropy_loss.py
2024-10-24 12:11:38 -07:00
Edd
bd6ed7343a
fix/transformers-unpack ( #1180 )
...
* Fix DPO, ORPO (#1177 )
* Fix TRL
* Update mistral.py
* Patch processing_class
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Installation guide (#1165 )
* chore: update chat_templates.py (#1166 )
orginal -> original
* Disable Flex Attention
* Update tokenizer_utils.py
* Update _utils.py
* n_items
* Update cross_entropy_loss.py
* Fix DPO, ORPO
* Update _utils.py
---------
Co-authored-by: timothelaborie <97834767+timothelaborie@users.noreply.github.com>
Co-authored-by: Ikko Eltociear Ashimine <eltociear@gmail.com>
* Add warning for missing Unpack and KwargsForCausalLM in older Transformers versions
---------
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
Co-authored-by: timothelaborie <97834767+timothelaborie@users.noreply.github.com>
Co-authored-by: Ikko Eltociear Ashimine <eltociear@gmail.com>
2024-10-24 12:10:52 -07:00
Daniel Han
ddde04459c
Update _utils.py
2024-10-24 01:11:20 -07:00
Daniel Han
6a90ba64b7
Fix DPO, ORPO ( #1177 )
...
* Fix TRL
* Update mistral.py
* Patch processing_class
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Installation guide (#1165 )
* chore: update chat_templates.py (#1166 )
orginal -> original
* Disable Flex Attention
* Update tokenizer_utils.py
* Update _utils.py
* n_items
* Update cross_entropy_loss.py
* Fix DPO, ORPO
* Update _utils.py
---------
Co-authored-by: timothelaborie <97834767+timothelaborie@users.noreply.github.com>
Co-authored-by: Ikko Eltociear Ashimine <eltociear@gmail.com>
2024-10-24 00:36:37 -07:00
Daniel Han
83c2c564b6
Update _utils.py
2024-10-24 00:25:28 -07:00
Daniel Han
7555e3979c
Merge branch 'main' into nightly
2024-10-24 00:24:27 -07:00
Daniel Han
cd1615dc64
Fix DPO, ORPO
2024-10-24 00:17:26 -07:00
Daniel Han
a048a791f4
Update cross_entropy_loss.py
2024-10-23 22:18:24 -07:00
Daniel Han
8390144201
n_items
2024-10-23 22:13:45 -07:00
Daniel Han
acef7d6a3d
Update _utils.py
2024-10-23 12:39:58 -07:00
Edd
f402d945ff
Fix/patch tokenizer ( #1171 )
...
* fix: correct tokenizer handling in patch_sft_trainer_tokenizer
* Revert "fix: correct tokenizer handling in patch_sft_trainer_tokenizer"
This reverts commit 7a98e465cbd4f980c8b364b0396d44f2d052090f.
* fix: correct condition for test_text assignment in patch_sft_trainer_tokenizer
2024-10-23 12:32:33 -07:00
Daniel Han
e5e67f83b4
Many bug fixes ( #1162 )
...
* Fix TRL
* Update mistral.py
* Patch processing_class
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Installation guide (#1165 )
* chore: update chat_templates.py (#1166 )
orginal -> original
* Disable Flex Attention
* Update tokenizer_utils.py
* Update _utils.py
---------
Co-authored-by: timothelaborie <97834767+timothelaborie@users.noreply.github.com>
Co-authored-by: Ikko Eltociear Ashimine <eltociear@gmail.com>
2024-10-23 03:14:57 -07:00
Daniel Han
93cee86d9c
Update _utils.py
2024-10-23 03:14:48 -07:00
Daniel Han
785a964f20
Update tokenizer_utils.py
2024-10-23 03:04:22 -07:00
Daniel Han
205b0437b3
Disable Flex Attention
2024-10-23 02:58:40 -07:00
Ikko Eltociear Ashimine
a8f79d46c6
chore: update chat_templates.py ( #1166 )
...
orginal -> original
2024-10-23 00:59:02 -07:00
timothelaborie
db0ce5d033
Installation guide ( #1165 )
2024-10-23 00:55:26 -07:00
Daniel Han
45b2689ad6
Update tokenizer_utils.py
2024-10-22 01:28:38 -07:00
Daniel Han
8f0a165a53
Update tokenizer_utils.py
2024-10-22 01:22:13 -07:00
Daniel Han
aaa151b15a
Update tokenizer_utils.py
2024-10-22 01:09:20 -07:00
Daniel Han
11e63948e5
Update tokenizer_utils.py
2024-10-22 01:05:41 -07:00
Daniel Han
a952e13d6a
Update tokenizer_utils.py
2024-10-22 00:57:36 -07:00