Daniel Han
8090b7c01a
Tied weights
2024-10-31 12:36:21 -07:00
Daniel Han
055eeb8c47
Update cross_entropy_loss.py
2024-10-31 02:00:04 -07:00
Daniel Han
d455751303
Update cross_entropy_loss.py
2024-10-31 01:56:03 -07:00
Daniel Han
7bf626b36b
Update cross_entropy_loss.py
2024-10-31 01:23:41 -07:00
Daniel Han
8aefcd0b50
Update cross_entropy_loss.py
2024-10-30 17:35:20 -07:00
Daniel Han
6db9d286d8
Update cross_entropy_loss.py
2024-10-30 17:26:34 -07:00
Daniel Han
54b901bc50
Update cross_entropy_loss.py
2024-10-30 17:19:30 -07:00
Daniel Han
30cdf652d3
Update cross_entropy_loss.py
2024-10-30 16:55:24 -07:00
Daniel Han
9f926ced28
Update cross_entropy_loss.py
2024-10-30 16:53:04 -07:00
Daniel Han
9920950b7f
Update cross_entropy_loss.py
2024-10-30 16:50:56 -07:00
Daniel Han
d86b20a50f
Update cross_entropy_loss.py
2024-10-30 16:46:33 -07:00
Daniel Han
6d7004b5ca
Update cross_entropy_loss.py
2024-10-30 16:37:37 -07:00
Daniel Han
07394c3436
Update cross_entropy_loss.py
2024-10-30 15:46:09 -07:00
Daniel Han
530c4958e8
Update cross_entropy_loss.py
2024-10-30 15:33:37 -07:00
Daniel Han
251ba77771
Update cross_entropy_loss.py
2024-10-30 14:57:14 -07:00
Daniel Han
526505c119
Update _utils.py
2024-10-30 14:01:49 -07:00
Daniel Han
74ab93c9da
Update _utils.py
2024-10-30 14:00:03 -07:00
Daniel Han
5b75e21a4b
Update _utils.py
2024-10-30 13:54:11 -07:00
Daniel Han
784dd13da7
Update _utils.py
2024-10-30 13:51:10 -07:00
Daniel Han
5f5fef8075
Update __init__.py
2024-10-30 13:47:21 -07:00
Daniel Han
95ecc5795d
Update __init__.py
2024-10-30 13:44:05 -07:00
Daniel Han
9ccbc0ed1f
Update _utils.py
2024-10-30 13:35:01 -07:00
Daniel Han
6bef8f1c3c
Update pyproject.toml
2024-10-30 13:16:00 -07:00
Daniel Han
7e1692ace1
Bug fixes
2024-10-30 13:11:43 -07:00
Daniel Han
20e38eda6b
Feat/all tmp ( #1219 )
...
* Update save.py
Check whether path is in /tmp dir for Kaggle environment
* Update save.py
Move temporary_location to /tmp in Kaggle
* Enhance Kaggle environment support in save and tokenizer utilities
---------
Co-authored-by: dendarrion <37800703+dendarrion@users.noreply.github.com>
Co-authored-by: Erland366 <erland.pg366@gmail.com>
2024-10-30 00:43:03 -07:00
Daniel Han
85a5f6098a
Update cross_entropy_loss.py
2024-10-28 15:01:04 -07:00
Daniel Han
5ee1189657
Update cross_entropy_loss.py
2024-10-28 14:47:11 -07:00
Daniel Han
cac56d112b
Update cross_entropy_loss.py
2024-10-28 14:30:06 -07:00
Daniel Han
c6e9af2e5b
Update _utils.py
2024-10-28 10:41:31 -07:00
Daniel Han
5541ab48fe
Update _utils.py
2024-10-28 01:12:33 -07:00
Daniel Han
2dfdba3493
More patching
2024-10-28 01:10:23 -07:00
Daniel Han
a8b37a320d
Revert "ignored labels"
...
This reverts commit 9d07be077b .
2024-10-27 22:18:05 -07:00
Daniel Han
9d07be077b
ignored labels
2024-10-27 22:10:59 -07:00
Daniel Han
02437a8391
Typo
2024-10-27 19:09:27 -07:00
Daniel Han
5286f19725
Update llama.py
2024-10-27 19:08:14 -07:00
Daniel Han
1c044da660
Fix pad token
2024-10-27 19:06:57 -07:00
Daniel Han
3acc5afad3
Update _utils.py
2024-10-27 17:34:33 -07:00
Daniel Han
7083a1d455
Unk token issues
2024-10-27 17:32:26 -07:00
Daniel Han
bf3b175d39
Merge branch 'main' into nightly
2024-10-27 16:24:20 -07:00
Daniel Han
a2f8db3e73
Merge branch 'main' of https://github.com/unslothai/unsloth
2024-10-27 15:09:42 -07:00
Daniel Han
007efc2751
Update _utils.py
2024-10-27 15:09:35 -07:00
Edd
fdf25b758a
Fix/casting continue pretraining ( #1200 )
...
* Bring back float32 if float16 instead of bfloat16
* Refactor mixed precision handling for lm_head and embed_tokens to ensure correct dtype usage
* Fix dtype retrieval for embed_tokens and lm_head in mixed precision training
* Fix dtype retrieval for embed_tokens and lm_head to use weight dtype in mixed precision training
* Fix dtype handling for embed_tokens and lm_head to ensure correct float32 usage in mixed precision training
* Fix dtype assignment for lm_head modules to ensure correct weight dtype usage in mixed precision training
2024-10-27 15:06:45 -07:00
Daniel Han
49ae619412
Update pyproject.toml
2024-10-26 18:05:55 -07:00
Daniel Han
8d46c0d4d6
Torch 2.5
2024-10-26 18:03:15 -07:00
Daniel Han
f94f7c1882
Merge branch 'main' into nightly
2024-10-26 01:22:21 -07:00
Daniel Han
d76eda4f66
Bug fixes ( #1195 )
...
* Fix TRL
* Update mistral.py
* Patch processing_class
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Installation guide (#1165 )
* chore: update chat_templates.py (#1166 )
orginal -> original
* Disable Flex Attention
* Update tokenizer_utils.py
* Update _utils.py
* n_items
* Update cross_entropy_loss.py
* Fix DPO, ORPO
* Update _utils.py
* Update _utils.py
* fix/transformers-unpack (#1180 )
* Fix DPO, ORPO (#1177 )
* Fix TRL
* Update mistral.py
* Patch processing_class
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Update tokenizer_utils.py
* Installation guide (#1165 )
* chore: update chat_templates.py (#1166 )
orginal -> original
* Disable Flex Attention
* Update tokenizer_utils.py
* Update _utils.py
* n_items
* Update cross_entropy_loss.py
* Fix DPO, ORPO
* Update _utils.py
---------
Co-authored-by: timothelaborie <97834767+timothelaborie@users.noreply.github.com>
Co-authored-by: Ikko Eltociear Ashimine <eltociear@gmail.com>
* Add warning for missing Unpack and KwargsForCausalLM in older Transformers versions
---------
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
Co-authored-by: timothelaborie <97834767+timothelaborie@users.noreply.github.com>
Co-authored-by: Ikko Eltociear Ashimine <eltociear@gmail.com>
* Update cross_entropy_loss.py
* Update _utils.py
* Update _utils.py
* donot upcast lm_head and embeddings to float32 (#1186 )
* Cleanup upcast logs (#1188 )
* Fix/phi-longrope (#1193 )
* Enhance rotary embedding handling in LlamaAttention and LongRopeRotaryEmbedding
* Typo
* Improve rotary embedding handling in LlamaAttention to prevent errors with short KV cache
* Update llama.py
* Update llama.py
---------
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
* Update transformers
---------
Co-authored-by: timothelaborie <97834767+timothelaborie@users.noreply.github.com>
Co-authored-by: Ikko Eltociear Ashimine <eltociear@gmail.com>
Co-authored-by: Edd <68678137+Erland366@users.noreply.github.com>
Co-authored-by: Datta Nimmaturi <datta.nimmaturi@nutanix.com>
2024-10-26 01:21:24 -07:00
Daniel Han
6f28d160b2
Update transformers
2024-10-26 01:20:37 -07:00
Edd
2bc189f490
Fix/phi-longrope ( #1193 )
...
* Enhance rotary embedding handling in LlamaAttention and LongRopeRotaryEmbedding
* Typo
* Improve rotary embedding handling in LlamaAttention to prevent errors with short KV cache
* Update llama.py
* Update llama.py
---------
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
2024-10-25 15:44:10 -07:00
Datta Nimmaturi
625209e11f
Cleanup upcast logs ( #1188 )
2024-10-25 12:17:54 -07:00
Datta Nimmaturi
67760559e4
donot upcast lm_head and embeddings to float32 ( #1186 )
2024-10-25 01:28:12 -07:00