Daniel Han
724cec30e9
Update mistral.py
2025-01-16 00:56:56 -08:00
AminWhat
e3a2dbbab6
Torch.Cuda Is Available Condition and Warning ( #1545 )
...
* check for torch.cuda and triton if available
on my machine(mac m3) the cuda were not available
* Update pyproject.toml
* Update __init__.py
---------
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
2025-01-15 23:02:23 -08:00
Daniel Han
bf0fe0e63e
Merge branch 'main' into nightly
2025-01-15 22:55:45 -08:00
Michael Han
d5771b94f4
Merge pull request #1542 from unslothai/shimmyshimmer-patch-3
...
Update README.md
2025-01-14 23:20:19 -08:00
Michael Han
6a4af87df3
Update README.md
...
Update to benchmark tables
2025-01-14 23:20:07 -08:00
Daniel Han
48ded877f1
Update llama.py
2025-01-14 22:32:44 -08:00
Daniel Han
29b6901cef
Merge branch 'main' into nightly
2025-01-14 22:32:02 -08:00
Daniel Han
22dee37138
Update issue templates
2025-01-14 03:13:35 -08:00
Daniel Han
f6415527d7
Update bug_report.md ( #1538 )
2025-01-14 03:12:17 -08:00
Daniel Han
999fd6e192
Update issue templates
2025-01-14 03:10:29 -08:00
Michael Han
2f08550068
Merge pull request #1529 from unslothai/shimmyshimmer-patch-2
...
Update README.md
2025-01-11 17:35:11 -08:00
Michael Han
007a67405f
Update README.md
2025-01-11 17:34:51 -08:00
Michael Han
3de57e0515
Merge pull request #1515 from unslothai/shimmyshimmer-patch-1
...
2025-01
Update README.md for Notebooks
2025-01-10 10:13:04 -08:00
Daniel Han
468f13cd96
Update mapper.py
2025-01-10 04:34:23 -08:00
Michael Han
f064dc942f
Update README.md
2025-01-09 16:59:43 -08:00
Michael Han
35977bb83a
Update README.md
2025-01-08 23:02:27 -08:00
Daniel Han
4a630f3191
Update Unsloth-Zoo
2025-01-08 16:46:04 -08:00
Daniel Han
67b504c8ca
Update _utils.py
2025-01-08 15:48:40 -08:00
Daniel Han
75b4e34ae4
Update tokenizer_utils.py
2025-01-08 15:48:11 -08:00
Daniel Han
3e8cc950b6
Phi-4 bug fix
2025-01-08 15:40:27 -08:00
Daniel Han
254ab1c411
Phi-4 ( #1523 )
...
* use exact model name
* Update save.py
* Update _utils.py
* Update _utils.py
* Update _utils.py
* Update _utils.py
* print
* Update _utils.py
* Update _utils.py
* Update llama.py
* Update _utils.py
* Update vision.py
* Update _utils.py
* Update _utils.py
* Update _utils.py
* Update _utils.py
* Update _utils.py
* Update _utils.py
* Update _utils.py
* Update _utils.py
* Update loader.py
* accurate_accumulation
* Update loader.py
* Update loader.py
* Update _utils.py
* Update loader.py
* Update loader.py
* Update loader.py
* Update loader.py
* Update pyproject.toml
* Update __init__.py
* Update pyproject.toml
* Update __init__.py
* Update __init__.py
* Fix Triton heuristics
https://github.com/triton-lang/triton/issues/5224
* Update __init__.py
* Update __init__.py
* Update __init__.py
* Update __init__.py
* Xformers
* Update loader.py
* Update loader.py
* Rewind
* Update _utils.py
* Update _utils.py
* requires grad
* Update loader.py
* Update _utils.py
* Update loader.py
* changing model to base_model if peft model is already used
* Improve debugging experience (#1512 )
* Create CONTRIBUTING.md (#1472 )
Creating contributing guidelines
* Update CONTRIBUTING.md
improved sentence
* Improve logging control in `unsloth_compile_transformers` by conditionally redirecting stdout based on UNSLOTH_DISABLE_LOGGER environment variable
---------
Co-authored-by: Michael Han <107991372+shimmyshimmer@users.noreply.github.com>
Co-authored-by: Nino Risteski <95188570+NinoRisteski@users.noreply.github.com>
* Update loader.py
* Update llama.py
* Update llama.py
* Revert "Update llama.py"
This reverts commit a8edd0931a .
* Update llama.py
* Update llama.py
* Update llama.py
* Update llama.py
* Update llama.py
* Update llama.py
* Update llama.py
* Update llama.py
* Update llama.py
* Update llama.py
* Update llama.py
* Update llama.py
* Update llama.py
* Auto change is_bfloat16_supported
* Update llama.py
* Force data-type
* Update llama.py
* All attention refactor fix (#1491 )
* change initilization of n_heads, n_kv_heads, hidden_size in llama.py
* do the same for cohere, mistral, gemma2, granite
* do the same for flexattention,cohere, mistral, granite
* Update llama.py
* Update llama.py
* Update granite to work with latest post_patch methods (#1502 )
* Update granite to work with latest post_patch methods
* Pass position_embeddings for granite even if transformers<4.47
* Update llama.py
---------
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
* Minor fixes for granite models (#1503 )
* Update granite.py
Grab residual multiplier directly from layer
* Update llama.py
Version should read >= 4.47.1 as that is the version requiring the changes
* Update granite.py
* Update llama.py
---------
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
* support modelscope models and datasets (#1481 )
* support modelscope
* change modelscope args
* remove useless import
* remove useless import
* fix
* wip
* fix
* remove useless code
* add readme
* add some comments
* change print to raise error
* update comment
* Update loader.py
---------
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
* Merge branch 'main' into nightly
* Phi 4
---------
Co-authored-by: Itsuro Tajima <tajima@georepublic.de>
Co-authored-by: Muhammad Osama <muhammadosama1994@gmail.com>
Co-authored-by: Edd <68678137+Erland366@users.noreply.github.com>
Co-authored-by: Michael Han <107991372+shimmyshimmer@users.noreply.github.com>
Co-authored-by: Nino Risteski <95188570+NinoRisteski@users.noreply.github.com>
Co-authored-by: Kareem <81531392+KareemMusleh@users.noreply.github.com>
Co-authored-by: Datta Nimmaturi <datta.nimmaturi@nutanix.com>
Co-authored-by: Z <coffeevampirebusiness@gmail.com>
Co-authored-by: tastelikefeet <58414341+tastelikefeet@users.noreply.github.com>
2025-01-08 15:10:46 -08:00
Daniel Han
2e33fda307
Merge branch 'main' into nightly
2025-01-08 15:10:13 -08:00
Daniel Han
40a1206e8e
Phi 4
2025-01-08 14:38:41 -08:00
Daniel Han
9811d45059
Merge branch 'main' into nightly
2025-01-08 12:42:18 -08:00
sebaxakerhtc
9413580114
Update __init__.py ( #1520 )
...
* Update __init__.py
This PR is solving the (issue)[https://github.com/unslothai/unsloth/issues/1518 ] with some GPUs
* Update __init__.py
---------
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
2025-01-07 14:51:17 -08:00
Daniel Han
4a216e8106
Update pyproject.toml
2025-01-07 04:29:09 -08:00
Daniel Han
346abbcdc1
Bug fixes ( #1516 )
...
* use exact model name
* Update save.py
* Update _utils.py
* Update _utils.py
* Update _utils.py
* Update _utils.py
* print
* Update _utils.py
* Update _utils.py
* Update llama.py
* Update _utils.py
* Update vision.py
* Update _utils.py
* Update _utils.py
* Update _utils.py
* Update _utils.py
* Update _utils.py
* Update _utils.py
* Update _utils.py
* Update _utils.py
* Update loader.py
* accurate_accumulation
* Update loader.py
* Update loader.py
* Update _utils.py
* Update loader.py
* Update loader.py
* Update loader.py
* Update loader.py
* Update pyproject.toml
* Update __init__.py
* Update pyproject.toml
* Update __init__.py
* Update __init__.py
* Fix Triton heuristics
https://github.com/triton-lang/triton/issues/5224
* Update __init__.py
* Update __init__.py
* Update __init__.py
* Update __init__.py
* Xformers
* Update loader.py
* Update loader.py
* Rewind
* Update _utils.py
* Update _utils.py
* requires grad
* Update loader.py
* Update _utils.py
* Update loader.py
* changing model to base_model if peft model is already used
* Improve debugging experience (#1512 )
* Create CONTRIBUTING.md (#1472 )
Creating contributing guidelines
* Update CONTRIBUTING.md
improved sentence
* Improve logging control in `unsloth_compile_transformers` by conditionally redirecting stdout based on UNSLOTH_DISABLE_LOGGER environment variable
---------
Co-authored-by: Michael Han <107991372+shimmyshimmer@users.noreply.github.com>
Co-authored-by: Nino Risteski <95188570+NinoRisteski@users.noreply.github.com>
* Update loader.py
* Update llama.py
* Update llama.py
* Revert "Update llama.py"
This reverts commit a8edd0931a .
* Update llama.py
* Update llama.py
* Update llama.py
* Update llama.py
* Update llama.py
* Update llama.py
* Update llama.py
* Update llama.py
* Update llama.py
* Update llama.py
* Update llama.py
* Update llama.py
* Update llama.py
* Auto change is_bfloat16_supported
* Update llama.py
* Force data-type
* Update llama.py
* All attention refactor fix (#1491 )
* change initilization of n_heads, n_kv_heads, hidden_size in llama.py
* do the same for cohere, mistral, gemma2, granite
* do the same for flexattention,cohere, mistral, granite
* Update llama.py
* Update llama.py
* Update granite to work with latest post_patch methods (#1502 )
* Update granite to work with latest post_patch methods
* Pass position_embeddings for granite even if transformers<4.47
* Update llama.py
---------
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
* Minor fixes for granite models (#1503 )
* Update granite.py
Grab residual multiplier directly from layer
* Update llama.py
Version should read >= 4.47.1 as that is the version requiring the changes
* Update granite.py
* Update llama.py
---------
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
* support modelscope models and datasets (#1481 )
* support modelscope
* change modelscope args
* remove useless import
* remove useless import
* fix
* wip
* fix
* remove useless code
* add readme
* add some comments
* change print to raise error
* update comment
* Update loader.py
---------
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
---------
Co-authored-by: Itsuro Tajima <tajima@georepublic.de>
Co-authored-by: Muhammad Osama <muhammadosama1994@gmail.com>
Co-authored-by: Edd <68678137+Erland366@users.noreply.github.com>
Co-authored-by: Michael Han <107991372+shimmyshimmer@users.noreply.github.com>
Co-authored-by: Nino Risteski <95188570+NinoRisteski@users.noreply.github.com>
Co-authored-by: Kareem <81531392+KareemMusleh@users.noreply.github.com>
Co-authored-by: Datta Nimmaturi <datta.nimmaturi@nutanix.com>
Co-authored-by: Z <coffeevampirebusiness@gmail.com>
Co-authored-by: tastelikefeet <58414341+tastelikefeet@users.noreply.github.com>
2025-01-07 04:23:14 -08:00
tastelikefeet
8d8fda4e65
support modelscope models and datasets ( #1481 )
...
* support modelscope
* change modelscope args
* remove useless import
* remove useless import
* fix
* wip
* fix
* remove useless code
* add readme
* add some comments
* change print to raise error
* update comment
* Update loader.py
---------
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
2025-01-07 04:09:36 -08:00
Z
7bacfbbaae
Minor fixes for granite models ( #1503 )
...
* Update granite.py
Grab residual multiplier directly from layer
* Update llama.py
Version should read >= 4.47.1 as that is the version requiring the changes
* Update granite.py
* Update llama.py
---------
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
2025-01-07 03:58:40 -08:00
Datta Nimmaturi
cad56301ec
Update granite to work with latest post_patch methods ( #1502 )
...
* Update granite to work with latest post_patch methods
* Pass position_embeddings for granite even if transformers<4.47
* Update llama.py
---------
Co-authored-by: Daniel Han <danielhanchen@gmail.com>
2025-01-07 03:49:11 -08:00
Daniel Han
639ac6f987
Update llama.py
2025-01-07 03:39:08 -08:00
Daniel Han
67a8162a4c
Update llama.py
2025-01-07 03:33:46 -08:00
Kareem
fccdf38c45
All attention refactor fix ( #1491 )
...
* change initilization of n_heads, n_kv_heads, hidden_size in llama.py
* do the same for cohere, mistral, gemma2, granite
* do the same for flexattention,cohere, mistral, granite
2025-01-07 02:41:15 -08:00
Michael Han
feb822755e
Update README.md
...
Notebook links
2025-01-07 02:02:59 -08:00
Daniel Han
a243ddb4f0
Update llama.py
2025-01-07 01:56:32 -08:00
Daniel Han
2ab9a55aff
Force data-type
2025-01-07 01:51:49 -08:00
Daniel Han
fc80163b37
Update llama.py
2025-01-07 01:43:20 -08:00
Daniel Han
5b2569e8fb
Auto change is_bfloat16_supported
2025-01-07 01:40:04 -08:00
Daniel Han
9514818c4b
Update llama.py
2025-01-07 01:10:41 -08:00
Daniel Han
5684cbe502
Update llama.py
2025-01-07 01:05:04 -08:00
Daniel Han
cc69d74d94
Update llama.py
2025-01-07 01:02:35 -08:00
Daniel Han
deb98e9b6b
Update llama.py
2025-01-07 01:02:24 -08:00
Daniel Han
67e3251577
Update llama.py
2025-01-07 00:50:56 -08:00
Daniel Han
2181a64ba9
Update llama.py
2025-01-07 00:45:22 -08:00
Daniel Han
ae3f2e05bb
Update llama.py
2025-01-07 00:41:26 -08:00
Daniel Han
c65fb54914
Update llama.py
2025-01-07 00:41:16 -08:00
Daniel Han
904ec69799
Update llama.py
2025-01-07 00:38:04 -08:00
Daniel Han
7719207b6a
Update llama.py
2025-01-07 00:34:46 -08:00
Daniel Han
fcda5ce87f
Update llama.py
2025-01-07 00:34:09 -08:00
Daniel Han
2d66722781
Update llama.py
2025-01-07 00:30:32 -08:00