Commit graph

1,148 commits

Author SHA1 Message Date
Datta Nimmaturi
cc33bd05dc Add dropout to granite to match HF's implementation (#1557)
Signed-off-by: datta0 <venkatadattasainimmaturi@gmail.com>
2025-01-19 03:54:12 -08:00
Daniel Han
76debb6817 Bug fixes 2025-01-16 03:09:02 -08:00
Daniel Han
69b09d80c8 Fix 2025-01-16 01:22:13 -08:00
Daniel Han
b99b2c5211 Update _utils.py 2025-01-16 01:18:15 -08:00
Daniel Han
7b144ebdaa Update _utils.py 2025-01-16 01:15:42 -08:00
Daniel Han
8cbfdd5aac Update _utils.py 2025-01-16 01:10:40 -08:00
Daniel Han
1262831bf2 Update _utils.py 2025-01-16 01:09:23 -08:00
Daniel Han
cf3291f278 Update _utils.py 2025-01-16 01:07:23 -08:00
Daniel Han
8ed010b4a5 Update mistral.py 2025-01-16 00:58:46 -08:00
Daniel Han
724cec30e9 Update mistral.py 2025-01-16 00:56:56 -08:00
AminWhat
e3a2dbbab6 Torch.Cuda Is Available Condition and Warning (#1545)
* check for torch.cuda and triton if available
on my machine(mac m3) the cuda were not available

* Update pyproject.toml

* Update __init__.py

---------

Co-authored-by: Daniel Han <danielhanchen@gmail.com>
2025-01-15 23:02:23 -08:00
Daniel Han
bf0fe0e63e Merge branch 'main' into nightly 2025-01-15 22:55:45 -08:00
Michael Han
d5771b94f4 Merge pull request #1542 from unslothai/shimmyshimmer-patch-3
Update README.md
2025-01-14 23:20:19 -08:00
Michael Han
6a4af87df3 Update README.md
Update to benchmark tables
2025-01-14 23:20:07 -08:00
Daniel Han
48ded877f1 Update llama.py 2025-01-14 22:32:44 -08:00
Daniel Han
29b6901cef Merge branch 'main' into nightly 2025-01-14 22:32:02 -08:00
Daniel Han
22dee37138 Update issue templates 2025-01-14 03:13:35 -08:00
Daniel Han
f6415527d7 Update bug_report.md (#1538) 2025-01-14 03:12:17 -08:00
Daniel Han
999fd6e192 Update issue templates 2025-01-14 03:10:29 -08:00
Michael Han
2f08550068 Merge pull request #1529 from unslothai/shimmyshimmer-patch-2
Update README.md
2025-01-11 17:35:11 -08:00
Michael Han
007a67405f Update README.md 2025-01-11 17:34:51 -08:00
Michael Han
3de57e0515 Merge pull request #1515 from unslothai/shimmyshimmer-patch-1 2025-01
Update README.md for Notebooks
2025-01-10 10:13:04 -08:00
Daniel Han
468f13cd96 Update mapper.py 2025-01-10 04:34:23 -08:00
Michael Han
f064dc942f Update README.md 2025-01-09 16:59:43 -08:00
Michael Han
35977bb83a Update README.md 2025-01-08 23:02:27 -08:00
Daniel Han
4a630f3191 Update Unsloth-Zoo 2025-01-08 16:46:04 -08:00
Daniel Han
67b504c8ca Update _utils.py 2025-01-08 15:48:40 -08:00
Daniel Han
75b4e34ae4 Update tokenizer_utils.py 2025-01-08 15:48:11 -08:00
Daniel Han
3e8cc950b6 Phi-4 bug fix 2025-01-08 15:40:27 -08:00
Daniel Han
254ab1c411 Phi-4 (#1523)
* use exact model name

* Update save.py

* Update _utils.py

* Update _utils.py

* Update _utils.py

* Update _utils.py

* print

* Update _utils.py

* Update _utils.py

* Update llama.py

* Update _utils.py

* Update vision.py

* Update _utils.py

* Update _utils.py

* Update _utils.py

* Update _utils.py

* Update _utils.py

* Update _utils.py

* Update _utils.py

* Update _utils.py

* Update loader.py

* accurate_accumulation

* Update loader.py

* Update loader.py

* Update _utils.py

* Update loader.py

* Update loader.py

* Update loader.py

* Update loader.py

* Update pyproject.toml

* Update __init__.py

* Update pyproject.toml

* Update __init__.py

* Update __init__.py

* Fix Triton heuristics

https://github.com/triton-lang/triton/issues/5224

* Update __init__.py

* Update __init__.py

* Update __init__.py

* Update __init__.py

* Xformers

* Update loader.py

* Update loader.py

* Rewind

* Update _utils.py

* Update _utils.py

* requires grad

* Update loader.py

* Update _utils.py

* Update loader.py

* changing model to base_model if peft model is already used

* Improve debugging experience (#1512)

* Create CONTRIBUTING.md (#1472)

Creating contributing guidelines

* Update CONTRIBUTING.md

improved sentence

* Improve logging control in `unsloth_compile_transformers` by conditionally redirecting stdout based on UNSLOTH_DISABLE_LOGGER environment variable

---------

Co-authored-by: Michael Han <107991372+shimmyshimmer@users.noreply.github.com>
Co-authored-by: Nino Risteski <95188570+NinoRisteski@users.noreply.github.com>

* Update loader.py

* Update llama.py

* Update llama.py

* Revert "Update llama.py"

This reverts commit a8edd0931a.

* Update llama.py

* Update llama.py

* Update llama.py

* Update llama.py

* Update llama.py

* Update llama.py

* Update llama.py

* Update llama.py

* Update llama.py

* Update llama.py

* Update llama.py

* Update llama.py

* Update llama.py

* Auto change is_bfloat16_supported

* Update llama.py

* Force data-type

* Update llama.py

* All attention refactor fix (#1491)

* change initilization of n_heads, n_kv_heads, hidden_size in llama.py

* do the same for cohere, mistral, gemma2, granite

* do the same for flexattention,cohere, mistral, granite

* Update llama.py

* Update llama.py

* Update granite to work with latest post_patch methods (#1502)

* Update granite to work with latest post_patch methods

* Pass position_embeddings for granite even if transformers<4.47

* Update llama.py

---------

Co-authored-by: Daniel Han <danielhanchen@gmail.com>

* Minor fixes for granite models (#1503)

* Update granite.py

Grab residual multiplier directly from layer

* Update llama.py

Version should read >= 4.47.1 as that is the version requiring the changes

* Update granite.py

* Update llama.py

---------

Co-authored-by: Daniel Han <danielhanchen@gmail.com>

* support modelscope models and datasets (#1481)

* support modelscope

* change modelscope args

* remove useless import

* remove useless import

* fix

* wip

* fix

* remove useless code

* add readme

* add some comments

* change print to raise error

* update comment

* Update loader.py

---------

Co-authored-by: Daniel Han <danielhanchen@gmail.com>

* Merge branch 'main' into nightly

* Phi 4

---------

Co-authored-by: Itsuro Tajima <tajima@georepublic.de>
Co-authored-by: Muhammad Osama <muhammadosama1994@gmail.com>
Co-authored-by: Edd <68678137+Erland366@users.noreply.github.com>
Co-authored-by: Michael Han <107991372+shimmyshimmer@users.noreply.github.com>
Co-authored-by: Nino Risteski <95188570+NinoRisteski@users.noreply.github.com>
Co-authored-by: Kareem <81531392+KareemMusleh@users.noreply.github.com>
Co-authored-by: Datta Nimmaturi <datta.nimmaturi@nutanix.com>
Co-authored-by: Z <coffeevampirebusiness@gmail.com>
Co-authored-by: tastelikefeet <58414341+tastelikefeet@users.noreply.github.com>
2025-01-08 15:10:46 -08:00
Daniel Han
2e33fda307 Merge branch 'main' into nightly 2025-01-08 15:10:13 -08:00
Daniel Han
40a1206e8e Phi 4 2025-01-08 14:38:41 -08:00
Daniel Han
9811d45059 Merge branch 'main' into nightly 2025-01-08 12:42:18 -08:00
sebaxakerhtc
9413580114 Update __init__.py (#1520)
* Update __init__.py

This PR is solving the (issue)[https://github.com/unslothai/unsloth/issues/1518] with some GPUs

* Update __init__.py

---------

Co-authored-by: Daniel Han <danielhanchen@gmail.com>
2025-01-07 14:51:17 -08:00
Daniel Han
4a216e8106 Update pyproject.toml 2025-01-07 04:29:09 -08:00
Daniel Han
346abbcdc1 Bug fixes (#1516)
* use exact model name

* Update save.py

* Update _utils.py

* Update _utils.py

* Update _utils.py

* Update _utils.py

* print

* Update _utils.py

* Update _utils.py

* Update llama.py

* Update _utils.py

* Update vision.py

* Update _utils.py

* Update _utils.py

* Update _utils.py

* Update _utils.py

* Update _utils.py

* Update _utils.py

* Update _utils.py

* Update _utils.py

* Update loader.py

* accurate_accumulation

* Update loader.py

* Update loader.py

* Update _utils.py

* Update loader.py

* Update loader.py

* Update loader.py

* Update loader.py

* Update pyproject.toml

* Update __init__.py

* Update pyproject.toml

* Update __init__.py

* Update __init__.py

* Fix Triton heuristics

https://github.com/triton-lang/triton/issues/5224

* Update __init__.py

* Update __init__.py

* Update __init__.py

* Update __init__.py

* Xformers

* Update loader.py

* Update loader.py

* Rewind

* Update _utils.py

* Update _utils.py

* requires grad

* Update loader.py

* Update _utils.py

* Update loader.py

* changing model to base_model if peft model is already used

* Improve debugging experience (#1512)

* Create CONTRIBUTING.md (#1472)

Creating contributing guidelines

* Update CONTRIBUTING.md

improved sentence

* Improve logging control in `unsloth_compile_transformers` by conditionally redirecting stdout based on UNSLOTH_DISABLE_LOGGER environment variable

---------

Co-authored-by: Michael Han <107991372+shimmyshimmer@users.noreply.github.com>
Co-authored-by: Nino Risteski <95188570+NinoRisteski@users.noreply.github.com>

* Update loader.py

* Update llama.py

* Update llama.py

* Revert "Update llama.py"

This reverts commit a8edd0931a.

* Update llama.py

* Update llama.py

* Update llama.py

* Update llama.py

* Update llama.py

* Update llama.py

* Update llama.py

* Update llama.py

* Update llama.py

* Update llama.py

* Update llama.py

* Update llama.py

* Update llama.py

* Auto change is_bfloat16_supported

* Update llama.py

* Force data-type

* Update llama.py

* All attention refactor fix (#1491)

* change initilization of n_heads, n_kv_heads, hidden_size in llama.py

* do the same for cohere, mistral, gemma2, granite

* do the same for flexattention,cohere, mistral, granite

* Update llama.py

* Update llama.py

* Update granite to work with latest post_patch methods (#1502)

* Update granite to work with latest post_patch methods

* Pass position_embeddings for granite even if transformers<4.47

* Update llama.py

---------

Co-authored-by: Daniel Han <danielhanchen@gmail.com>

* Minor fixes for granite models (#1503)

* Update granite.py

Grab residual multiplier directly from layer

* Update llama.py

Version should read >= 4.47.1 as that is the version requiring the changes

* Update granite.py

* Update llama.py

---------

Co-authored-by: Daniel Han <danielhanchen@gmail.com>

* support modelscope models and datasets (#1481)

* support modelscope

* change modelscope args

* remove useless import

* remove useless import

* fix

* wip

* fix

* remove useless code

* add readme

* add some comments

* change print to raise error

* update comment

* Update loader.py

---------

Co-authored-by: Daniel Han <danielhanchen@gmail.com>

---------

Co-authored-by: Itsuro Tajima <tajima@georepublic.de>
Co-authored-by: Muhammad Osama <muhammadosama1994@gmail.com>
Co-authored-by: Edd <68678137+Erland366@users.noreply.github.com>
Co-authored-by: Michael Han <107991372+shimmyshimmer@users.noreply.github.com>
Co-authored-by: Nino Risteski <95188570+NinoRisteski@users.noreply.github.com>
Co-authored-by: Kareem <81531392+KareemMusleh@users.noreply.github.com>
Co-authored-by: Datta Nimmaturi <datta.nimmaturi@nutanix.com>
Co-authored-by: Z <coffeevampirebusiness@gmail.com>
Co-authored-by: tastelikefeet <58414341+tastelikefeet@users.noreply.github.com>
2025-01-07 04:23:14 -08:00
tastelikefeet
8d8fda4e65 support modelscope models and datasets (#1481)
* support modelscope

* change modelscope args

* remove useless import

* remove useless import

* fix

* wip

* fix

* remove useless code

* add readme

* add some comments

* change print to raise error

* update comment

* Update loader.py

---------

Co-authored-by: Daniel Han <danielhanchen@gmail.com>
2025-01-07 04:09:36 -08:00
Z
7bacfbbaae Minor fixes for granite models (#1503)
* Update granite.py

Grab residual multiplier directly from layer

* Update llama.py

Version should read >= 4.47.1 as that is the version requiring the changes

* Update granite.py

* Update llama.py

---------

Co-authored-by: Daniel Han <danielhanchen@gmail.com>
2025-01-07 03:58:40 -08:00
Datta Nimmaturi
cad56301ec Update granite to work with latest post_patch methods (#1502)
* Update granite to work with latest post_patch methods

* Pass position_embeddings for granite even if transformers<4.47

* Update llama.py

---------

Co-authored-by: Daniel Han <danielhanchen@gmail.com>
2025-01-07 03:49:11 -08:00
Daniel Han
639ac6f987 Update llama.py 2025-01-07 03:39:08 -08:00
Daniel Han
67a8162a4c Update llama.py 2025-01-07 03:33:46 -08:00
Kareem
fccdf38c45 All attention refactor fix (#1491)
* change initilization of n_heads, n_kv_heads, hidden_size in llama.py

* do the same for cohere, mistral, gemma2, granite

* do the same for flexattention,cohere, mistral, granite
2025-01-07 02:41:15 -08:00
Michael Han
feb822755e Update README.md
Notebook links
2025-01-07 02:02:59 -08:00
Daniel Han
a243ddb4f0 Update llama.py 2025-01-07 01:56:32 -08:00
Daniel Han
2ab9a55aff Force data-type 2025-01-07 01:51:49 -08:00
Daniel Han
fc80163b37 Update llama.py 2025-01-07 01:43:20 -08:00
Daniel Han
5b2569e8fb Auto change is_bfloat16_supported 2025-01-07 01:40:04 -08:00
Daniel Han
9514818c4b Update llama.py 2025-01-07 01:10:41 -08:00
Daniel Han
5684cbe502 Update llama.py 2025-01-07 01:05:04 -08:00
Daniel Han
cc69d74d94 Update llama.py 2025-01-07 01:02:35 -08:00