Commit graph

4,859 commits

Author SHA1 Message Date
Daniel Han
c4dc08309e refix rope 2024-07-23 10:53:31 -07:00
Daniel Han
f285c33046 patch RoPE 2024-07-23 10:48:45 -07:00
Daniel Han
d587ce218d hack for rotary 2024-07-23 10:43:36 -07:00
Daniel Han
c17e8ca33d Update llama.py 2024-07-23 10:36:06 -07:00
Daniel Han
9a18fee63f Update llama.py 2024-07-23 10:35:03 -07:00
Daniel Han
19cf853157 Update _utils.py 2024-07-23 10:33:07 -07:00
Daniel Han
5efbd701ba Llama 3.1 2024-07-23 10:27:36 -07:00
Daniel Han
daa4d13564 Update _utils.py 2024-07-22 23:01:18 -07:00
Daniel Han
eda2343056 Llama 3.1 2024-07-22 22:58:02 -07:00
Daniel Han
0690914c62 Update tokenizer_utils.py 2024-07-20 13:25:59 -07:00
Daniel Han
71c4aed1be Update tokenizer_utils.py 2024-07-20 13:22:36 -07:00
Daniel Han
50c51e6ec1 Update llama.py 2024-07-20 12:47:36 -07:00
Daniel Han
07228828a0 Update llama.py 2024-07-20 11:53:32 -07:00
Daniel Han
11dcf38761 Merge branch 'main' into nightly 2024-07-20 09:47:22 -07:00
Daniel Han
32466f7bc4 Update mistral.py 2024-07-19 09:32:27 -07:00
Daniel Han
bffc936663 Fix Gemma 2024-07-19 09:27:18 -07:00
Daniel Han
256b55fcdd Update README.md 2024-07-19 03:05:15 -07:00
Daniel Han
b8e6560b8d Update README.md 2024-07-19 03:03:50 -07:00
Daniel Han
100ac9c052 Nightly (#784)
* Update __init__.py

* dynamic RoPE

* Update mistral.py

* Update llama.py

* Update tokenizer_utils.py

* Update mistral.py

* Update llama.py

* Update __init__.py

* Update flex_attention.py

* Update llama.py

* Update llama.py

* Mistral Nemo

* Update tokenizer_utils.py

* Update tokenizer_utils.py

* Update tokenizer_utils.py
2024-07-19 01:39:08 -07:00
Daniel Han
e6002b1b32 Merge branch 'main' into nightly 2024-07-19 01:38:47 -07:00
Daniel Han
8ae30938d3 Update tokenizer_utils.py 2024-07-19 01:29:52 -07:00
Daniel Han
f6d47c99df Nightly (#783)
* Update __init__.py

* dynamic RoPE

* Update mistral.py

* Update llama.py

* Update tokenizer_utils.py

* Update mistral.py

* Update llama.py

* Update __init__.py

* Update flex_attention.py

* Update llama.py

* Update llama.py

* Mistral Nemo

* Update tokenizer_utils.py

* Update tokenizer_utils.py
2024-07-19 01:24:46 -07:00
Daniel Han
f302c074b7 Update tokenizer_utils.py 2024-07-19 01:06:41 -07:00
Daniel Han
14144ad6fc Update tokenizer_utils.py 2024-07-19 01:03:51 -07:00
Daniel Han
881ee0ed37 Merge branch 'main' into nightly 2024-07-19 00:59:48 -07:00
Daniel Han
8783596962 Update tokenizer_utils.py 2024-07-19 00:41:35 -07:00
Daniel Han
47e08076e6 Mistral Nemo (#782)
* Update __init__.py

* dynamic RoPE

* Update mistral.py

* Update llama.py

* Update tokenizer_utils.py

* Update mistral.py

* Update llama.py

* Update __init__.py

* Update flex_attention.py

* Update llama.py

* Update llama.py

* Mistral Nemo
2024-07-19 00:14:24 -07:00
Daniel Han
2f9556c428 Mistral Nemo 2024-07-18 22:53:23 -07:00
Daniel Han
187157f548 Update llama.py 2024-07-18 22:07:23 -07:00
Daniel Han
ebbbf6be52 Update llama.py 2024-07-18 21:57:33 -07:00
Daniel Han
e1dc32c2a6 Merge branch 'main' into nightly 2024-07-18 21:55:10 -07:00
Daniel Han
742a7629c2 Fix bugs (#779)
* Update __init__.py

* dynamic RoPE

* Update mistral.py

* Update llama.py

* Update tokenizer_utils.py

* Update mistral.py

* Update llama.py

* Update __init__.py

* Update flex_attention.py
2024-07-18 18:19:24 -07:00
Daniel Han
0da004c70e Update flex_attention.py 2024-07-18 18:18:09 -07:00
Daniel Han
d4fa9a0cdf Update __init__.py 2024-07-18 14:43:03 -07:00
Daniel Han
765a7a9330 Update llama.py 2024-07-18 14:31:35 -07:00
Daniel Han
125b3727ff Update mistral.py 2024-07-18 13:33:13 -07:00
Daniel Han
fcac73786c Update tokenizer_utils.py 2024-07-18 13:25:30 -07:00
Daniel Han
1144bbb15c Update llama.py 2024-07-18 12:33:22 -07:00
Daniel Han
72d9e5f5a0 Update mistral.py 2024-07-18 12:08:49 -07:00
Daniel Han
54dd81de67 dynamic RoPE 2024-07-18 12:06:38 -07:00
Daniel Han
5dc52e6b2e Update __init__.py 2024-07-18 11:07:32 -07:00
Daniel Han
66ce2d401a Update pyproject.toml 2024-07-18 10:59:09 -07:00
Daniel Han
1a7c3e1b3c Update __init__.py 2024-07-18 10:58:12 -07:00
Daniel Han
6a437e43f5 Mistral Nemo 12b (#777)
* Update gemma2.py

* Update llama.py

* Update llama.py

* Update gemma2.py

* init

* Update gemma2.py

* Update gemma2.py

* Update _utils.py

* Update _utils.py

* Update _utils.py

* Update _utils.py

* Update _utils.py

* Update gemma2.py

* Update gemma2.py

* Update gemma2.py

* All RoPE Scaling support

* cleanup

* Update llama.py

* Update llama.py

* Update _utils.py

* Update _utils.py

* exec

* exec

* Attention_Module

* attention_module

* imports

* exec

* Update llama.py

* Update llama.py

* boolean mask

* revert masking

* Update llama.py

* Update save.py

* Update llama.py

* Update gemma2.py

* Update gemma2.py

* Update gemma2.py

* Update utils.py

* retry

* Update gemma2.py

* Update gemma2.py

* Update gemma2.py

* Update _utils.py

* Update _utils.py

* Update gemma2.py

* Update chat_templates.py

* Gemma 2 Ollama support

* Update llama.py

* Update llama.py

* error handling

* Update _utils.py

* Update _utils.py

* Stats for debugging

* Update _utils.py

* Update _utils.py

* Debugging

* Update tokenizer_utils.py

* Update _utils.py

* Update cross_entropy_loss.py

* Update cross_entropy_loss.py

* Update cross_entropy_loss.py

* Update rms_layernorm.py

* Update rms_layernorm.py

* Update rms_layernorm.py

* Update rms_layernorm.py

* Update rms_layernorm.py

* Update rms_layernorm.py

* Update rms_layernorm.py

* Check exec, eval

* Update _utils.py

* Update _utils.py

* Images

* Bug fixes

* Update pyproject.toml

* Bug fixes

* Update _utils.py

* Update _utils.py

* Deprecation fix

* Update chat_templates.py

* Now permitting use of pre-installed llama.cpp (#763)

* Now permitting use of pre-installed llama.cpp

* Update save.py

---------

Co-authored-by: Giuseppe Strafforello <giuseppe.strafforello@titantechnologies.com>
Co-authored-by: Daniel Han <danielhanchen@gmail.com>

* Update save.py

* Deprecation & compile

* typo

* Update chat_templates.py

* Update chat_templates.py

* train_on_responses_only

* Update llama.py

* Update llama.py

* Update save.py

* Update gemma2.py

* Flex Attention

* typos

* Update _utils.py

* Update llama.py

* Update __init__.py

* Update flex_attention.py

* Update llama.py

* Update llama.py

* emulation

* Update __init__.py

* Update rope_embedding.py

* Update flex_attention.py

* Update flex_attention.py

* Update rope_embedding.py

* libdevice

* triton_tanh

* Update flex_attention.py

* Update flex_attention.py

* Update flex_attention.py

* Update flex_attention.py

* Update flex_attention.py

* score

* Update llama.py

* Update flex_attention.py

* Update flex_attention.py

* Update flex_attention.py

* Update flex_attention.py

* Update flex_attention.py

* Update llama.py

* Update flex_attention.py

* Update flex_attention.py

* Update flex_attention.py

* Update flex_attention.py

* Flex Attention removal

* upload tensorboard training stats to hub if available (#773)

* causal_mask

* Update llama.py

* Update llama.py

* Update flex_attention.py

* Update _utils.py

* Update mapper.py

* Update _utils.py

---------

Co-authored-by: pepistrafforello <pepi.strafforello@gmail.com>
Co-authored-by: Giuseppe Strafforello <giuseppe.strafforello@titantechnologies.com>
Co-authored-by: Sébastien De Greef <sebdg@binarycompany.com>
2024-07-18 10:51:10 -07:00
Daniel Han
fa893e7d67 Chat templates 2024-07-15 14:36:44 -07:00
Daniel Han
ca6c3dcc99 Train on responses only (#770)
* Update gemma2.py

* Update llama.py

* Update llama.py

* Update gemma2.py

* init

* Update gemma2.py

* Update gemma2.py

* Update _utils.py

* Update _utils.py

* Update _utils.py

* Update _utils.py

* Update _utils.py

* Update gemma2.py

* Update gemma2.py

* Update gemma2.py

* All RoPE Scaling support

* cleanup

* Update llama.py

* Update llama.py

* Update _utils.py

* Update _utils.py

* exec

* exec

* Attention_Module

* attention_module

* imports

* exec

* Update llama.py

* Update llama.py

* boolean mask

* revert masking

* Update llama.py

* Update save.py

* Update llama.py

* Update gemma2.py

* Update gemma2.py

* Update gemma2.py

* Update utils.py

* retry

* Update gemma2.py

* Update gemma2.py

* Update gemma2.py

* Update _utils.py

* Update _utils.py

* Update gemma2.py

* Update chat_templates.py

* Gemma 2 Ollama support

* Update llama.py

* Update llama.py

* error handling

* Update _utils.py

* Update _utils.py

* Stats for debugging

* Update _utils.py

* Update _utils.py

* Debugging

* Update tokenizer_utils.py

* Update _utils.py

* Update cross_entropy_loss.py

* Update cross_entropy_loss.py

* Update cross_entropy_loss.py

* Update rms_layernorm.py

* Update rms_layernorm.py

* Update rms_layernorm.py

* Update rms_layernorm.py

* Update rms_layernorm.py

* Update rms_layernorm.py

* Update rms_layernorm.py

* Check exec, eval

* Update _utils.py

* Update _utils.py

* Images

* Bug fixes

* Update pyproject.toml

* Bug fixes

* Update _utils.py

* Update _utils.py

* Deprecation fix

* Update chat_templates.py

* Now permitting use of pre-installed llama.cpp (#763)

* Now permitting use of pre-installed llama.cpp

* Update save.py

---------

Co-authored-by: Giuseppe Strafforello <giuseppe.strafforello@titantechnologies.com>
Co-authored-by: Daniel Han <danielhanchen@gmail.com>

* Update save.py

* Deprecation & compile

* typo

* Update chat_templates.py

* Update chat_templates.py

* train_on_responses_only

* Update llama.py

* Update llama.py

* Update save.py

* Update gemma2.py

---------

Co-authored-by: pepistrafforello <pepi.strafforello@gmail.com>
Co-authored-by: Giuseppe Strafforello <giuseppe.strafforello@titantechnologies.com>
2024-07-14 22:41:04 -07:00
Daniel Han
f176cbd36a Many bug fixes (#754)
* Update gemma2.py

* Update llama.py

* Update llama.py

* Update gemma2.py

* init

* Update gemma2.py

* Update gemma2.py

* Update _utils.py

* Update _utils.py

* Update _utils.py

* Update _utils.py

* Update _utils.py

* Update gemma2.py

* Update gemma2.py

* Update gemma2.py

* All RoPE Scaling support

* cleanup

* Update llama.py

* Update llama.py

* Update _utils.py

* Update _utils.py

* exec

* exec

* Attention_Module

* attention_module

* imports

* exec

* Update llama.py

* Update llama.py

* boolean mask

* revert masking

* Update llama.py

* Update save.py

* Update llama.py

* Update gemma2.py

* Update gemma2.py

* Update gemma2.py

* Update utils.py

* retry

* Update gemma2.py

* Update gemma2.py

* Update gemma2.py

* Update _utils.py

* Update _utils.py

* Update gemma2.py

* Update chat_templates.py

* Gemma 2 Ollama support

* Update llama.py

* Update llama.py

* error handling

* Update _utils.py

* Update _utils.py

* Stats for debugging

* Update _utils.py

* Update _utils.py

* Debugging

* Update tokenizer_utils.py

* Update _utils.py

* Update cross_entropy_loss.py

* Update cross_entropy_loss.py

* Update cross_entropy_loss.py

* Update rms_layernorm.py

* Update rms_layernorm.py

* Update rms_layernorm.py

* Update rms_layernorm.py

* Update rms_layernorm.py

* Update rms_layernorm.py

* Update rms_layernorm.py

* Check exec, eval

* Update _utils.py

* Update _utils.py

* Images

* Bug fixes

* Update pyproject.toml

* Bug fixes

* Update _utils.py

* Update _utils.py
2024-07-10 01:59:06 -07:00
Daniel Han
316aaefdf2 Update llama.py 2024-07-08 10:44:19 -07:00
Daniel Han
2eb950872a Update llama.py 2024-07-08 10:38:49 -07:00
Daniel Han
1f1211fbd6 Update _utils.py 2024-07-08 10:01:20 -07:00