From 4aae8a678cae377ce5b98f7df9948db15b2e282f Mon Sep 17 00:00:00 2001 From: Daniel Han Date: Wed, 21 Feb 2024 03:58:03 +1100 Subject: [PATCH] Updated Home (markdown) --- Home.md | 15 +++++++++++++++ 1 file changed, 15 insertions(+) diff --git a/Home.md b/Home.md index 8b534e9..9df827a 100644 --- a/Home.md +++ b/Home.md @@ -167,6 +167,21 @@ tokenizer = get_chat_template( ) ``` +### 2x Faster Inference +Unsloth supports natively 2x faster inference. All QLoRA, LoRA and non LoRA inference paths are 2x faster. This requires no change of code or any new dependencies. +```python +from unsloth import FastLanguageModel +model, tokenizer = FastLanguageModel.from_pretrained( + model_name = "lora_model", # YOUR MODEL YOU USED FOR TRAINING + max_seq_length = max_seq_length, + dtype = dtype, + load_in_4bit = load_in_4bit, +) +FastLanguageModel.for_inference(model) # Enable native 2x faster inference +text_streamer = TextStreamer(tokenizer) +_ = model.generate(**inputs, streamer = text_streamer, max_new_tokens = 64) +``` + ### NotImplementedError: A UTF-8 locale is required. Got ANSI See https://github.com/googlecolab/colabtools/issues/3409