Phase 2 results (`scripts/benchmarks/results/grpo_equivalence.md`): - vLLM: 74.4s train, 4.14s median step, 158 GB peak (100%) - unsloth_fi_false: 355s train, 23.95s median, 10.7 GB peak (17%) - cb_paged (via sdpa_paged load): 466s train, 36s median, 55.6 GB peak (11.5%) Coherence gate passes on all three backends: losses finite, rewards in the expected early-GRPO range, KL trajectories qualitatively matched between vLLM and unsloth_fi_false in [0, 0.015]. Memory story is striking: unsloth_fi_false uses 15x less memory than vLLM. qwen3_grpo_unified.py fixes: - Auto-adjust per_device_train_batch_size -> num_generations for vanilla-HF backends (Unsloth's loader does this automatically; TRL on the HF path doesn't and crashes on the divisibility check). - cb_paged now loads with sdpa_paged (not paged_attention). The FA4 paged_attention kernel requires cu_seq_lens_q on every forward, but the GRPO training forward feeds a dense batch without them. sdpa_paged gracefully falls back to plain SDPA in that case and still exercises the paged path during the CB rollout. cb_sync_driver.py fixes: - FIFOScheduler no longer accepts manual_eviction in its signature; dropped. - drive_until_empty used to check has_pending_requests() before calling prepare_next_batch(), which returned False at startup because nothing had yet been pulled from the input_queue. Now the loop drains the input queue first and exits only when both queues + scheduler are empty. Smoke test on GPU 1 (8 prompts, 64 tokens): eager path produces 512 correct tokens; CUDA-graph path hangs during first-step capture (PagedAttentionCache probably allocates on first use). Tracked for the next commit.
93 lines
3.3 KiB
Python
93 lines
3.3 KiB
Python
"""One-shot: materialize a rank-32 LoRA adapter on `unsloth/Qwen3-4B-Base`.
|
|
|
|
Writes a PEFT-style directory so every backend (vLLM `LoRARequest`,
|
|
`peft.PeftModel.from_pretrained`, Unsloth `FastLanguageModel.get_peft_model`)
|
|
can load the SAME weights. Random-init is fine for throughput measurement --
|
|
the goal is to have LoRA kernels active during generation, not a trained
|
|
model.
|
|
|
|
Run:
|
|
CUDA_VISIBLE_DEVICES=6 python scripts/benchmarks/make_lora_adapter.py \
|
|
--output outputs/lora_rank32_fresh
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import argparse
|
|
import os
|
|
from pathlib import Path
|
|
|
|
|
|
def parse_args():
|
|
p = argparse.ArgumentParser()
|
|
p.add_argument("--model_name", default="unsloth/Qwen3-4B-Base")
|
|
p.add_argument("--output", default="outputs/lora_rank32_fresh")
|
|
p.add_argument("--rank", type=int, default=32)
|
|
return p.parse_args()
|
|
|
|
|
|
def main():
|
|
args = parse_args()
|
|
out_dir = Path(args.output).resolve()
|
|
out_dir.mkdir(parents=True, exist_ok=True)
|
|
|
|
# Use vanilla HF -- PEFT's save_pretrained yields the canonical
|
|
# adapter_config.json + adapter_model.safetensors that vLLM's LoRARequest
|
|
# expects. Loading via Unsloth would leak Unsloth-specific LoRA wrappers.
|
|
import torch
|
|
from transformers import AutoModelForCausalLM, AutoTokenizer
|
|
from peft import LoraConfig, get_peft_model
|
|
|
|
tok = AutoTokenizer.from_pretrained(args.model_name)
|
|
if tok.pad_token is None:
|
|
tok.pad_token = tok.eos_token
|
|
|
|
# bf16 base; we only need structure + save. Keep on CPU to avoid a GPU load
|
|
# just for `save_pretrained`.
|
|
print(f"[make_lora_adapter] Loading {args.model_name} on CPU...")
|
|
model = AutoModelForCausalLM.from_pretrained(args.model_name, dtype=torch.bfloat16)
|
|
|
|
peft_cfg = LoraConfig(
|
|
r=args.rank,
|
|
lora_alpha=args.rank * 2,
|
|
target_modules=["q_proj", "k_proj", "v_proj", "o_proj",
|
|
"gate_proj", "up_proj", "down_proj"],
|
|
bias="none",
|
|
task_type="CAUSAL_LM",
|
|
lora_dropout=0.0,
|
|
)
|
|
peft_model = get_peft_model(model, peft_cfg)
|
|
peft_model.print_trainable_parameters()
|
|
|
|
# Ensure both A and B matrices are non-zero. PEFT initializes A with
|
|
# kaiming_uniform and B with zeros -- which makes the adapter a no-op and
|
|
# would mask LoRA kernels on some backends. Seed B with tiny random values.
|
|
n_reinit = 0
|
|
with torch.no_grad():
|
|
for name, p in peft_model.named_parameters():
|
|
if "lora_B" in name:
|
|
p.normal_(mean=0.0, std=1e-4)
|
|
n_reinit += 1
|
|
print(f"[make_lora_adapter] Reinitialized {n_reinit} lora_B matrices with tiny gaussian.")
|
|
|
|
peft_model.save_pretrained(str(out_dir))
|
|
tok.save_pretrained(str(out_dir))
|
|
|
|
# Sanity: verify safetensors file present and non-trivial.
|
|
from safetensors import safe_open
|
|
st_path = out_dir / "adapter_model.safetensors"
|
|
n_zero_tensors = 0
|
|
n_tensors = 0
|
|
with safe_open(str(st_path), framework="pt") as f:
|
|
for key in f.keys():
|
|
t = f.get_tensor(key)
|
|
n_tensors += 1
|
|
if (t == 0).all().item():
|
|
n_zero_tensors += 1
|
|
print(f"[make_lora_adapter] Wrote {n_tensors} tensors to {st_path} "
|
|
f"({n_zero_tensors} all-zero).")
|
|
print(f"[make_lora_adapter] Adapter saved to {out_dir}")
|
|
|
|
|
|
if __name__ == "__main__":
|
|
main()
|