unsloth/tests/utils/test_batched_leftpad_generation_gpu.py
Daniel Han 184141db99
Tests + CI guard: batched left-padded generation can never silently regress again (#1066, #3699) (#6145)
* Add regression guard for batched left-padded generation (#1066, #3699)

Three layers of tests plus a path-filtered CI workflow so the left-padding
position_ids / attention-mask bug class cannot silently return:

- tests/utils/test_prepare_inputs_ast_guard.py: import-free AST checks on
  _fast_prepare_inputs_for_generation (cumsum-from-mask branch present,
  cache_position only as fallback, no mask truncation, model families wired)
- tests/utils/test_prepare_inputs_leftpad.py: CPU behavioral unit test with
  synthetic left-padded masks and fake caches; exact expected position_ids
  for prefill and cached decode
- tests/utils/test_batched_leftpad_generation_gpu.py: optional GPU e2e,
  solo vs batched prefix match, skipped without CUDA
- .github/workflows/batch-inference-guard.yml: ubuntu-latest CPU job running
  the two deterministic layers on PRs touching unsloth/models/**

Validated: all pass on main; both CPU layers fail at 6d0f8643~1 (pre #4100)
and at 332eabf3~1 (pre #2216), reproducing the historical bug signatures.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Cite staging proof in batch-inference-guard header (staging-2 PRs 170/171)

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Fold left-padding guard into consolidated Core CI; merge AST + behavioral tests

No new workflow and no new CI job: the guard now runs as one HARD GATE step
inside consolidated-tests-ci.yml, right after the callback signature drift
detector, where the CPU torch stack is already installed. The AST structural
checks and the behavioral unit tests live in a single file
(tests/utils/test_prepare_inputs_leftpad.py); the AST layer stays stdlib-only
with unsloth imported lazily inside the behavioral tests, so import breakage
cannot mask the structural checks.

Revalidated after the merge: 11 assertions pass on main, 8 fail at
6d0f8643~1 (pre #4100).

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Update staging proof reference for consolidated gate (PRs 170/172)

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-06-10 08:00:28 -07:00

107 lines
3.8 KiB
Python

"""End-to-end GPU guard for batched left-padded generation (issues #1066, #3699).
For each prompt, greedy generation inside a left-padded batch must match
generating that prompt alone at batch size 1 for the first PREFIX_TOKENS
tokens, and the full output must not be gibberish. The bug class (#1066,
#3699) makes padded rows diverge immediately into garbage; in contrast,
benign batch-size-dependent kernel numerics can flip a greedy near-tie deep
into the sequence, so an exact full-length match would be flaky. Uses a small
instruct model (chat-templated prompts have high-margin argmaxes).
Skipped automatically when CUDA is unavailable, so CPU CI is unaffected.
Run manually on any GPU box:
python -m pytest tests/utils/test_batched_leftpad_generation_gpu.py -v
"""
import pytest
import torch
cuda_available = torch.cuda.is_available()
pytestmark = pytest.mark.skipif(not cuda_available, reason = "requires a CUDA GPU")
MODEL_NAME = "unsloth/Qwen2.5-0.5B-Instruct"
MAX_NEW_TOKENS = 32
PREFIX_TOKENS = 16
PROMPTS = [
"Give me a short introduction to large language model.",
"Here is an experiment log: "
+ " ".join(f"run {i} completed with stable throughput and no anomalies;" for i in range(1, 41))
+ " In one sentence, what is the overall conclusion?",
]
@pytest.fixture(scope = "module")
def model_and_tokenizer():
from unsloth import FastLanguageModel
model, tokenizer = FastLanguageModel.from_pretrained(
model_name = MODEL_NAME,
max_seq_length = 2048,
load_in_4bit = True,
)
FastLanguageModel.for_inference(model)
tokenizer.padding_side = "left"
if tokenizer.pad_token_id is None:
tokenizer.pad_token_id = tokenizer.eos_token_id
return model, tokenizer
def _chat(tokenizer, prompt):
return tokenizer.apply_chat_template(
[{"role": "user", "content": prompt}],
tokenize = False,
add_generation_prompt = True,
)
def _generate(model, tokenizer, texts):
inputs = tokenizer(texts, return_tensors = "pt", padding = True, add_special_tokens = False).to(
"cuda"
)
with torch.inference_mode():
out = model.generate(
**inputs,
max_new_tokens = MAX_NEW_TOKENS,
do_sample = False,
temperature = None,
top_p = None,
top_k = None,
use_cache = True,
pad_token_id = tokenizer.pad_token_id,
)
suffixes = out[:, inputs["input_ids"].shape[1] :]
return [row.tolist() for row in suffixes]
def _looks_gibberish(text):
if not text.strip():
return True
exclam = text.count("!") / max(len(text), 1)
nonascii = sum(1 for c in text if ord(c) > 0x2FFF) / max(len(text), 1)
return exclam > 0.3 or nonascii > 0.5
def test_batched_leftpad_matches_solo_generation(model_and_tokenizer):
model, tokenizer = model_and_tokenizer
texts = [_chat(tokenizer, p) for p in PROMPTS]
solo = [_generate(model, tokenizer, [t])[0] for t in texts]
batched = _generate(model, tokenizer, texts)
for i, prompt in enumerate(PROMPTS):
solo_text = tokenizer.decode(solo[i], skip_special_tokens = True)
batch_text = tokenizer.decode(batched[i], skip_special_tokens = True)
assert batched[i][:PREFIX_TOKENS] == solo[i][:PREFIX_TOKENS], (
f"prompt {i} ({prompt[:30]!r}...) diverged from solo generation "
f"within the first {PREFIX_TOKENS} tokens inside a left-padded "
"batch; batched left-padded generation is broken again "
f"(issues #1066, #3699).\n"
f"solo : {solo_text!r}\nbatched: {batch_text!r}"
)
assert not _looks_gibberish(batch_text), (
f"prompt {i} produced gibberish in a left-padded batch "
f"(issues #1066, #3699): {batch_text!r}"
)