Run #8 (matrix) failures: - Cells 2 & 3: RecursionError in patch_tiled_mlp shim. Root cause: tests/_zoo_aggressive_cuda_spoof.py routed torch.cuda.manual_seed and manual_seed_all back through torch.manual_seed, but torch.manual_seed internally calls torch.cuda.manual_seed_all -> infinite recursion. Fix: no-op the cuda seed APIs (callers already paid the CPU-RNG cost via torch.manual_seed; CUDA-side seeding has no meaning on a GPU-less runner). Same fix for cuda.set_rng_state / get_rng_state and initial_seed / seed / seed_all. Locally re-validated tiled MLP shim: diff = 0.000e+00, no recursion. - Cell 1: unsloth_zoo's test_every_patched_moe_experts_class_has_lora_extractor fails on transformers==4.57.6 because the MoE class surface unsloth_zoo patches is newer. That's the real drift signal the matrix is supposed to surface; the bug is upstream, not in CI. Keeping it as-is. Per-step `continue-on-error: true` added on every test step so a cell running into one failure (like cell 1's MoE test) still runs the remaining steps (test_apply_fused_lm_head, static checks, runtime patch ledger, tiled MLP, llama-cli smoke). The job-level continue-on-error remains. Drop `pip install --upgrade 'transformers>=4.51,<5.5'` and `'trl>=0.13,<1'` in the static-check steps -- those upgrades would override the matrix-selected versions and defeat the matrix's purpose. The static checks now use whatever versions the runtime-deps step installed for that cell. |
||
|---|---|---|
| .. | ||
| ISSUE_TEMPLATE | ||
| workflows | ||
| CODEOWNERS | ||
| dependabot.yml | ||
| FUNDING.yml | ||