Added the final two entries to results/flex_vs_vllm.md: | Batch | LoRA | vLLM tok/s | flex tok/s | flex/vLLM | |-------|------|-----------:|-----------:|----------:| | 256 | no | 21170 | 7074 | 33 %| flex scales sub-linearly past batch ~128 (7074 @ 256 vs 6501 @ 128 is only a 9 % jump for 2x batch), while vLLM keeps climbing (14996 -> 21170). That's expected: vLLM's chunked prefill + per-step kernel packing is more efficient at huge batches. For GRPO's realistic batch range (4-64 concurrent seqs) the flex path sits at 30-55 % of vLLM. Also tried `FLEX_COMPILE_MODE=max-autotune-no-cudagraphs` on the flex_attention compile (gated via env var). Same throughput as default compile (2186 vs 2189 at batch 32). max-autotune with cudagraphs crashes because it nests its own cudagraph_trees inside our CUDA graph capture and hits `Cannot prepare for replay during capturing stage`. |
||
|---|---|---|
| .. | ||
| benchmarks | ||
| enforce_kwargs_spacing.py | ||
| install_gemma4_mlx.sh | ||
| install_qwen3_6_mlx.sh | ||
| run_ruff_format.py | ||