..
flex_8_tuned.json
FlexKernelOptions sweep: flex reaches 72% of vLLM at batch 64 + LoRA
2026-04-20 23:30:05 +00:00
flex_16_tuned.json
FlexKernelOptions sweep: flex reaches 72% of vLLM at batch 64 + LoRA
2026-04-20 23:30:05 +00:00
flex_32_tuned.json
FlexKernelOptions sweep: flex reaches 72% of vLLM at batch 64 + LoRA
2026-04-20 23:30:05 +00:00
flex_32x512_cudagraph.json
Breakthrough: flex_attention + paged KV + CUDA graphs = 35-48% of vLLM
2026-04-20 15:57:41 +00:00
flex_32x512_lora_cudagraph.json
Flex+CUDA-graph closes gap to vLLM across batch sizes
2026-04-20 16:09:47 +00:00
flex_64_lora_autotune.json
flex: test FA4 prefill + Inductor autotune replay (both regress)
2026-04-21 00:25:11 +00:00
flex_64_lora_autotune_tma.json
flex: test FA4 prefill + Inductor autotune replay (both regress)
2026-04-21 00:25:11 +00:00
flex_64_lora_fa4prefill.json
flex: test FA4 prefill + Inductor autotune replay (both regress)
2026-04-21 00:25:11 +00:00
flex_64_lora_pinned_blocks.json
flex: test FA4 prefill + Inductor autotune replay (both regress)
2026-04-21 00:25:11 +00:00
flex_64_lora_torch211_10rounds.json
flex: test FA4 prefill + Inductor autotune replay (both regress)
2026-04-21 00:25:11 +00:00
flex_64_lora_torch211_baseline.json
flex: test FA4 prefill + Inductor autotune replay (both regress)
2026-04-21 00:25:11 +00:00
flex_64_lora_torch211_repeat.json
flex: test FA4 prefill + Inductor autotune replay (both regress)
2026-04-21 00:25:11 +00:00
flex_64_lora_tuned.json
FlexKernelOptions sweep: flex reaches 72% of vLLM at batch 64 + LoRA
2026-04-20 23:30:05 +00:00
flex_64_tuned.json
FlexKernelOptions sweep: flex reaches 72% of vLLM at batch 64 + LoRA
2026-04-20 23:30:05 +00:00
flex_64x512_lora_cudagraph.json
Flex+CUDA-graph closes gap to vLLM across batch sizes
2026-04-20 16:09:47 +00:00
flex_128_tuned.json
FlexKernelOptions sweep: flex reaches 72% of vLLM at batch 64 + LoRA
2026-04-20 23:30:05 +00:00
flex_128x512_cudagraph.json
Flex+CUDA-graph closes gap to vLLM across batch sizes
2026-04-20 16:09:47 +00:00
grpo_cb_paged_10.json
Phase 2 vibe (10-step): vllm vs unsloth_fi_false vs cb_paged + Phase 3 fixes
2026-04-20 14:23:51 +00:00
grpo_cb_paged_10.summary.json
Phase 2 vibe (10-step): vllm vs unsloth_fi_false vs cb_paged + Phase 3 fixes
2026-04-20 14:23:51 +00:00
grpo_cb_paged_30.json
Phase 2: 30-step equivalence + pairwise diffs vs vLLM
2026-04-20 15:03:32 +00:00
grpo_cb_paged_30.summary.json
Phase 2: 30-step equivalence + pairwise diffs vs vLLM
2026-04-20 15:03:32 +00:00
grpo_fi_false_30.json
Phase 2: 30-step equivalence + pairwise diffs vs vLLM
2026-04-20 15:03:32 +00:00
grpo_fi_false_30.summary.json
Phase 2: 30-step equivalence + pairwise diffs vs vLLM
2026-04-20 15:03:32 +00:00
grpo_unsloth_fi_false_10.json
Phase 2 vibe (10-step): vllm vs unsloth_fi_false vs cb_paged + Phase 3 fixes
2026-04-20 14:23:51 +00:00
grpo_unsloth_fi_false_10.summary.json
Phase 2 vibe (10-step): vllm vs unsloth_fi_false vs cb_paged + Phase 3 fixes
2026-04-20 14:23:51 +00:00
grpo_vllm_10.json
Phase 2 vibe (10-step): vllm vs unsloth_fi_false vs cb_paged + Phase 3 fixes
2026-04-20 14:23:51 +00:00
grpo_vllm_10.summary.json
Phase 2 vibe (10-step): vllm vs unsloth_fi_false vs cb_paged + Phase 3 fixes
2026-04-20 14:23:51 +00:00
grpo_vllm_30.json
Phase 2: 30-step equivalence + pairwise diffs vs vLLM
2026-04-20 15:03:32 +00:00
grpo_vllm_30.summary.json
Phase 2: 30-step equivalence + pairwise diffs vs vLLM
2026-04-20 15:03:32 +00:00
lora_cb_paged_fa4_gen.json
Phase 1+3: LoRA rollout benchmarks + CB sync driver + Phase 4 scaffold
2026-04-20 14:01:16 +00:00
lora_cb_sdpa_paged_gen.json
Phase 1+3: LoRA rollout benchmarks + CB sync driver + Phase 4 scaffold
2026-04-20 14:01:16 +00:00
lora_unsloth_fi_false_gen.json
Phase 1+3: LoRA rollout benchmarks + CB sync driver + Phase 4 scaffold
2026-04-20 14:01:16 +00:00
lora_vllm_gen.json
Phase 1+3: LoRA rollout benchmarks + CB sync driver + Phase 4 scaffold
2026-04-20 14:01:16 +00:00
notebook_ref_10.json
Add Phase 0+1 GRPO backend comparison scaffolding
2026-04-20 13:54:06 +00:00
vllm_8.json
FlexKernelOptions sweep: flex reaches 72% of vLLM at batch 64 + LoRA
2026-04-20 23:30:05 +00:00
vllm_16.json
FlexKernelOptions sweep: flex reaches 72% of vLLM at batch 64 + LoRA
2026-04-20 23:30:05 +00:00
vllm_32.json
FlexKernelOptions sweep: flex reaches 72% of vLLM at batch 64 + LoRA
2026-04-20 23:30:05 +00:00
vllm_64.json
FlexKernelOptions sweep: flex reaches 72% of vLLM at batch 64 + LoRA
2026-04-20 23:30:05 +00:00
vllm_64x512_lora.json
Flex+CUDA-graph closes gap to vLLM across batch sizes
2026-04-20 16:09:47 +00:00
vllm_128x512.json
Flex+CUDA-graph closes gap to vLLM across batch sizes
2026-04-20 16:09:47 +00:00
vllm_256x512.json
Batch-size sweep: flex@256 vs vLLM@256 + max-autotune check
2026-04-20 16:17:19 +00:00