Added test_mla_defaults_n_kv_to_1_when_heads_absent to verify MLA path uses n_kv=1 (not n_heads) when head_count_kv is absent. |
||
|---|---|---|
| .. | ||
| __init__.py | ||
| conftest.py | ||
| test_data_recipe_seed.py | ||
| test_gpu_selection.py | ||
| test_gpu_selection_sandbox.py | ||
| test_kv_cache_estimation.py | ||
| test_transformers_version.py | ||
| test_utils.py | ||
| test_vram_estimation.py | ||