unsloth/studio/backend/tests
Daniel Han b48d68f8bf Fix Mistral seed mapping, raise default OAI-compat stop cap, thread sampling through GGUF direct path
Mistral chat completions uses random_seed not seed; map the field via a new seed_field on the provider registry so the new seed control actually works on Mistral. Default for other providers stays seed.

DeepSeek and Mistral both accept up to 16 stop sequences but the default OAI-compat branch was hard-capping at 4 (the OpenAI Chat limit). Studio routes the openai provider through /v1/responses not /v1/chat/completions so the 4-cap only applies if we explicitly added an openai entry. Raise the default to 16 and let per-provider stop_max overrides tighten if needed.

The local GGUF direct chat path (gguf_generate / gguf_generate_with_tools) bypassed _build_openai_passthrough_body and therefore dropped frequency_penalty, seed, stop, and parallel_tool_calls on the floor for users on the default no-tools and with-tools paths. Thread the new fields through LlamaCppBackend.generate_chat_completion and generate_chat_completion_with_tools and the two callsites that invoke them.

Also tighten comments to drop review-process narration that crept in and to remove the em dashes I had introduced in this PR's earlier commits.

Tests pin the Mistral random_seed rename, the DeepSeek 16-cap, and confirm the openai-compat default cap is 16.
2026-05-24 14:32:36 +00:00
..
__init__.py
conftest.py
test_anthropic_cache_ttl.py Studio: support Anthropic 1h cache TTL via prompt_cache_ttl (#5685) 2026-05-22 06:03:32 -07:00
test_anthropic_code_execution.py Studio: per-model Anthropic server-side tool versions (#5679) 2026-05-22 06:03:27 -07:00
test_anthropic_compaction.py Studio: wire Anthropic server-side context compaction (#5686) 2026-05-22 06:19:09 -07:00
test_anthropic_messages.py Studio: Claude Code Anthropic API tool compatibility (#5390) 2026-05-21 16:45:05 +04:00
test_anthropic_thinking_translation.py studio: API external provider support for chat (OpenAI, Mistral, Gemini, Cohere, Anthropic, OpenRouter, DeepSeek, custom providers) (#4706) 2026-05-14 16:13:59 +04:00
test_anthropic_tool_versions.py Studio: per-model Anthropic server-side tool versions (#5679) 2026-05-22 06:03:27 -07:00
test_anthropic_web_fetch.py Studio: wire Anthropic server-side context compaction (#5686) 2026-05-22 06:19:09 -07:00
test_browse_folders_route.py
test_cache_case_resolution.py
test_cached_gguf_routes.py
test_chat_history_routes.py Studio: persist chat history in backend storage (#5272) 2026-05-22 06:18:05 -07:00
test_chat_history_storage.py Studio: persist chat history in backend storage (#5272) 2026-05-22 06:18:05 -07:00
test_cleanup_cancelled_checkpoints.py studio: scope cancel-cleanup to in-flight tmp dirs; walk back tool_call_id (#5488) 2026-05-18 00:01:48 -07:00
test_data_recipe_github_progress.py
test_data_recipe_seed.py
test_desktop_auth.py Studio: persist chat history in backend storage (#5272) 2026-05-22 06:18:05 -07:00
test_detect_mmproj_file.py fix(studio/mmproj): block cross-family projectors in flat local GGUF dirs (#5347) (#5350) 2026-05-14 20:31:20 -07:00
test_export_log_cursor.py
test_external_provider_usage_chunk.py Studio: surface prompt-cache token counts in /v1/chat/completions usage chunk (#5670) 2026-05-22 06:02:52 -07:00
test_gguf_metadata.py fix(studio/mmproj): block cross-family projectors in flat local GGUF dirs (#5347) (#5350) 2026-05-14 20:31:20 -07:00
test_gguf_reload_inheritance.py studio: add --spec-draft-n-max toggle for MTP speculative decoding (#5582) 2026-05-19 06:17:04 -07:00
test_gpu_selection.py Update VRAM estimator to cater to broader model configs (#5175) 2026-05-05 04:12:36 -07:00
test_gpu_selection_sandbox.py
test_host_defaults.py Default Studio host to 127.0.0.1 and prompt before auto-start (#5267) 2026-05-04 13:03:16 +04:00
test_inference_model_validation.py studio: scope cancel-cleanup to in-flight tmp dirs; walk back tool_call_id (#5488) 2026-05-18 00:01:48 -07:00
test_kv_cache_estimation.py studio: reserve VRAM headroom for the MTP draft cache in auto-fit (#5585) 2026-05-19 06:19:02 -07:00
test_llama_cpp_cache_aware_disk_check.py
test_llama_cpp_context_fit.py Studio: pin GPU at 95% headroom and warn on silent CPU fallback (#5323) 2026-05-13 04:48:15 -07:00
test_llama_cpp_freshness.py Studio: warn when llama.cpp prebuilt is at least 3 days behind (#5529) 2026-05-18 00:21:50 -07:00
test_llama_cpp_load_progress.py
test_llama_cpp_load_progress_live.py
test_llama_cpp_load_progress_matrix.py
test_llama_cpp_max_context_threshold.py fix KVCache estimates for gemma4 style sliding window models (#5225) 2026-05-05 04:06:46 -07:00
test_llama_cpp_mtp_detection.py studio: add --spec-draft-n-max toggle for MTP speculative decoding (#5582) 2026-05-19 06:17:04 -07:00
test_llama_cpp_no_context_shift.py
test_llama_cpp_wait_for_health.py tests/studio: lock in Windows GPU detection fix (#5106) with a synthetic CI test (#5376) 2026-05-18 00:06:01 -07:00
test_llama_cpp_wait_for_vram_settle.py studio: settle GPU VRAM after killing llama-server before the next reload (#5693) 2026-05-22 05:50:39 -07:00
test_llama_cpp_windows_nvidia_path.py Studio: add torch's pip nvidia DLL dirs to PATH on Windows (#5324) 2026-05-11 05:42:09 -07:00
test_llama_server_args.py studio: emit one comma-chained --spec-type for CPU/Mac MTP path (#5575) 2026-05-19 03:16:05 -07:00
test_log_filter_no_truncation.py Studio: stop truncating long log lines as suspected base64 (#5335) 2026-05-08 13:07:18 +04:00
test_login_rate_limit.py studio: proxy-aware login rate-limit; allow google favicons in CSP (#5489) 2026-05-18 00:02:15 -07:00
test_middleware.py studio: proxy-aware login rate-limit; allow google favicons in CSP (#5489) 2026-05-18 00:02:15 -07:00
test_mlx_inference_backend.py Studio: tools, thinking blocks, code execution and web search for safetensors (#5520) 2026-05-19 06:30:17 -07:00
test_mlx_training_worker_config.py studio: skip flash-attn install on Blackwell GPUs (sm_100+) (#5420) 2026-05-14 18:13:50 +04:00
test_models_get_model_config_case_resolution.py
test_multimodal_document.py Studio: PDF / document attachments for Anthropic + OpenAI (#5689) 2026-05-22 06:22:57 -07:00
test_native_context_length.py
test_offline_gguf_cache_fallback.py studio: load cached GGUF models when fully offline (#5505) 2026-05-17 21:25:39 -07:00
test_offline_inference_parent.py studio: extend offline DNS auto-detect to inference parent + training (#5512) 2026-05-18 00:31:33 -07:00
test_openai_code_execution.py fix(studio): handle expired OpenAI shell-tool containers without surfacing error in chat (#5547) 2026-05-18 05:47:57 -07:00
test_openai_compaction.py Studio: wire OpenAI Responses server-side context compaction (#5687) 2026-05-22 06:20:45 -07:00
test_openai_container_crud.py tests/openai: patch httpx.AsyncClient ctor so delete tests hit mock (#5469) 2026-05-15 15:53:54 -07:00
test_openai_image_generation.py Studio: wire OpenAI image_generation tool (#5688) 2026-05-22 06:03:38 -07:00
test_openai_responses_translation.py Studio: o3 reasoning summary payload (#5426) 2026-05-15 17:13:28 +04:00
test_openai_tool_passthrough.py Studio: assert promoted fields on attribute path in test_extra_fields_accepted 2026-05-23 15:33:14 +00:00
test_pricing.py Studio: per-session cost calculator + /api/providers/pricing endpoint (#5690) 2026-05-22 06:03:43 -07:00
test_providers_api.py studio: API external provider support for chat (OpenAI, Mistral, Gemini, Cohere, Anthropic, OpenRouter, DeepSeek, custom providers) (#4706) 2026-05-14 16:13:59 +04:00
test_pytorch_mirror.py
test_recommended_folders_permission.py Fix /recommended-folders 500 on unreadable model directories (Python 3.12+) (#5523) 2026-05-18 00:16:14 +04:00
test_responses_api.py
test_responses_tool_passthrough.py
test_safetensors_capability_advertise.py Revert "studio: tool calling for Llama-3, Mistral, Gemma 4 on safetensors + MLX (#5615)" (#5619) 2026-05-19 07:26:39 -07:00
test_safetensors_tool_loop.py Revert "studio: tool calling for Llama-3, Mistral, Gemma 4 on safetensors + MLX (#5615)" (#5619) 2026-05-19 07:26:39 -07:00
test_sampling_params_routing.py Fix Mistral seed mapping, raise default OAI-compat stop cap, thread sampling through GGUF direct path 2026-05-24 14:32:36 +00:00
test_sandbox_tools.py studio: tighten sandbox blocklist precision (bash, hf upload, NOFILE) (#5487) 2026-05-18 00:01:17 -07:00
test_studio_api.py
test_studio_train_validation.py studio: security and hardening pass (auth rate-limit, sandbox, path containment, schema validation, headers) (#5375) 2026-05-13 06:12:18 -07:00
test_tool_policy_gates.py unsloth run: add --enable-tools/--disable-tools server-side tool policy (#5277) 2026-05-05 12:45:15 +04:00
test_tool_policy_state.py unsloth run: add --enable-tools/--disable-tools server-side tool policy (#5277) 2026-05-05 12:45:15 +04:00
test_trained_model_scan.py studio: security and hardening pass (auth rate-limit, sandbox, path containment, schema validation, headers) (#5375) 2026-05-13 06:12:18 -07:00
test_training_history_update.py Studio: Dark theme refactor, right sidebar redesign, and chat UI polish (#5150) 2026-05-07 14:33:31 +04:00
test_training_raw_support.py studio: drop unused max_grad_value schema + route plumbing (#5424) 2026-05-14 05:43:58 -07:00
test_training_worker_flash_attn.py studio: install flash-linear-attention and tilelang for Qwen3.5 family (#5434) 2026-05-18 03:49:06 -07:00
test_transformers_version.py
test_utils.py
test_vision_cache.py
test_vram_estimation.py Update VRAM estimator to cater to broader model configs (#5175) 2026-05-05 04:12:36 -07:00
test_windows_gpu_detection_mock.py tests/studio: lock in Windows GPU detection fix (#5106) with a synthetic CI test (#5376) 2026-05-18 00:06:01 -07:00