unsloth/studio/backend/tests
Daniel Han 9a737facaf
Studio: wire Anthropic server-side context compaction (#5686)
* Studio: wire Anthropic server-side context compaction

Anthropic ships server-side context compaction as a beta
(`compact-2026-01-12`). When the rendered prompt crosses the
configured input-token threshold, Anthropic runs an extra LLM pass
that summarises older turns and the request continues against the
compacted prefix. The response carries the original top-level fields
plus a new `context_management` block (with `applied_edits`) and
`usage.iterations[]` accounting per pass.

Per the docs the feature is currently supported on Opus 4.6, Opus 4.7,
Sonnet 4.6, and Mythos preview. The minimum threshold is 50k tokens;
under-50k requests 400.

Changes:

- Add prefix gate + helper `_anthropic_supports_compaction` plus
  constants `_ANTHROPIC_COMPACTION_PREFIXES`, `_ANTHROPIC_COMPACTION_BETA`,
  `_ANTHROPIC_COMPACTION_TYPE`, `_ANTHROPIC_COMPACTION_MIN`.
- Add `compaction_threshold: Optional[int]` to ChatCompletionRequest
  (50k ge bound, 2M le bound). Thread through `routes/inference.py`
  -> `stream_chat_completion` -> `_stream_anthropic`.
- In `_stream_anthropic`, when threshold is set AND the model
  accepts compaction, attach `context_management.edits[{type:
  "compact_20260112", trigger:{type:"input_tokens", value:N}}]` to
  the outbound body. Sub-50k values are clamped up to 50k to keep
  the request well-formed.
- Refactor the anthropic-beta header builder to merge any combination
  of `code-execution-2025-08-25` + `compact-2026-01-12` flags into
  one header value. Unrelated betas added at the registry level still
  pass through.
- Add `test_anthropic_compaction.py` with 16 cases: gate matrix
  (every doc-listed model), correct body shape, threshold clamping,
  beta header merge with code execution, silent no-op on unsupported
  models, omitted-threshold pass-through.

Live verified end-to-end against the real Anthropic API:
`compact_20260112` accepted on Opus 4.7, response carries
`context_management.applied_edits` + `usage.iterations[]` as
documented. (The first WebFetch-summarised version of these docs
suggested `compact_20260120`; the actual API only accepts
`compact_20260112`, matching the beta-header date. Worth pinning
behind a test so a future doc update can't drift back.)

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Address review: drop ge=50_000 clamp + parse usage.iterations[]

Two reviewer follow-ups on the compaction PR:

1. Pydantic ge=50_000 on compaction_threshold was dead code.
   FastAPI rejected sub-50k threshold values with a 422 before the
   `max(int(...), _ANTHROPIC_COMPACTION_MIN)` clamp in
   _stream_anthropic could ever fire. Relaxed the floor to ge=1 so
   the in-helper clamp actually does its job; the schema comment
   now explains why this is intentional. Added a regression test
   that posts a value of 1 and 49_999 through the real request
   schema.

2. Anthropic publishes per-iteration token counts in
   `usage.iterations[]` whenever a fresh compaction has run, and
   the top-level input_tokens / output_tokens cover only the
   `message` iteration -- billing must add the compaction
   iterations on top. Aggregate compaction iteration tokens into
   `last_usage["compaction_input_tokens" / "compaction_output_tokens"]`
   so the cost surface (PR 5690) can read them without re-walking
   the array, and surface both figures in the closing stream
   summary log. Added two tests: one that pins the aggregation on a
   compacted turn and one that pins `None` when no fresh
   iterations land (so re-applied compaction blocks don't double-bill).

Sourcing: https://platform.claude.com/docs/en/build-with-claude/compaction

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Address review: round-trip Anthropic compaction blocks across turns

Codex P1: once context_management is enabled and Anthropic runs
server-side compaction mid-stream, the response carries a
`{type:"compaction", content:"<summary>"}` content block on the
assistant message. The translator only handled text_delta and
input_json_delta on content_block_delta, so the compaction block
was silently dropped. Worse, the request schema's ContentPart
discriminated Union didn't accept `type:"compaction"`, and
_build_external_messages didn't pass it through, so even a
hand-crafted assistant message carrying the block would 422 at
parse time. Net result: Anthropic re-compacted from scratch on
every subsequent turn, wasting input tokens and reasoning budget.

End-to-end backend wiring of the round-trip:

1. SSE translator. _stream_anthropic now tracks a `current_compaction`
   state slot. content_block_start with type=="compaction" seeds it
   (Anthropic may include the summary on the start event AND/OR
   stream it via text_delta events on the same block index --
   handle both). text_delta inside a compaction block routes into
   the compaction buffer instead of the user-visible content
   stream, since the summary is opaque internal state, not
   assistant prose. content_block_stop emits a `compaction_block`
   tool_event carrying the full summary so the chat-adapter can
   persist it. compaction_blocks_seen is surfaced in the closing
   summary log.

2. Pydantic schema. Added CompactionContentPart with Tag("compaction")
   on the ContentPart Union so requests carrying the block parse
   cleanly. Required `content` field with a docstring pointing at
   the Anthropic docs.

3. Message builder. _build_external_messages forwards compaction
   parts on both vision and non-vision paths; the per-provider
   stream helper decides whether to forward to the wire (Anthropic
   does; other providers ignore the part). When a non-vision route
   ends up with a single text part, collapse back to a string
   so providers that don't accept content arrays still get the
   expected shape.

4. _stream_anthropic outbound translator. {type:"compaction"} parts
   on an assistant message land on the wire verbatim. Empty/missing
   `content` is skipped so a malformed stored block can't 400
   Anthropic.

Tests added (5): stream emits compaction_block tool event with the
summary intact; user-visible content stream does NOT carry the
summary text; outbound body forwards compaction parts verbatim on
the next turn; Pydantic schema accepts the part; builder passes
it through on both vision and non-vision provider routes.

Frontend follow-up: the chat-adapter needs to persist the
compaction_block tool_event onto the stored assistant message so
turn N+1 includes it in payload.messages. Pinned in the PR
description.

Sourcing: https://platform.claude.com/docs/en/build-with-claude/compaction

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Address review: gate compaction-part passthrough to Anthropic only

Codex P1: my previous round-trip change preserved {type:"compaction"}
parts on every provider route in _build_external_messages. That
meant a chat history with prior compaction state silently leaked
the Anthropic-specific block to OpenAI/DeepSeek/Mistral/Gemini/
Kimi/OpenRouter on a provider switch, where generic
/chat/completions passthrough hands the unknown content type to
the upstream API and 400s the whole turn.

Added a `provider_type` kwarg to _build_external_messages and
gated the compaction forwarder on `provider_type == "anthropic"`.
Every other value (including the legacy None for callers that
don't pass it yet) strips the part. The Anthropic stream helper
still maps it to a native `compaction` block on the wire.

Threaded provider_type through from _proxy_to_external_provider's
call site.

Tests updated: vision + provider="anthropic" still forwards; six
non-anthropic providers strip the part; missing provider_type
strips defensively; non-vision + anthropic still forwards; non-vision
+ non-anthropic collapses back to a text string.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-22 06:19:09 -07:00
..
__init__.py Final cleanup 2026-03-12 18:28:04 +00:00
conftest.py Studio: Expose openai and anthropic compatible external API end points (#4956) 2026-04-13 21:08:11 +04:00
test_anthropic_cache_ttl.py Studio: support Anthropic 1h cache TTL via prompt_cache_ttl (#5685) 2026-05-22 06:03:32 -07:00
test_anthropic_code_execution.py Studio: per-model Anthropic server-side tool versions (#5679) 2026-05-22 06:03:27 -07:00
test_anthropic_compaction.py Studio: wire Anthropic server-side context compaction (#5686) 2026-05-22 06:19:09 -07:00
test_anthropic_messages.py Studio: Claude Code Anthropic API tool compatibility (#5390) 2026-05-21 16:45:05 +04:00
test_anthropic_thinking_translation.py studio: API external provider support for chat (OpenAI, Mistral, Gemini, Cohere, Anthropic, OpenRouter, DeepSeek, custom providers) (#4706) 2026-05-14 16:13:59 +04:00
test_anthropic_tool_versions.py Studio: per-model Anthropic server-side tool versions (#5679) 2026-05-22 06:03:27 -07:00
test_anthropic_web_fetch.py Studio: wire Anthropic server-side context compaction (#5686) 2026-05-22 06:19:09 -07:00
test_browse_folders_route.py Studio: add folder browser modal for Custom Folders (#5035) 2026-04-15 08:04:33 -07:00
test_cache_case_resolution.py Add tests for cache case resolution (from PR #4822) (#4823) 2026-04-03 13:58:26 -07:00
test_cached_gguf_routes.py Studio: support GGUF variant selection for non-suffixed repos (#5023) 2026-04-15 15:32:01 +04:00
test_chat_history_routes.py Studio: persist chat history in backend storage (#5272) 2026-05-22 06:18:05 -07:00
test_chat_history_storage.py Studio: persist chat history in backend storage (#5272) 2026-05-22 06:18:05 -07:00
test_cleanup_cancelled_checkpoints.py studio: scope cancel-cleanup to in-flight tmp dirs; walk back tool_call_id (#5488) 2026-05-18 00:01:48 -07:00
test_data_recipe_github_progress.py Studio: add github_repo seed reader and GitHub Support Bot recipe (#5169) 2026-04-24 12:02:03 -07:00
test_data_recipe_seed.py fix(seed): disable remote code execution in seed inspect dataset loads (#4275) 2026-03-13 19:37:43 +04:00
test_desktop_auth.py Studio: persist chat history in backend storage (#5272) 2026-05-22 06:18:05 -07:00
test_detect_mmproj_file.py fix(studio/mmproj): block cross-family projectors in flat local GGUF dirs (#5347) (#5350) 2026-05-14 20:31:20 -07:00
test_export_log_cursor.py studio: stream export worker output into the export dialog (#4897) 2026-04-14 08:55:43 -07:00
test_external_provider_usage_chunk.py Studio: surface prompt-cache token counts in /v1/chat/completions usage chunk (#5670) 2026-05-22 06:02:52 -07:00
test_gguf_metadata.py fix(studio/mmproj): block cross-family projectors in flat local GGUF dirs (#5347) (#5350) 2026-05-14 20:31:20 -07:00
test_gguf_reload_inheritance.py studio: add --spec-draft-n-max toggle for MTP speculative decoding (#5582) 2026-05-19 06:17:04 -07:00
test_gpu_selection.py Update VRAM estimator to cater to broader model configs (#5175) 2026-05-05 04:12:36 -07:00
test_gpu_selection_sandbox.py [Studio] multi gpu finetuning/inference via "balanced_low0/sequential" device_map (#4602) 2026-03-30 02:33:15 -07:00
test_host_defaults.py Default Studio host to 127.0.0.1 and prompt before auto-start (#5267) 2026-05-04 13:03:16 +04:00
test_inference_model_validation.py studio: scope cancel-cleanup to in-flight tmp dirs; walk back tool_call_id (#5488) 2026-05-18 00:01:48 -07:00
test_kv_cache_estimation.py studio: reserve VRAM headroom for the MTP draft cache in auto-fit (#5585) 2026-05-19 06:19:02 -07:00
test_llama_cpp_cache_aware_disk_check.py Studio: make GGUF disk-space preflight cache-aware (#5012) 2026-04-14 08:53:37 -07:00
test_llama_cpp_context_fit.py Studio: pin GPU at 95% headroom and warn on silent CPU fallback (#5323) 2026-05-13 04:48:15 -07:00
test_llama_cpp_freshness.py Studio: warn when llama.cpp prebuilt is at least 3 days behind (#5529) 2026-05-18 00:21:50 -07:00
test_llama_cpp_load_progress.py Studio: live model-load progress + rate/ETA on download and load (#5017) 2026-04-14 09:46:22 -07:00
test_llama_cpp_load_progress_live.py Studio: split model-load progress label across two rows (#5020) 2026-04-14 10:58:16 -07:00
test_llama_cpp_load_progress_matrix.py Studio: split model-load progress label across two rows (#5020) 2026-04-14 10:58:16 -07:00
test_llama_cpp_max_context_threshold.py fix KVCache estimates for gemma4 style sliding window models (#5225) 2026-05-05 04:06:46 -07:00
test_llama_cpp_mtp_detection.py studio: add --spec-draft-n-max toggle for MTP speculative decoding (#5582) 2026-05-19 06:17:04 -07:00
test_llama_cpp_no_context_shift.py Studio: hard-stop at n_ctx with a 'Context limit reached' toast (#5021) 2026-04-14 10:58:20 -07:00
test_llama_cpp_wait_for_health.py tests/studio: lock in Windows GPU detection fix (#5106) with a synthetic CI test (#5376) 2026-05-18 00:06:01 -07:00
test_llama_cpp_wait_for_vram_settle.py studio: settle GPU VRAM after killing llama-server before the next reload (#5693) 2026-05-22 05:50:39 -07:00
test_llama_cpp_windows_nvidia_path.py Studio: add torch's pip nvidia DLL dirs to PATH on Windows (#5324) 2026-05-11 05:42:09 -07:00
test_llama_server_args.py studio: emit one comma-chained --spec-type for CPU/Mac MTP path (#5575) 2026-05-19 03:16:05 -07:00
test_log_filter_no_truncation.py Studio: stop truncating long log lines as suspected base64 (#5335) 2026-05-08 13:07:18 +04:00
test_login_rate_limit.py studio: proxy-aware login rate-limit; allow google favicons in CSP (#5489) 2026-05-18 00:02:15 -07:00
test_middleware.py studio: proxy-aware login rate-limit; allow google favicons in CSP (#5489) 2026-05-18 00:02:15 -07:00
test_mlx_inference_backend.py Studio: tools, thinking blocks, code execution and web search for safetensors (#5520) 2026-05-19 06:30:17 -07:00
test_mlx_training_worker_config.py studio: skip flash-attn install on Blackwell GPUs (sm_100+) (#5420) 2026-05-14 18:13:50 +04:00
test_models_get_model_config_case_resolution.py Add tests for cache case resolution (from PR #4822) (#4823) 2026-04-03 13:58:26 -07:00
test_native_context_length.py Studio: Fix chat template disappearing after browser refresh (#5209) 2026-05-01 08:19:09 -07:00
test_offline_gguf_cache_fallback.py studio: load cached GGUF models when fully offline (#5505) 2026-05-17 21:25:39 -07:00
test_offline_inference_parent.py studio: extend offline DNS auto-detect to inference parent + training (#5512) 2026-05-18 00:31:33 -07:00
test_openai_code_execution.py fix(studio): handle expired OpenAI shell-tool containers without surfacing error in chat (#5547) 2026-05-18 05:47:57 -07:00
test_openai_container_crud.py tests/openai: patch httpx.AsyncClient ctor so delete tests hit mock (#5469) 2026-05-15 15:53:54 -07:00
test_openai_image_generation.py Studio: wire OpenAI image_generation tool (#5688) 2026-05-22 06:03:38 -07:00
test_openai_responses_translation.py Studio: o3 reasoning summary payload (#5426) 2026-05-15 17:13:28 +04:00
test_openai_tool_passthrough.py Fix GGUF multi-image chat handling (#5508) 2026-05-19 04:36:20 -07:00
test_pricing.py Studio: per-session cost calculator + /api/providers/pricing endpoint (#5690) 2026-05-22 06:03:43 -07:00
test_providers_api.py studio: API external provider support for chat (OpenAI, Mistral, Gemini, Cohere, Anthropic, OpenRouter, DeepSeek, custom providers) (#4706) 2026-05-14 16:13:59 +04:00
test_pytorch_mirror.py Add configurable PyTorch mirror via UNSLOTH_PYTORCH_MIRROR env var (#5024) 2026-04-15 11:39:11 +04:00
test_recommended_folders_permission.py Fix /recommended-folders 500 on unreadable model directories (Python 3.12+) (#5523) 2026-05-18 00:16:14 +04:00
test_responses_api.py Studio: Expose openai and anthropic compatible external API end points (#4956) 2026-04-13 21:08:11 +04:00
test_responses_tool_passthrough.py Studio: forward standard OpenAI tools / tool_choice on /v1/responses (Codex compat) (#5122) 2026-04-21 13:17:20 +04:00
test_safetensors_capability_advertise.py Revert "studio: tool calling for Llama-3, Mistral, Gemma 4 on safetensors + MLX (#5615)" (#5619) 2026-05-19 07:26:39 -07:00
test_safetensors_tool_loop.py Revert "studio: tool calling for Llama-3, Mistral, Gemma 4 on safetensors + MLX (#5615)" (#5619) 2026-05-19 07:26:39 -07:00
test_sandbox_tools.py studio: tighten sandbox blocklist precision (bash, hf upload, NOFILE) (#5487) 2026-05-18 00:01:17 -07:00
test_studio_api.py Studio: forward standard OpenAI tools / tool_choice to llama-server (#5099) 2026-04-18 12:53:23 +04:00
test_studio_train_validation.py studio: security and hardening pass (auth rate-limit, sandbox, path containment, schema validation, headers) (#5375) 2026-05-13 06:12:18 -07:00
test_tool_policy_gates.py unsloth run: add --enable-tools/--disable-tools server-side tool policy (#5277) 2026-05-05 12:45:15 +04:00
test_tool_policy_state.py unsloth run: add --enable-tools/--disable-tools server-side tool policy (#5277) 2026-05-05 12:45:15 +04:00
test_trained_model_scan.py studio: security and hardening pass (auth rate-limit, sandbox, path containment, schema validation, headers) (#5375) 2026-05-13 06:12:18 -07:00
test_training_history_update.py Studio: Dark theme refactor, right sidebar redesign, and chat UI polish (#5150) 2026-05-07 14:33:31 +04:00
test_training_raw_support.py studio: drop unused max_grad_value schema + route plumbing (#5424) 2026-05-14 05:43:58 -07:00
test_training_worker_flash_attn.py studio: install flash-linear-attention and tilelang for Qwen3.5 family (#5434) 2026-05-18 03:49:06 -07:00
test_transformers_version.py split venv_t5 into tiered 5.3.0/5.5.0 and fix trust_remote_code (#4878) 2026-04-07 20:05:01 +04:00
test_utils.py Add AMD ROCm/HIP support across installer and hardware detection (#4720) 2026-04-10 01:56:12 -07:00
test_vision_cache.py Studio: split vision-cache exception test to match transient vs permanent (#5145) 2026-04-23 00:22:40 -07:00
test_vram_estimation.py Update VRAM estimator to cater to broader model configs (#5175) 2026-05-05 04:12:36 -07:00
test_windows_gpu_detection_mock.py tests/studio: lock in Windows GPU detection fix (#5106) with a synthetic CI test (#5376) 2026-05-18 00:06:01 -07:00