unsloth/studio/backend/tests/test_pricing.py
Daniel Han 7d1b68079e
Studio: Anthropic fast_mode toggle and streaming refusal handling (#5715)
* Studio: add Anthropic fast_mode toggle + surface streaming refusals

Fast mode (beta `fast-mode-2026-02-01`) lets Claude Opus 4.6 and 4.7
generate output tokens up to 2.5x faster at 6x standard Opus
pricing. The toggle lives in Configuration → Provider when the
selected Anthropic model is Opus 4.6 or 4.7 and is otherwise
hidden. Backend gates the same prefixes a second time so a stale
frontend cannot make Anthropic 400 the request, and the
`fast-mode-2026-02-01` beta header is merged onto whatever other
betas the request already needed (code-execution, compaction).

Streaming refusals (`message_delta.delta.stop_reason="refusal"` on
Claude 4 models) now surface a short user-facing notice in the
assistant message before the translated OpenAI chunk emits the
existing `finish_reason="content_filter"`. Previously the chat
bubble truncated silently because the SSE stopped mid-stream with
no visible explanation. Per the upstream docs the conversation
must be reset before continuing, so the notice tells the user
exactly that.

Reference:
- https://platform.claude.com/docs/en/build-with-claude/fast-mode
- https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/handle-streaming-refusals

Tests:
- studio/backend/tests/test_anthropic_fast_mode_and_refusal.py (8 cases
  pinning fast_mode pass-through on 4.6/4.7, silent drop on Sonnet /
  Haiku / older Opus / None / False, and the refusal notice + finish
  reason on a synthetic refusal stream).

* Studio: drop refused Anthropic turns from the next request

Anthropic's streaming-refusal guidance says the refused assistant
turn must be removed or updated before the next call -- otherwise
the safety classifier keeps refusing. The PR only added a
user-visible notice; the partial assistant output (plus the notice
itself) still rode the next request via toOpenAIMessage.

Tag the refusal turn with an HTML-comment sentinel emitted alongside
the notice. The chat-adapter checks for that sentinel in
toOpenAIMessage and returns null, so the refused turn is excluded
from outboundMessages. The notice still renders in the transcript
(HTML comments don't display), so users keep the explanation.

* Studio: filter None finish_reason entries in test helper

test_refusal_maps_to_content_filter expects only ['content_filter']
in the finish_reasons list, but the post-PR refusal path emits a
user-visible content notice chunk first. Every _content_chunk
carries 'finish_reason: None' by construction; the helper was
appending those, so the assertion saw [None, 'content_filter']
instead of ['content_filter'].

None is not a finish reason -- it's just mid-stream delta noise.
Skip None values in _finish_reasons so the helper reflects what
the test names actually claim to check. Same fix applies cleanly
to the other helper usages (pause_turn test expects [] and the
sibling stop test expects ['stop'], both unaffected).

* Studio: cover Anthropic fast-mode edge cases

Adds 19 cases on top of the 9 in test_anthropic_fast_mode_and_refusal.
The base file pins the happy path; this file fills in the cliffs:

* Dated-snapshot prefix matching: claude-opus-4-7-2026-02-01 and
  claude-opus-4-6-2026-02-01 still gate fast_mode through, while
  claude-opus-4-5-2025-08-01 and claude-sonnet-4-6-2026-02-01 do not.
* Strict opt-in: a future claude-opus-4-8 or claude-opus-5 does NOT
  auto-enable fast_mode -- the prefix tuple must be bumped explicitly
  when a new family is whitelisted upstream.
* Beta-header merge: fast_mode coexists with code-execution-2025-08-25
  and compact-2026-01-12 in one comma-separated anthropic-beta header
  with no duplicates and no truncation. Pins the value to the exact
  fast-mode-2026-02-01 docs token so a typo would fail CI.
* Non-destruction: fast_mode=None produces byte-identical outbound
  body and headers to the version that omits the argument entirely.
  Same for fast_mode=False. Guarantees the upgrade path is
  non-breaking on existing Anthropic streams.
* Refusal stream ordering: the user-visible notice precedes the
  finish_reason chunk so a streaming UI paints text before flipping
  to content_filter. Refusal sentinel emitted exactly once. Notice
  rides a normal content delta chunk with finish_reason still null.
  Partial assistant deltas survive before the notice.
* Provider-side refusal coverage: a refusal on Sonnet (not just Opus)
  still emits the notice + sentinel + content_filter mapping, since
  refusal handling is not gated on fast-mode capability.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Persist fastMode, drop refused user message on retry

Two follow-ups on #5715:

1) sanitizeInferenceParams stripped fastMode. fastMode is in
   PERSISTED_INFERENCE_PARAM_KEYS but the storage sanitizer only kept
   numeric fields plus systemPrompt and trustRemoteCode, so the new
   toggle was silently dropped on reload and on the
   /api/chat/settings round-trip. Save it the same way trustRemoteCode
   is saved.

2) Refusal recovery now also drops the triggering user turn.
   Returning null from toOpenAIMessage on the assistant side left the
   user prompt that caused the refusal in the outbound history, so
   the very next request would re-trigger the same classifier.
   Anthropic's refusal-handling guidance is explicit on this: remove
   the refused turn AND the user message that triggered it before
   the next call. Implemented via a pre-pass that pops the trailing
   user message when an assistant carries the refusal sentinel.

Typecheck clean.

* Studio: out-of-band refusal signal + fast-mode prefix/usage/pricing fixes

The text sentinel for the Anthropic refusal drop signal was spoofable:
any assistant message containing the literal
<!--studio:anthropic-refusal--> would prune the prior user + assistant
pair on the next request. Move the signal onto a separate _toolEvent
chunk that the chat adapter latches into
assistant.metadata.custom.anthropicRefusal; assistant text can no
longer control the pruner.

Tighten the fast-mode model gate (backend + frontend) to require a "-"
family boundary so claude-opus-4-70 / claude-opus-4-7b style IDs do
not get speed: "fast" on a naive startswith match.

Use survivingMessages for the image / audio attachment scan so a
refused user turn does not gate or mis-attribute the next non-refused
turn.

Propagate Anthropic usage.speed onto the OpenAI-style usage chunk and
apply the documented 6x fast-mode multiplier in the cost calculator
(stacks with prompt-cache multipliers per the docs); expose the new
multiplier on the pricing snapshot for the UI tooltip.

Tests cover the tool-event chunk shape, the prefix-collision rejects,
usage.speed propagation, the 6x pricing math, and that the visible
refusal text carries no embedded sentinel.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Shorten fast-mode and refusal comments for PR #5715

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-05-25 23:37:12 -07:00

487 lines
17 KiB
Python

# SPDX-License-Identifier: AGPL-3.0-only
# Copyright 2026-present the Unsloth AI Inc. team. All rights reserved. See /studio/LICENSE.AGPL-3.0
"""Unit tests for the per-session cost calculator.
Pricing inputs are baked into ``core/inference/pricing.py``; this
test verifies the math (with multipliers from the prompt-caching
docs) and that unknown models / empty usage degrade gracefully.
"""
import math
from core.inference.pricing import (
ANTHROPIC_CACHE_5M_WRITE_MULT,
ANTHROPIC_CACHE_1H_WRITE_MULT,
ANTHROPIC_CACHE_READ_MULT,
ANTHROPIC_FAST_MODE_MULT,
ANTHROPIC_PRICING,
OPENAI_CACHE_READ_MULT,
OPENAI_CONTAINER_USD_PER_HOUR,
OPENAI_PRICING,
OPENAI_WEB_SEARCH_USD_PER_1K,
calculate_cost,
pricing_snapshot,
)
def _isclose(a, b, tol = 1e-6):
return math.isclose(a, b, rel_tol = tol, abs_tol = tol)
# ── unknown model -> priced=False, totals zero, tokens still report ──
def test_unknown_model_priced_false():
out = calculate_cost(
"anthropic",
"made-up-model-9000",
{"input_tokens": 100, "output_tokens": 50},
)
assert out["priced"] is False
assert out["total_usd"] == 0.0
assert out["billable_input_tokens"] == 100
assert out["billable_output_tokens"] == 50
# ── Anthropic base math (Opus 4.7: 5/25 per MTok) ────────────────────
def test_anthropic_opus_4_7_input_and_output_math():
out = calculate_cost(
"anthropic",
"claude-opus-4-7",
{"input_tokens": 1_000_000, "output_tokens": 1_000_000},
)
assert _isclose(out["input_usd"], 5.0)
assert _isclose(out["output_usd"], 25.0)
assert _isclose(out["total_usd"], 30.0)
# ── Anthropic fast-mode 6x multiplier (Opus 4.6 / 4.7 only) ─────────
def test_anthropic_fast_mode_charges_6x_standard_opus():
"""6x on input + output when ``usage.speed == "fast"``.
https://platform.claude.com/docs/en/build-with-claude/fast-mode"""
out = calculate_cost(
"anthropic",
"claude-opus-4-7",
{
"input_tokens": 1_000_000,
"output_tokens": 1_000_000,
"speed": "fast",
},
)
assert _isclose(out["input_usd"], 5.0 * ANTHROPIC_FAST_MODE_MULT)
assert _isclose(out["output_usd"], 25.0 * ANTHROPIC_FAST_MODE_MULT)
assert _isclose(out["total_usd"], 30.0 * ANTHROPIC_FAST_MODE_MULT)
assert "(fast)" in out["model_priced"], out["model_priced"]
def test_anthropic_fast_mode_does_not_affect_standard_speed():
"""``speed: "standard"`` (or missing) keeps the base rates."""
out_standard = calculate_cost(
"anthropic",
"claude-opus-4-7",
{
"input_tokens": 1_000_000,
"output_tokens": 1_000_000,
"speed": "standard",
},
)
out_missing = calculate_cost(
"anthropic",
"claude-opus-4-7",
{"input_tokens": 1_000_000, "output_tokens": 1_000_000},
)
assert _isclose(out_standard["total_usd"], out_missing["total_usd"])
assert _isclose(out_standard["total_usd"], 30.0)
def test_anthropic_fast_mode_stacks_with_cache_read_multiplier():
"""Cache multipliers apply on top of fast-mode (per docs)."""
base = ANTHROPIC_PRICING["claude-opus-4-7"]["input_per_mtok"]
out = calculate_cost(
"anthropic",
"claude-opus-4-7",
{
"input_tokens": 0,
"output_tokens": 0,
"cache_read_input_tokens": 1_000_000,
"speed": "fast",
},
)
expected = base * ANTHROPIC_FAST_MODE_MULT * ANTHROPIC_CACHE_READ_MULT
assert _isclose(out["cache_read_usd"], expected)
# ── Anthropic cache write 5m + read multipliers ──────────────────────
def test_anthropic_cache_5m_and_read_use_correct_multipliers():
base = ANTHROPIC_PRICING["claude-opus-4-7"]["input_per_mtok"]
out = calculate_cost(
"anthropic",
"claude-opus-4-7",
{
"input_tokens": 0,
"output_tokens": 0,
"cache_creation_input_tokens": 1_000_000,
"cache_read_input_tokens": 1_000_000,
"cache_creation": {
"ephemeral_5m_input_tokens": 1_000_000,
"ephemeral_1h_input_tokens": 0,
},
},
)
assert _isclose(out["cache_write_usd"], base * ANTHROPIC_CACHE_5M_WRITE_MULT)
assert _isclose(out["cache_read_usd"], base * ANTHROPIC_CACHE_READ_MULT)
# billable_input_tokens = input + cache_create + cache_read
assert out["billable_input_tokens"] == 2_000_000
def test_anthropic_cache_1h_write_uses_2x_multiplier():
base = ANTHROPIC_PRICING["claude-opus-4-7"]["input_per_mtok"]
out = calculate_cost(
"anthropic",
"claude-opus-4-7",
{
"input_tokens": 0,
"output_tokens": 0,
"cache_creation_input_tokens": 1_000_000,
"cache_read_input_tokens": 0,
"cache_creation": {
"ephemeral_5m_input_tokens": 0,
"ephemeral_1h_input_tokens": 1_000_000,
},
},
)
assert _isclose(out["cache_write_usd"], base * ANTHROPIC_CACHE_1H_WRITE_MULT)
def test_anthropic_cache_5m_default_when_no_breakdown():
# When the docs/response doesn't surface the 5m/1h split, treat
# the full cache_creation bucket as 5m (the upstream default pool).
base = ANTHROPIC_PRICING["claude-opus-4-7"]["input_per_mtok"]
out = calculate_cost(
"anthropic",
"claude-opus-4-7",
{
"input_tokens": 0,
"output_tokens": 0,
"cache_creation_input_tokens": 500_000,
},
)
expected = 0.5 * base * ANTHROPIC_CACHE_5M_WRITE_MULT
assert _isclose(out["cache_write_usd"], expected)
# ── Anthropic server-tool surcharges ────────────────────────────────
def test_anthropic_web_search_charged_per_thousand():
out = calculate_cost(
"anthropic",
"claude-opus-4-7",
{
"input_tokens": 0,
"output_tokens": 0,
"server_tool_use": {"web_search_requests": 250},
},
)
assert _isclose(out["server_tools_usd"], 2.5) # $10/1000 * 250
def test_anthropic_code_exec_charged_per_hour():
out = calculate_cost(
"anthropic",
"claude-opus-4-7",
{
"input_tokens": 0,
"output_tokens": 0,
"server_tool_use": {"code_execution_hours": 2.0},
},
)
assert _isclose(out["server_tools_usd"], 0.10) # $0.05/hr * 2
def test_anthropic_dated_id_falls_back_to_canonical_prefix():
# Hypothetical dated snapshot of claude-opus-4-7 should still
# inherit the canonical-id pricing via the prefix-match fallback.
out = calculate_cost(
"anthropic",
"claude-opus-4-7-20260712",
{"input_tokens": 1_000_000, "output_tokens": 0},
)
assert out["priced"] is True
assert _isclose(out["input_usd"], 5.0)
# ── OpenAI base math (gpt-5.5: 5/30 per MTok) ────────────────────────
def test_openai_gpt55_input_output_math():
# Sub-272k input keeps us in the short-context tier ($5/$30).
# The dedicated long-context tests below exercise the crossover.
out = calculate_cost(
"openai",
"gpt-5.5",
{"input_tokens": 200_000, "output_tokens": 50_000},
)
assert _isclose(out["input_usd"], 200_000 / 1_000_000.0 * 5.0)
assert _isclose(out["output_usd"], 50_000 / 1_000_000.0 * 30.0)
assert _isclose(out["total_usd"], 1.0 + 1.5)
def test_openai_cache_read_subtracted_from_input_at_discount():
# OpenAI folds cached tokens into input_tokens, unlike Anthropic.
# The calculator must subtract cached_tokens from the "full price"
# bucket and re-bill them at 0.1x. Use a sub-272k total so the
# short-context tier applies (long-context crossover is exercised
# in its own test below).
base = OPENAI_PRICING["gpt-5.5"]["input_per_mtok"]
out = calculate_cost(
"openai",
"gpt-5.5",
{
"input_tokens": 100_000,
"output_tokens": 0,
"input_tokens_details": {"cached_tokens": 80_000},
},
)
# 20k charged at full price, 80k charged at 0.1x
assert _isclose(out["input_usd"], 20_000 / 1_000_000.0 * base)
assert _isclose(
out["cache_read_usd"], 80_000 / 1_000_000.0 * base * OPENAI_CACHE_READ_MULT
)
def test_openai_billable_input_tokens_does_not_double_count_cache_read():
# OpenAI's input_tokens already includes cached_tokens, so the
# billable counter must NOT add cache_read on top -- otherwise the
# tooltip says 180k input when the bill is for 100k.
out = calculate_cost(
"openai",
"gpt-5.5",
{
"input_tokens": 100_000,
"output_tokens": 0,
"input_tokens_details": {"cached_tokens": 80_000},
},
)
assert out["billable_input_tokens"] == 100_000
def test_openai_dated_snapshot_inherits_canonical_pricing():
# Sub-272k stays in the short-context tier; the prefix-match
# fallback is what proves the dated snapshot inherits gpt-5.5
# pricing.
out = calculate_cost(
"openai",
"gpt-5.5-2026-04-23",
{"input_tokens": 200_000, "output_tokens": 0},
)
assert out["priced"] is True
assert _isclose(out["input_usd"], 200_000 / 1_000_000.0 * 5.0)
def test_openai_gpt54_family_uses_verified_prices():
# Spot-check the lower-tier rows that previously underbilled.
# gpt-5.4 has a long-context tier so the input has to stay
# below 272k; the mini/nano/codex rows have no crossover so
# 1M tokens is fine.
cases = {
# (input_tokens, expected_input_usd, expected_output_usd)
"gpt-5.4": (200_000, 200_000 / 1_000_000.0 * 2.5, 200_000 / 1_000_000.0 * 15.0),
"gpt-5.4-mini": (1_000_000, 0.75, 4.5),
"gpt-5.4-nano": (1_000_000, 0.20, 1.25),
"gpt-5.3-codex": (1_000_000, 1.75, 14.0),
}
for model, (in_tokens, exp_in, exp_out) in cases.items():
out = calculate_cost(
"openai",
model,
{"input_tokens": in_tokens, "output_tokens": in_tokens},
)
assert out["priced"] is True, model
assert _isclose(out["input_usd"], exp_in), model
assert _isclose(out["output_usd"], exp_out), model
def test_openai_unlisted_model_priced_false_not_zero_default():
# o-series / gpt-4.5 are no longer on the pricing page, so we
# intentionally drop them rather than silently underbill at $0.
for model in ("o3", "o4-mini", "gpt-4.5", "gpt-4.5-preview"):
out = calculate_cost(
"openai",
model,
{"input_tokens": 1_000_000, "output_tokens": 1_000_000},
)
assert out["priced"] is False, model
assert out["total_usd"] == 0.0, model
# Token counts still report so the UI can render usage.
assert out["billable_input_tokens"] == 1_000_000, model
assert out["billable_output_tokens"] == 1_000_000, model
# ── canonical Anthropic 4.5 ids now resolve to a price ─────────────
def test_anthropic_canonical_4_5_ids_are_priced():
# Codex P1: claude-opus-4-5 (no date) is the canonical id used
# in backend defaults but was missing from the table, so the
# calculator returned priced=False + zero cost. Pin the aliases.
cases = {
"claude-opus-4-5": (5.0, 25.0),
"claude-sonnet-4-5": (3.0, 15.0),
"claude-haiku-4-5": (1.0, 5.0),
# Opus 4.1 has the same problem.
"claude-opus-4-1": (15.0, 75.0),
}
for model, (inp, outp) in cases.items():
out = calculate_cost(
"anthropic",
model,
{"input_tokens": 1_000_000, "output_tokens": 1_000_000},
)
assert out["priced"] is True, model
assert _isclose(out["input_usd"], inp), model
assert _isclose(out["output_usd"], outp), model
# ── OpenAI long-context tier crossover ──────────────────────────────
def test_openai_gpt55_short_context_under_272k_uses_base_rates():
out = calculate_cost(
"openai",
"gpt-5.5",
{"input_tokens": 100_000, "output_tokens": 5_000},
)
assert _isclose(out["input_usd"], 100_000 / 1_000_000.0 * 5.0)
assert _isclose(out["output_usd"], 5_000 / 1_000_000.0 * 30.0)
# No long-context marker on the model id when we stayed under.
assert "long-context" not in out["model_priced"], out["model_priced"]
def test_openai_gpt55_long_context_crossover_uses_higher_rates():
# 300k billable input > 272k threshold -> long-context tier
# applies to the WHOLE turn, not a per-token blend.
out = calculate_cost(
"openai",
"gpt-5.5",
{"input_tokens": 300_000, "output_tokens": 10_000},
)
assert _isclose(out["input_usd"], 300_000 / 1_000_000.0 * 10.0)
assert _isclose(out["output_usd"], 10_000 / 1_000_000.0 * 45.0)
assert "long-context" in out["model_priced"], out["model_priced"]
def test_openai_gpt54_long_context_crossover():
out = calculate_cost(
"openai",
"gpt-5.4",
{"input_tokens": 500_000, "output_tokens": 20_000},
)
assert _isclose(out["input_usd"], 500_000 / 1_000_000.0 * 5.0)
assert _isclose(out["output_usd"], 20_000 / 1_000_000.0 * 22.5)
def test_openai_gpt54_mini_has_no_long_context_tier():
# Mini/nano/codex don't publish a long-context price; the base
# rate must keep applying even at very large prompts.
out = calculate_cost(
"openai",
"gpt-5.4-mini",
{"input_tokens": 500_000, "output_tokens": 0},
)
assert _isclose(out["input_usd"], 500_000 / 1_000_000.0 * 0.75)
assert "long-context" not in out["model_priced"], out["model_priced"]
# ── OpenAI server-tool surcharges ──────────────────────────────────
def test_openai_web_search_charged_per_thousand():
out = calculate_cost(
"openai",
"gpt-5.5",
{
"input_tokens": 0,
"output_tokens": 0,
"openai_tool_use": {"web_search_requests": 250},
},
)
assert _isclose(
out["server_tools_usd"], 250 / 1_000.0 * OPENAI_WEB_SEARCH_USD_PER_1K
)
assert _isclose(out["total_usd"], 250 / 1_000.0 * OPENAI_WEB_SEARCH_USD_PER_1K)
def test_openai_container_hours_charged():
out = calculate_cost(
"openai",
"gpt-5.5",
{
"input_tokens": 0,
"output_tokens": 0,
"openai_tool_use": {"container_hours": 1.5},
},
)
assert _isclose(out["server_tools_usd"], 1.5 * OPENAI_CONTAINER_USD_PER_HOUR)
def test_openai_tool_surcharges_added_to_total():
# End-to-end: input + output + web_search + container in one
# turn. Total must sum all four buckets.
out = calculate_cost(
"openai",
"gpt-5.5",
{
"input_tokens": 100_000,
"output_tokens": 5_000,
"openai_tool_use": {
"web_search_requests": 3,
"container_hours": 0.25,
},
},
)
expected_input = 100_000 / 1_000_000.0 * 5.0
expected_output = 5_000 / 1_000_000.0 * 30.0
expected_tools = (
3 / 1_000.0 * OPENAI_WEB_SEARCH_USD_PER_1K
+ 0.25 * OPENAI_CONTAINER_USD_PER_HOUR
)
assert _isclose(
out["total_usd"],
round(expected_input + expected_output + expected_tools, 6),
)
# ── snapshot endpoint includes the multipliers ───────────────────────
def test_snapshot_contains_provider_buckets_and_multipliers():
snap = pricing_snapshot()
assert set(snap.keys()) == {"anthropic", "openai"}
a = snap["anthropic"]
o = snap["openai"]
assert "models" in a and "claude-opus-4-7" in a["models"]
assert a["cache_5m_write_mult"] == ANTHROPIC_CACHE_5M_WRITE_MULT
assert a["cache_1h_write_mult"] == ANTHROPIC_CACHE_1H_WRITE_MULT
assert a["cache_read_mult"] == ANTHROPIC_CACHE_READ_MULT
assert a["fast_mode_mult"] == ANTHROPIC_FAST_MODE_MULT
assert "web_search_usd_per_1k" in a
assert "code_execution_usd_per_hour" in a
assert "models" in o and "gpt-5.5" in o["models"]
assert o["cache_read_mult"] == OPENAI_CACHE_READ_MULT
# OpenAI tool surcharge constants are also exposed so the frontend
# tooltip can render the per-call rate.
assert o["web_search_usd_per_1k"] == OPENAI_WEB_SEARCH_USD_PER_1K
assert o["container_usd_per_hour"] == OPENAI_CONTAINER_USD_PER_HOUR
# Long-context tier metadata travels with the model row.
gpt55 = o["models"]["gpt-5.5"]
assert gpt55["long_context_threshold"] == 272_000
assert gpt55["long_context_input_per_mtok"] == 10.0
assert gpt55["long_context_output_per_mtok"] == 45.0