unsloth/studio/backend/core
Daniel Han a35fbe22ea
Studio: serialize non-streaming responses once and pool the proxy client (#6393)
* Studio: serialize non-streaming responses once and pool the proxy client

Two safe latency wins on the OpenAI/Anthropic-compatible endpoints that leave
the streaming generation paths untouched (they keep Connection: close and
max_keepalive_connections=0 so a client disconnect still stops GPU decode).

1. Non-streaming responses used JSONResponse(content=model.model_dump()), which
   builds a dict and then re-runs json.dumps. Serialize once with
   model.model_dump_json() via a small _model_json_response helper. The body is
   byte-identical (nulls preserved), about 3x faster to encode in a microbench.

2. The non-streaming completions and embeddings proxies built a fresh
   httpx.AsyncClient per request. Route them through one pooled client
   (core/inference/llama_http) closed on shutdown; streaming generation keeps
   its own per-request close-only client. About 5x faster per call to the
   local llama-server in a microbench.

The existing API-monitor tests for the non-streaming completions, embeddings
and passthrough paths now patch nonstreaming_client instead of httpx.AsyncClient
to match the pooled client, so they stay deterministic.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

* Studio: make the pooled non-streaming client per event loop

Review follow-up on the shared httpx client. It was a single module-global
instance, which has two lifecycle problems the per-request client did not:

1. After aclose() in lifespan shutdown, nonstreaming_client() kept handing back
   the closed client, so a second lifespan in the same process (repeated
   TestClient, embedded restart) failed with "client has been closed".

2. An httpx client binds its transport to the loop it first runs on, so reuse
   from another loop could raise "Event loop is closed".

Hold one client per running loop in a WeakKeyDictionary, recreate when missing
or closed, and close all on shutdown. Single-loop production is unchanged.

* [pre-commit.ci] auto fixes from pre-commit.com hooks

for more information, see https://pre-commit.ci

---------

Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
2026-06-17 22:38:02 -07:00
..
data_recipe Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
export fix: respect absolute export paths to prevent cross-drive copy failures (WinError 112) (#6088) 2026-06-12 12:52:57 +02:00
inference Studio: serialize non-streaming responses once and pool the proxy client (#6393) 2026-06-17 22:38:02 -07:00
rag Studio: project sources backed by RAG (#6205) 2026-06-12 15:42:51 +02:00
training Studio: Xet-primary model downloads with automatic HTTP fallback on stall (#6372) 2026-06-16 06:17:54 -07:00
__init__.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
_torchao_stub.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
tool_healing.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00