unsloth/studio/backend/core/inference
Unsloth 6492d36719 Address review round 4: refcount concurrent guards, distrust proxy TCP
- Overlapping requests lost offline mid-flight. A later guard saw the
  HF_HUB_OFFLINE that an earlier one had set and took the no-op branch, so
  when the earlier guard exited it restored the constants and sessions while
  the later request was still resolving hub files, dropping it back onto the
  retry path. Each guard now holds its own reference on the refcounted
  force_hf_offline window. A user-supplied offline variable is still left
  untouched, told apart via force_hf_offline_active().
- The socket-timeout fallback trusted a TCP handshake to the proxy, which
  only proves the proxy is up, not that it can reach the hub. A live proxy
  with a blackholed upstream therefore read as reachable. With a proxy
  configured the timeout now stays unreachable; the TCP check is only
  evidence when connecting to the endpoint directly.

Verified: second guard engages and offline survives the first guard's exit,
state fully restored after both; dead-upstream proxy reads unreachable while
a slow direct endpoint still reads reachable; 9 concurrent metadata requests
against an unreachable hub all return 200 in 5.1s total.
2026-07-29 01:49:45 -07:00
..
sandbox_site Pin utf-8 on shipping-code text I/O instead of the operator locale (#7486) 2026-07-27 02:14:20 -07:00
__init__.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
_html_to_md.py Studio: stream live tool output with SSE heartbeats, fix web page extraction, and surface interrupted turns (#7083) 2026-07-15 08:41:00 -07:00
_vulkan_probe.py Vulkan GPUs: real device names and selectable ordinals (rebase of #7356 onto #7476) (#7498) 2026-07-27 05:21:48 -07:00
anthropic_compat.py Fix Claude client tools under server tool policy (#7518) 2026-07-28 03:14:05 -07:00
api_monitor.py Studio: add option to disable the in-memory API monitor (#7156) 2026-07-27 22:47:24 -03:00
audio_codecs.py Studio: add configurable model download location (#7274) 2026-07-23 01:34:38 -07:00
chat_eos.py Studio: stop chat generation on the assistant-turn-end token (fixes Qwen3.5 loop) (#6804) 2026-07-06 10:07:56 -07:00
chat_template_helpers.py Studio: split parallel tool calls for Llama 3.x chat templates (#7426) 2026-07-27 20:35:34 +01:00
chat_templates.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
defaults.py Add DeepSeek-V4-Flash-GGUF to Studio with none/high/max reasoning (#6908) 2026-07-07 06:13:43 -07:00
external_provider.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
inference.py feat(studio): run chats in parallel in the Chat tab (#7455) 2026-07-28 04:40:38 -07:00
key_exchange.py Reduce and tighten code comments and docstrings repo-wide (#6095) 2026-06-08 23:09:51 -07:00
llama_admission.py Studio: bound how many tool approvals may park their slot (#7496) 2026-07-28 14:49:21 -07:00
llama_cpp.py Address review round 4: refcount concurrent guards, distrust proxy TCP 2026-07-29 01:49:45 -07:00
llama_http.py fix(studio/llama_cpp): disable trust_env on the loopback health probe (#6750) (#6752) 2026-06-30 19:09:26 +02:00
llama_keepwarm.py Studio: tighten the comments added by the OpenAI model-admission work (#7501) 2026-07-27 05:59:03 -07:00
llama_server_args.py feat(studio): adjustable llama-server parallel slots from the web UI (#7447) 2026-07-28 18:03:28 -07:00
llama_stats.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
local_model_resolver.py Studio: tighten the comments added by the OpenAI model-admission work (#7501) 2026-07-27 05:59:03 -07:00
mcp_client.py Studio: pass raise_on_error=False on the stdio MCP call path (#7517) 2026-07-28 19:41:47 -03:00
mcp_config_import.py studio: show MCP "Import config" on the add-server form (#6030) 2026-06-11 16:17:22 +01:00
message_content.py fix(studio): handle multimodal list content in inference text paths (#4383) (#6480) 2026-06-23 01:26:11 -07:00
mlx_inference.py feat(studio): run chats in parallel in the Chat tab (#7455) 2026-07-28 04:40:38 -07:00
model_ids.py Studio: tighten the comments added by the OpenAI model-admission work (#7501) 2026-07-27 05:59:03 -07:00
openai_auto_download.py Studio: tighten the comments added by the OpenAI model-admission work (#7501) 2026-07-27 05:59:03 -07:00
orchestrator.py feat(studio): run chats in parallel in the Chat tab (#7455) 2026-07-28 04:40:38 -07:00
passthrough_healing.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
presence_penalty.py Studio: apply presence_penalty on the safetensors and MLX inference paths (#6923) 2026-07-06 22:24:47 -07:00
pricing.py Replace standalone Studio wording with Unsloth (#7221) 2026-07-19 00:47:04 -07:00
providers.py Allow API key for Ollama connections (#7173) 2026-07-18 22:47:00 -07:00
runtime_context.py Expose runtime context length for hub models (#6154) 2026-06-11 22:13:53 +03:00
safetensors_agentic.py Studio: surface the tool-call nudge in the chat UI (#7559) 2026-07-28 18:20:50 -07:00
stt_ggml_sidecar.py Studio: add local speech-to-text dictation engine (#7095) 2026-07-23 01:39:03 -07:00
stt_sidecar.py Studio STT: only load safetensors weights for custom dictation models (RCE fix) (#7364) 2026-07-23 03:15:45 -07:00
tensor_fallback.py studio: deterministic VRAM auto-fit for GGUF (MTP reserve, compute buffer, total-based budget) (#6312) 2026-06-17 03:10:22 -07:00
tool_call_parser.py Studio: surface the tool-call nudge in the chat UI (#7559) 2026-07-28 18:20:50 -07:00
tool_loop_controller.py feat(studio): run chats in parallel in the Chat tab (#7455) 2026-07-28 04:40:38 -07:00
tool_stream_exec.py Studio: stream live tool output with SSE heartbeats, fix web page extraction, and surface interrupted turns (#7083) 2026-07-15 08:41:00 -07:00
tools.py Gate the sed commands that run a shell (#7483) 2026-07-28 05:49:51 -07:00
web_access_policy.py Studio: add Deep Research (#7219) 2026-07-26 23:36:02 -07:00
worker.py fix(studio): activate MLX inference sidecar before detection (#7402) 2026-07-27 18:27:31 -03:00