unsloth/studio/backend
Daniel Han fe05b700dc
studio: fix slow cancellation of GGUF generation (#4352)
The streaming loop used response.iter_text() with timeout=None, which
blocks until the next chunk arrives from llama-server. On large models
like Qwen3.5-27B where each token takes seconds, pressing Stop in the
UI would not take effect until the next token was produced.

Fix by using a 0.5s read timeout and a new _iter_text_cancellable()
helper that checks cancel_event between timeout windows and explicitly
closes the response when cancelled. Applied to both the regular chat
completion and tool-calling streaming paths.
2026-03-17 01:47:21 -07:00
..
assets studio: web search, KV cache dtype, training progress, inference fixes 2026-03-17 00:30:01 -07:00
auth fix: remove old comments (#4292) 2026-03-14 16:50:13 +04:00
core studio: fix slow cancellation of GGUF generation (#4352) 2026-03-17 01:47:21 -07:00
loggers Final cleanup 2026-03-12 18:28:04 +00:00
models studio: web search, KV cache dtype, training progress, inference fixes 2026-03-17 00:30:01 -07:00
plugins Final cleanup 2026-03-12 18:28:04 +00:00
requirements studio: web search, KV cache dtype, training progress, inference fixes 2026-03-17 00:30:01 -07:00
routes studio: fix stale GGUF metadata, update helper model, auth improvements (#4346) 2026-03-17 01:22:08 -07:00
state Final cleanup 2026-03-12 18:28:04 +00:00
tests fix(seed): disable remote code execution in seed inspect dataset loads (#4275) 2026-03-13 19:37:43 +04:00
utils studio: fix stale GGUF metadata, update helper model, auth improvements (#4346) 2026-03-17 01:22:08 -07:00
__init__.py Final cleanup 2026-03-12 18:28:04 +00:00
colab.py Fix/colab comment edits (#4317) 2026-03-16 16:15:46 -07:00
main.py studio: web search, KV cache dtype, training progress, inference fixes 2026-03-17 00:30:01 -07:00
run.py fix: Ctrl+C not terminating backend on Linux (#4316) 2026-03-16 11:58:09 +04:00