Daniel Han
51bf500f57
Remove Blackwell flex attention disable workaround from studio ( #4273 )
...
The studio was disabling flex attention entirely on Blackwell+ GPUs
(sm_120 and above) by setting UNSLOTH_ENABLE_FLEX_ATTENTION=0 at
startup. This was a workaround for the flex_attention backward kernel
exceeding shared memory limits on these GPUs.
The root cause is now fixed in unsloth-zoo (PR #542 ) which patches the
backward kernel config selection to generate safe fallback configs that
fit within the GPU's shared memory limit. With that fix, flex attention
works correctly on Blackwell GPUs and provides a ~1.3x speedup over
the SDPA fallback.
2026-03-13 01:35:17 -07:00
Roland Tannous
47654cb91c
Final cleanup
2026-03-12 18:28:04 +00:00
Roland Tannous
a2baf80511
Update license headers
2026-03-12 17:23:10 +00:00
Roland Tannous
9dac1bedf9
Merge remote-tracking branch 'origin/nightly' into feature/llm-assist-detection
2026-03-11 16:23:09 +00:00
Roland Tannous
817f2e8dcc
feat: integrate structlog, configure workers for prod logging, and migrate print statements
2026-03-11 12:33:16 +00:00
Roland Tannous
f7ca361c5c
feat: add LLM-assisted dataset detection using ephemeral GGUF helper
...
Uses Qwen2.5-3B-Instruct Q8_0 via LlamaCppBackend to complement
heuristic-based dataset detection when heuristics are uncertain.
- New llm_assist.py: VLM instruction generation, column classification,
and user-friendly warning generation for dataset issues
- Pre-cache helper GGUF on FastAPI startup (background thread)
- Reorder training pipeline: dataset processing runs BEFORE model load
to avoid VRAM contention (detect → dataset → model → train)
- Add pre_detect_and_load_tokenizer() for lightweight detection
- LLM warnings on VLM conversion failures (broken URLs, missing images)
- LLM column classification fallback when heuristics return unknown
- Graceful degradation: all paths unchanged when helper unavailable
2026-03-10 09:20:45 +00:00
Roland Tannous
d882678fe4
Add AGPL-3.0 SPDX headers to all source files
2026-03-09 20:17:45 +00:00
Roland Tannous
d910759121
feat: add OpenAI-compatible /v1/chat/completions endpoint
2026-03-06 07:48:09 +00:00
Roland Tannous
1167be2798
refactor: consolidate version switching to .venv_t5, remove .venv_overlay
...
All version switching now uses .venv_t5/ (pre-installed by setup.sh).
The old .venv_overlay/ with runtime pip installs is removed.
ensure_transformers_version() (used only by export) now does a
lightweight sys.path swap instead of pip installing at runtime.
2026-03-06 04:37:06 +00:00
Roland Tannous
81b4928e99
Merge nightly into feature/transformers-v5-support
2026-03-05 06:49:44 +00:00
Roland Tannous
a6f1153f9a
fix: replace FileResponse with Response for index.html to prevent Content-Length mismatch and add path traversal guard
2026-02-25 01:05:04 +00:00
Shine1i
8739a01f56
Merge branch 'nightly' into feature/canvas-lab
2026-02-23 21:54:35 +01:00
Roland Tannous
5416bdd4e6
added shutil import to main.py
2026-02-23 07:44:59 +00:00
Roland Tannous
f16a7f2d17
Merge nightly into feature/transformers-v5-support
2026-02-23 07:40:28 +00:00
Roland Tannous
dbbcdb4f09
feat: clear unsloth_compiled_cache on startup, shutdown, and between model loads
2026-02-23 07:26:22 +00:00
Roland Tannous
3fe85d36cc
Remove stale .venv_overlay on server startup to prevent transformers version conflicts
2026-02-23 05:08:27 +00:00
Wasim Yousef Said
6dd0e11439
Merge branch 'nightly' into feature/canvas-lab
2026-02-20 01:23:44 -08:00
Roland Tannous
d57b2742ab
renamed UNSLOTH_FLEX_ATTENTION to UNSLOTH_ENABLE_FLEX_ATTENTION
2026-02-18 09:21:53 +00:00
Roland Tannous
5a02ed4f0f
Disable flex attention on Blackwell+ GPUs (sm_120+) at startup
2026-02-18 08:58:25 +00:00
Roland Tannous
b20d50e8d0
feat: add GET /api/system/hardware endpoint for GPU info and package versions
2026-02-16 10:29:21 +00:00
Shine1i
43c66da783
merge nightly
2026-02-15 14:52:07 +01:00
Shine1i
85653237ea
feat: add Data Recipe core functionality with job manager, API routes, and validation services
2026-02-15 13:43:46 +01:00
Roland Tannous
f453791916
Merge pull request #29 from unslothai/feature/export
...
Added the export routes and pydantic models
2026-02-15 12:42:01 +04:00
Roland Tannous
2840efcc08
fix: add no-cache headers to index.html to prevent stale frontend after rebuild
2026-02-15 06:06:27 +00:00
sshah229
ee703dd6c6
added router in main
2026-02-11 18:51:53 -07:00
sshah229
40bfe42974
added the pydantic models and routes for export
2026-02-11 18:34:12 -07:00
Roland Tannous
7db31723b9
reset DEVICE type on fastapi lifespan exit
2026-02-11 15:58:13 +00:00
Roland Tannous
59d5f24eb5
integrate global hardware detection at lifespan entrypoint
2026-02-11 15:34:26 +00:00
Roland Tannous
28c7df5925
remove unsloth_compiled_cache folder on fastapi lifespan exit
2026-02-11 12:30:05 +00:00
Roland Tannous
01fcb4f713
authentication refactor - added setup token and token refresh mechanism
2026-02-11 12:09:47 +00:00
sshah229
50ff5626f1
refactored the code for username/password and added pydantic models and routes for the same
2026-02-06 03:15:30 -07:00
Roland Tannous
75bb6c08a5
Add datasets check-format endpoint
2026-02-03 20:42:25 +00:00
Roland Tannous
b4ec0389f0
refactor/inference-api-routes-part-1
2026-02-03 16:57:57 +00:00
Roland Tannous
544d6944d1
root studio folder
2026-02-02 09:13:49 +00:00