unsloth/studio/backend/utils/datasets
Daniel Han 34cdbf42bf studio: training progress visibility + deferred llama.cpp compilation
Training progress:
- Show row counts in status messages: "Loaded dataset from HuggingFace:
  Open-Orca/OpenOrca (4,233,923 rows)" instead of just the dataset name
- Emit "Formatting dataset (N rows)..." and "Applying chat template
  (N rows)..." status updates so users see progress during the
  preprocessing stages that previously appeared stuck

Deferred llama.cpp compilation:
- Add LlamaCppBuilder that runs cmake build in a background thread
  at server startup if the llama-server binary is missing
- Studio starts immediately and is usable for training/non-GGUF tasks
  while llama.cpp compiles in the background
- GGUF model loads wait for the build to finish with a helpful message
- Add /api/inference/llama-cpp-status endpoint for build status
- Frontend shows "Waiting for llama.cpp to compile..." toast when
  loading a GGUF while build is in progress
2026-03-16 14:02:20 +00:00
..
__init__.py Update license headers 2026-03-12 17:23:10 +00:00
chat_templates.py Final cleanup 2026-03-12 18:28:04 +00:00
data_collators.py Final cleanup 2026-03-12 18:28:04 +00:00
dataset_utils.py studio: training progress visibility + deferred llama.cpp compilation 2026-03-16 14:02:20 +00:00
format_conversion.py Final cleanup 2026-03-12 18:28:04 +00:00
format_detection.py Final cleanup 2026-03-12 18:28:04 +00:00
llm_assist.py Improve AI Assist: Update default model, model output parsing, logging, and dataset mapping UX (#4323) 2026-03-16 16:04:35 +04:00
model_mappings.py Final cleanup 2026-03-12 18:28:04 +00:00
vlm_processing.py Final cleanup 2026-03-12 18:28:04 +00:00