* Rebuild Studio branch on top of main * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Fix security and code quality issues for Studio PR #4237 - Validate models_dir query param against allowed directory roots to prevent path traversal in /api/models/local endpoint - Replace string startswith() with Path.is_relative_to() for frontend path traversal check in serve_frontend - Sanitize SSE error messages to not leak exception details to clients (4 locations in inference.py) - Bind port-discovery socket to 127.0.0.1 instead of all interfaces in llama_cpp backend - Import datasets_root and resolve_output_dir in embedding training function to fix NameError and use managed output directory - Remove stale .gitignore entries for package-lock.json and test directories so tests can be tracked in version control - Add venv-reexecution logic to ui CLI command matching the studio command behavior * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Move models_dir path validation before try/except block The HTTPException(403) was inside the try/except Exception handler, so it would be caught and re-raised as a 500. Moving the validation before the try block ensures the 403 is returned directly and also makes the control flow clearer for static analysis (path is validated before any filesystem operations). * Use os.path.realpath + startswith for models_dir validation CodeQL py/path-injection does not recognize Path.is_relative_to() as a sanitizer. Switched to os.path.realpath + str.startswith which is a recognized sanitizer pattern in CodeQL's taint analysis. The startswith check uses root_str + os.sep to prevent prefix collisions (e.g. /app/models_evil matching /app/models). * Never pass user input to Path constructor in models_dir validation CodeQL traces taint through Path(resolved) even after a startswith barrier guard. Fix: the user-supplied models_dir is only used as a string for comparison against allowed roots. The Path object passed to _scan_models_dir comes from the trusted allowed_roots list, not from user input. This fully breaks the taint chain. --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
42 lines
889 B
YAML
42 lines
889 B
YAML
model: unsloth/Qwen2.5-0.5B
|
|
|
|
data:
|
|
dataset: tatsu-lab/alpaca
|
|
format_type: auto
|
|
|
|
training:
|
|
training_type: lora
|
|
max_seq_length: 2048
|
|
load_in_4bit: true
|
|
output_dir: outputs
|
|
num_epochs: 1
|
|
learning_rate: 0.0002
|
|
batch_size: 2
|
|
gradient_accumulation_steps: 4
|
|
warmup_steps: 5
|
|
max_steps: 0
|
|
save_steps: 0
|
|
weight_decay: 0.01
|
|
random_seed: 3407
|
|
packing: false
|
|
train_on_completions: false
|
|
gradient_checkpointing: "unsloth"
|
|
|
|
lora:
|
|
lora_r: 64
|
|
lora_alpha: 16
|
|
lora_dropout: 0.0
|
|
target_modules: "q_proj,k_proj,v_proj,o_proj,gate_proj,up_proj,down_proj"
|
|
vision_all_linear: false
|
|
use_rslora: false
|
|
use_loftq: false
|
|
finetune_vision_layers: true
|
|
finetune_language_layers: true
|
|
finetune_attention_modules: true
|
|
finetune_mlp_modules: true
|
|
|
|
logging:
|
|
enable_wandb: false
|
|
wandb_project: unsloth-training
|
|
enable_tensorboard: false
|
|
tensorboard_dir: runs
|