1. n_gpu_layers kwarg: accept (and ignore) in load_model signature so callers like llm_assist.py don't get TypeError 2. mmproj exclusion: filter out mmproj files in _find_smallest_fitting_variant so fallback doesn't pick a tiny vision projection as the "model" 3. Shard preservation after fallback: re-discover shards for the fallback variant instead of resetting to empty list, so split GGUFs download all shards 4. Orphan cleanup safety: only kill llama-server processes whose cmdline contains ".unsloth/", avoiding termination of unrelated llama-server instances on the same machine 5. Path expression sanitization: validate repo_id format before using it in cache directory lookups |
||
|---|---|---|
| .. | ||
| __init__.py | ||
| audio_codecs.py | ||
| inference.py | ||
| llama_cpp.py | ||
| orchestrator.py | ||
| worker.py | ||