* Add Agents settings tab for unsloth start Adds a Settings > Agents tab documenting the `unsloth start` command: quickstart, supported agents with click-to-copy commands, model selection, common options, remote Studio setup, argument pass-through, and a dry-run preview. Agent CLIs found on PATH are badged as installed. Also removes the "New" badge from the System and Chat tabs. * Use official brand logos for agents, invert Ollama and OpenRouter in dark mode Claude Code and OpenAI Codex now use the Anthropic and OpenAI logos from the provider-logos registry; agents without an official asset keep the monogram tile. Also inverts the Ollama and OpenRouter logos in dark mode so their monochrome marks stay visible. * Title Agents tab "Agents (unsloth start)" and move it below Connections The in-tab header now reads "Agents (unsloth start)" while the sidebar label stays "Agents". Reorders the tab to sit below Connections. * Address review: guard PATH detection, fix copy timeout, OS-aware remote snippet - Only probe agent PATH in the desktop app on a loopback backend, so Installed badges are not driven by a remote server's environment. - Show the "none found" note only when detection actually ran and returned empty, not when the call failed. - Share one copy hook that resets its timeout on rapid clicks and clears it on unmount. - Render the Remote Studio snippet with PowerShell syntax on Windows. - Note that --no-launch can still load a model when --model is set. - Drop unused quickstart translation keys. * Add interactive Agents command builder * Add local subagent command guidance * Add official coding agent icons * Use client OS for remote commands, fix copy a11y and model wording (#7303) - Pick the remote snippet shell from the client platform, not the server deviceType - Single-line the model examples so they paste in POSIX, PowerShell and cmd - Split the pass-through block into independent one-command copies - Derive detection visibility instead of clearing state in the effect - Announce copy success to assistive tech - Correct the quickstart/model copy: bare start uses the loaded model * Shell-quote the model, forward the HF token, and fix the quant placeholder - Quote the --model value in the generated and subagent commands so a local path with spaces or metacharacters stays a single argument (client-OS aware) - Pass the saved Hugging Face token to listGgufVariants so gated repos resolve - Show 'No separate quantization' instead of a stuck 'Loading quantizations...' when a model has no variants; clear the failure once a later request succeeds * Fix Agents command discovery and routing * Unsloth start improvements: download progress, server reuse, and safe model switching (#7313) * Improve unsloth start runtime lifecycle * Remove speculative Gemma prompt override * Polish model download progress output * Refine unsloth start status output * Clarify unsloth readiness banner * Clarify model reuse and switching output * Queue model switches behind active inference * Tighten unsloth start model switching * Reduce model switch bookkeeping * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Fix Studio re-exec compatibility * Recheck sidecar reservation after inference drain * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Pass start marker through child environment * Fix key redaction, switch-waiter ordering, and stop/messaging gaps for PR #7313 - Redact minted sk-unsloth keys from the startup-failure log tail: the early key marker lands in the server log before the model load finishes, so a load-phase crash printed a live key to the terminal - Deregister a finished switch waiter before releasing the swap gate so a swap on another event loop cannot count it as still queued and unload the model the finished request is about to generate against - Warn on same-repo quant switches: an explicit variant replaces the resident weights for every attached session, but the repo ids match so no switch warning was printed - Note the agent exit code when it is nonzero so the server keep-alive message does not read as a successful session - Use taskkill /T in unsloth studio stop so llama-server children stop too * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Tighten comments in start, studio, and inference changes --------- Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com> Co-authored-by: Daniel Han <danielhanchen@gmail.com> * Unsloth start: add local subagents for Claude Code, Codex, OpenCode and Pi (#7326) Bring the local-subagent support onto main. The original change (#7316) merged into the stacked pr/daniel-unsloth-start-audit branch rather than main, and #7313 reached main via squash, so these files never landed on main. Adds --as-subagent for claude, codex, opencode and pi: the parent agent keeps its own cloud model while a locally served GGUF is registered as a delegated subagent, using ephemeral per-session config that never touches the user's real agent config. * Fix Agents builder defaults and flag validation * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Fix Agents variant and provider fallbacks * Fix local model and Pi subagent edge cases * Agents tab: flag the Codex row when the loaded model is not GGUF * Agents tab: target the active Studio server, wrap narrow rows, index the tab's search terms * Agents tab: build copied commands from the browser-reachable Studio and show the key placeholder * Preserve cache load ids and path variants in built commands for PR #7312 A GGUF outside the active Hugging Face cache only loads by its snapshot path, so keep that load_id for --model while still listing the row by repo id. Path based models carry their quant in --gguf-variant rather than a ":variant" suffix, and the active selection now keeps the variant inference status reports for them. * Agents tab: index the intro for agent-name searches and keep long commands inside the panel * List GGUF variants from the cache the command loads from for PR #7312 A snapshot outside the active Hugging Face cache was offering the remote variant list, so a quant absent from that snapshot could be selected and the generated command would fail to load it. * Agents tab: omit --api-key so the CLI can replay a saved key for the base * Agents tab: label the indexed heading rows and fall back to the active desktop API base * Agents tab: name every supported agent in the indexed intro for PR #7303 * Send the cached GGUF load path and fix the agents tab search targets for PR #7312 * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Tighten the agents tab comments for PR #7303 * Build the agents tab example commands from the active Studio base for PR #7303 * Keep the resident model on its active cache load for PR #7312 * Tighten the agents tab and cached GGUF comments for PR #7312 * Take the agent command shell from the Studio host for PR #7303 * Stop emitting snapshot paths as --model and keep unsloth start searchable for PR #7312 * Pick the command shell from where the CLI runs for PR #7303 * Match a path load by its advertised id and follow the resident model for PR #7312 * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Keep an explicit quantization and retire superseded native-grant labels for PR #7312 * Scope the remembered quant, stop following unloaded models and keep local GGUF paths for PR #7312 * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Stop shadowing the path classifier, match snapshot ordering and sequence status polls for PR #7312 * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Release stale native-grant picks, keep local GGUF identities and index snapshot aliases for PR #7312 * Index inactive-cache snapshots, widen local GGUF detection and clear retired quants for PR #7312 * Classify cached repos by snapshot, merge repo ids case-insensitively and keep loose GGUFs variantless for PR #7312 * Fix snapshot alias, partial split and mmproj-only handling for PR #7312 * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Trust scanned model_format and drop incomplete snapshot ids for PR #7312 * Exclude mmproj and partial downloads, keep path case and drop duplicate scan for PR #7312 * [pre-commit.ci] auto fixes from pre-commit.com hooks for more information, see https://pre-commit.ci * Restrict revision aliases and require complete snapshot variants for PR #7312 * Index revisions individually and hide partial variants for PR #7312 --------- Co-authored-by: shimmyshimmer <107991372+shimmyshimmer@users.noreply.github.com> Co-authored-by: Daniel Han <danielhanchen@gmail.com> Co-authored-by: oobabooga <oobabooga4@gmail.com> Co-authored-by: pre-commit-ci[bot] <66853113+pre-commit-ci[bot]@users.noreply.github.com>
409 lines
12 KiB
TypeScript
409 lines
12 KiB
TypeScript
import { spawn, type ChildProcess } from "node:child_process";
|
|
import * as fs from "node:fs";
|
|
import * as path from "node:path";
|
|
import { fileURLToPath } from "node:url";
|
|
import type { ExtensionAPI } from "@earendil-works/pi-coding-agent";
|
|
import { Type } from "typebox";
|
|
|
|
// Distinct from the normal `unsloth` provider: subagent mode preserves the user's Pi config.
|
|
const provider = "unsloth-studio-subagent";
|
|
const maxResultCharacters = 100_000;
|
|
const maxParallelAgents = 4;
|
|
const cancelGraceMilliseconds = 2_000;
|
|
const configPath = process.env.UNSLOTH_PI_SUBAGENT_CONFIG || "";
|
|
delete process.env.UNSLOTH_PI_SUBAGENT_CONFIG;
|
|
let config: Record<string, unknown> = {};
|
|
if (configPath) {
|
|
try {
|
|
const parsed = JSON.parse(fs.readFileSync(configPath, "utf8"));
|
|
if (!parsed || typeof parsed !== "object" || Array.isArray(parsed)) {
|
|
throw new Error("expected a JSON object");
|
|
}
|
|
config = parsed;
|
|
} catch (error) {
|
|
throw new Error(`Could not read Unsloth subagent configuration: ${error}`);
|
|
}
|
|
}
|
|
const model = typeof config.model === "string" ? config.model : "";
|
|
const baseUrl = typeof config.baseUrl === "string" ? config.baseUrl : "";
|
|
const apiKey = typeof config.apiKey === "string" ? config.apiKey : "";
|
|
const approve = config.approve === true;
|
|
const contextWindow = positiveInt(config.contextWindow, 32768);
|
|
const maxTokens = positiveInt(config.maxTokens, Math.min(Math.floor(contextWindow / 4), 8192));
|
|
let activeAgents = 0;
|
|
const waitingAgents: Array<() => boolean> = [];
|
|
|
|
function positiveInt(value: unknown, fallback: number): number {
|
|
const parsed = Number.parseInt(typeof value === "string" ? value : String(value || ""), 10);
|
|
return Number.isFinite(parsed) && parsed > 0 ? parsed : fallback;
|
|
}
|
|
|
|
function finalText(message: any): string {
|
|
if (message?.role !== "assistant" || !Array.isArray(message.content)) return "";
|
|
return message.content
|
|
.filter((part: any) => part?.type === "text" && typeof part.text === "string")
|
|
.map((part: any) => part.text)
|
|
.join("\n")
|
|
.trim();
|
|
}
|
|
|
|
function boundedResult(text: string): string {
|
|
if (text.length <= maxResultCharacters) return text;
|
|
return `${text.slice(0, maxResultCharacters)}\n\n[Local agent output truncated]`;
|
|
}
|
|
|
|
function agentSlotRelease(): () => void {
|
|
let released = false;
|
|
return () => {
|
|
if (released) return;
|
|
released = true;
|
|
while (waitingAgents.length) {
|
|
if (waitingAgents.shift()!()) return;
|
|
}
|
|
activeAgents -= 1;
|
|
};
|
|
}
|
|
|
|
function acquireAgentSlot(signal: AbortSignal | undefined): Promise<() => void> {
|
|
if (signal?.aborted) return Promise.reject(new Error("The local Unsloth agent was cancelled."));
|
|
if (activeAgents < maxParallelAgents) {
|
|
activeAgents += 1;
|
|
return Promise.resolve(agentSlotRelease());
|
|
}
|
|
return new Promise((resolve, reject) => {
|
|
let waiting = true;
|
|
const grant = () => {
|
|
if (!waiting) return false;
|
|
waiting = false;
|
|
signal?.removeEventListener("abort", cancel);
|
|
resolve(agentSlotRelease());
|
|
return true;
|
|
};
|
|
const cancel = () => {
|
|
if (!waiting) return;
|
|
waiting = false;
|
|
const index = waitingAgents.indexOf(grant);
|
|
if (index >= 0) waitingAgents.splice(index, 1);
|
|
reject(new Error("The local Unsloth agent was cancelled."));
|
|
};
|
|
waitingAgents.push(grant);
|
|
signal?.addEventListener("abort", cancel, { once: true });
|
|
});
|
|
}
|
|
|
|
function piInvocation(args: string[]): { command: string; args: string[] } {
|
|
const currentScript = process.argv[1];
|
|
const bunVirtualScript = currentScript?.startsWith("/$bunfs/root/");
|
|
if (currentScript && !bunVirtualScript && fs.existsSync(currentScript)) {
|
|
return { command: process.execPath, args: [currentScript, ...args] };
|
|
}
|
|
const executable = path.basename(process.execPath).toLowerCase();
|
|
if (!/^(node|bun)(\.exe)?$/.test(executable)) return { command: process.execPath, args };
|
|
return { command: "pi", args };
|
|
}
|
|
|
|
function signalProcessGroup(child: ChildProcess, signal: NodeJS.Signals): void {
|
|
if (!child.pid) return;
|
|
try {
|
|
process.kill(-child.pid, signal);
|
|
} catch {
|
|
try {
|
|
child.kill(signal);
|
|
} catch {
|
|
// The process tree already exited.
|
|
}
|
|
}
|
|
}
|
|
|
|
async function stopChildTree(child: ChildProcess): Promise<void> {
|
|
if (!child.pid) return;
|
|
if (process.platform === "win32") {
|
|
await new Promise<void>((resolve) => {
|
|
const killer = spawn("taskkill", ["/PID", String(child.pid), "/T", "/F"], {
|
|
shell: false,
|
|
stdio: "ignore",
|
|
windowsHide: true,
|
|
});
|
|
killer.once("error", () => {
|
|
try {
|
|
child.kill("SIGKILL");
|
|
} catch {
|
|
// The child already exited.
|
|
}
|
|
resolve();
|
|
});
|
|
killer.once("close", (code) => {
|
|
if (code !== 0) {
|
|
try {
|
|
child.kill("SIGKILL");
|
|
} catch {
|
|
// The child already exited.
|
|
}
|
|
}
|
|
resolve();
|
|
});
|
|
});
|
|
return;
|
|
}
|
|
|
|
signalProcessGroup(child, "SIGTERM");
|
|
await new Promise((resolve) => setTimeout(resolve, cancelGraceMilliseconds));
|
|
signalProcessGroup(child, "SIGKILL");
|
|
}
|
|
|
|
interface LocalAgentResult {
|
|
task: string;
|
|
response: string;
|
|
transcript: any[];
|
|
error?: string;
|
|
}
|
|
|
|
async function runLocalAgent(
|
|
task: string,
|
|
cwd: string,
|
|
signal: AbortSignal | undefined,
|
|
onProgress: (result: LocalAgentResult) => void,
|
|
): Promise<LocalAgentResult> {
|
|
const extension = fileURLToPath(import.meta.url);
|
|
const args = [
|
|
"--mode",
|
|
"json",
|
|
"--print",
|
|
"--no-session",
|
|
...(approve ? ["--approve"] : []),
|
|
"--provider",
|
|
provider,
|
|
"--model",
|
|
model,
|
|
"--no-extensions",
|
|
"--extension",
|
|
extension,
|
|
`Task: ${task}`,
|
|
];
|
|
const invocation = piInvocation(args);
|
|
let output = "";
|
|
let stderr = "";
|
|
let childError = "";
|
|
let aborted = false;
|
|
const result: LocalAgentResult = { task, response: "", transcript: [] };
|
|
const transcriptEntries = new Set<string>();
|
|
const appendTranscript = (messages: any[]): boolean => {
|
|
let changed = false;
|
|
for (const message of messages) {
|
|
const entry = JSON.stringify(message);
|
|
if (transcriptEntries.has(entry)) continue;
|
|
transcriptEntries.add(entry);
|
|
result.transcript.push(message);
|
|
changed = true;
|
|
}
|
|
return changed;
|
|
};
|
|
const processLine = (line: string) => {
|
|
try {
|
|
const event = JSON.parse(line);
|
|
if (event.type === "message_end" && event.message && appendTranscript([event.message])) {
|
|
onProgress(result);
|
|
}
|
|
if (
|
|
event.type === "turn_end" &&
|
|
Array.isArray(event.toolResults) &&
|
|
event.toolResults.length &&
|
|
appendTranscript(event.toolResults)
|
|
) {
|
|
onProgress(result);
|
|
}
|
|
if (event.type !== "message_end") return;
|
|
const message = event.message;
|
|
// Pi reports model/API failures as message_end events while still
|
|
// exiting 0, so the exit status alone cannot surface them.
|
|
if (message?.stopReason === "error" || message?.stopReason === "aborted") {
|
|
childError =
|
|
(typeof message.errorMessage === "string" && message.errorMessage) ||
|
|
`The local Unsloth agent stopped: ${message.stopReason}.`;
|
|
return;
|
|
}
|
|
const response = finalText(message);
|
|
if (response) {
|
|
result.response = boundedResult(response);
|
|
childError = "";
|
|
}
|
|
} catch {
|
|
// Ignore non-JSON diagnostic lines. The exit status still reports failures.
|
|
}
|
|
};
|
|
|
|
const exitCode = await new Promise<number>((resolve, reject) => {
|
|
const child = spawn(invocation.command, invocation.args, {
|
|
cwd,
|
|
detached: process.platform !== "win32",
|
|
shell: false,
|
|
stdio: ["ignore", "pipe", "pipe"],
|
|
env: {
|
|
...process.env,
|
|
UNSLOTH_PI_SUBAGENT_CHILD: "1",
|
|
UNSLOTH_PI_SUBAGENT_CONFIG: configPath,
|
|
},
|
|
});
|
|
let cleanup: Promise<void> | undefined;
|
|
const cancel = () => {
|
|
if (aborted) return;
|
|
aborted = true;
|
|
cleanup = stopChildTree(child);
|
|
};
|
|
child.on("error", (error) => {
|
|
signal?.removeEventListener("abort", cancel);
|
|
reject(error);
|
|
});
|
|
child.stdout.on("data", (chunk) => {
|
|
output += chunk.toString();
|
|
const lines = output.split("\n");
|
|
output = lines.pop() || "";
|
|
for (const line of lines) processLine(line);
|
|
});
|
|
child.stderr.on("data", (chunk) => {
|
|
stderr = (stderr + chunk.toString()).slice(-100_000);
|
|
});
|
|
child.on("close", async (code) => {
|
|
signal?.removeEventListener("abort", cancel);
|
|
await cleanup;
|
|
if (output.trim()) processLine(output);
|
|
resolve(code ?? 1);
|
|
});
|
|
signal?.addEventListener("abort", cancel, { once: true });
|
|
if (signal?.aborted) cancel();
|
|
});
|
|
|
|
if (aborted) throw new Error("The local Unsloth agent was cancelled.");
|
|
if (exitCode !== 0) {
|
|
result.error = stderr.trim() || `The local Unsloth agent exited with code ${exitCode}.`;
|
|
}
|
|
if (childError) result.error = boundedResult(childError);
|
|
if (!result.response && !result.error) result.response = "The local agent returned no text.";
|
|
return result;
|
|
}
|
|
|
|
export default function unslothSubagent(pi: ExtensionAPI): void {
|
|
if (!model || !baseUrl || !apiKey || !configPath) {
|
|
throw new Error("Unsloth subagent configuration is incomplete.");
|
|
}
|
|
|
|
pi.registerProvider(provider, {
|
|
name: "Unsloth Studio",
|
|
baseUrl,
|
|
apiKey,
|
|
api: "openai-completions",
|
|
authHeader: true,
|
|
models: [
|
|
{
|
|
id: model,
|
|
name: `${model} via Unsloth`,
|
|
reasoning: false,
|
|
input: ["text"],
|
|
cost: { input: 0, output: 0, cacheRead: 0, cacheWrite: 0 },
|
|
contextWindow,
|
|
maxTokens,
|
|
},
|
|
],
|
|
});
|
|
|
|
if (process.env.UNSLOTH_PI_SUBAGENT_CHILD === "1") return;
|
|
|
|
pi.registerTool({
|
|
name: "unsloth_agent",
|
|
label: "Unsloth agent",
|
|
description:
|
|
"Run local coding agents powered by Unsloth for debugging, implementation, and codebase research. Use task for one agent. To run multiple independent agents, use tasks; up to four run concurrently. The tool returns only after every requested agent finishes.",
|
|
parameters: Type.Object({
|
|
task: Type.Optional(
|
|
Type.String({ description: "The complete task for one local Unsloth agent." }),
|
|
),
|
|
tasks: Type.Optional(
|
|
Type.Array(Type.String({ description: "A complete task for one local Unsloth agent." }), {
|
|
description: "Independent tasks to run concurrently, one local agent per task.",
|
|
minItems: 2,
|
|
maxItems: maxParallelAgents,
|
|
}),
|
|
),
|
|
}),
|
|
executionMode: "parallel",
|
|
async execute(_toolCallId, params, signal, onUpdate, ctx) {
|
|
const singleTask = typeof params.task === "string" && params.task.trim() ? params.task.trim() : "";
|
|
const parallelTasks = Array.isArray(params.tasks)
|
|
? params.tasks.map((task) => task.trim()).filter(Boolean)
|
|
: [];
|
|
if (Boolean(singleTask) === Boolean(parallelTasks.length)) {
|
|
throw new Error("Provide exactly one of task or tasks.");
|
|
}
|
|
if (parallelTasks.length > maxParallelAgents) {
|
|
throw new Error(`At most ${maxParallelAgents} local agents can run concurrently.`);
|
|
}
|
|
if (parallelTasks.length === 1) {
|
|
throw new Error("Use task for one local agent, or tasks for two to four agents.");
|
|
}
|
|
|
|
const tasks = singleTask ? [singleTask] : parallelTasks;
|
|
const results: Array<LocalAgentResult | undefined> = new Array(tasks.length);
|
|
let completed = 0;
|
|
const details = () => ({
|
|
provider,
|
|
model,
|
|
mode: tasks.length === 1 ? "single" : "parallel",
|
|
results: results.filter((result): result is LocalAgentResult => Boolean(result)),
|
|
});
|
|
const emitUpdate = () => {
|
|
onUpdate?.({
|
|
content: [
|
|
{
|
|
type: "text",
|
|
text: `Local agents: ${completed}/${tasks.length} completed`,
|
|
},
|
|
],
|
|
details: details(),
|
|
});
|
|
};
|
|
await Promise.all(
|
|
tasks.map(async (task, index) => {
|
|
let releaseAgentSlot: (() => void) | undefined;
|
|
try {
|
|
releaseAgentSlot = await acquireAgentSlot(signal);
|
|
results[index] = await runLocalAgent(task, ctx.cwd, signal, (partial) => {
|
|
results[index] = partial;
|
|
emitUpdate();
|
|
});
|
|
} catch (error) {
|
|
results[index] = {
|
|
task,
|
|
response: "",
|
|
transcript: results[index]?.transcript || [],
|
|
error: String(error),
|
|
};
|
|
} finally {
|
|
releaseAgentSlot?.();
|
|
completed += 1;
|
|
emitUpdate();
|
|
}
|
|
}),
|
|
);
|
|
if (signal?.aborted) throw new Error("The local Unsloth agent was cancelled.");
|
|
const completedResults = results.filter(
|
|
(result): result is LocalAgentResult => Boolean(result),
|
|
);
|
|
const succeeded = completedResults.filter((result) => !result.error).length;
|
|
const response =
|
|
completedResults.length === 1
|
|
? completedResults[0].error || completedResults[0].response
|
|
: [
|
|
`Parallel: ${succeeded}/${tasks.length} local agents succeeded`,
|
|
...completedResults.map(
|
|
(result, index) =>
|
|
`\n### Agent ${index + 1}${result.error ? " failed" : ""}\n\n${result.error || result.response}`,
|
|
),
|
|
].join("\n");
|
|
if (succeeded !== completedResults.length) throw new Error(response);
|
|
return {
|
|
content: [{ type: "text", text: response }],
|
|
details: details(),
|
|
};
|
|
},
|
|
});
|
|
}
|