9.6 KiB
bigboy-alma-deploy
Sovereign AI inference server: RTX 5060 Ti GPU (16GB VRAM), 5-drive btrfs storage on AlmaLinux 10.2
Quick Links
- Architecture & Design:
./planning/ARCHITECTURE.md - Deployment Status:
./planning/STATUS.md - Hardware Manifest:
./planning/HARDWARE.md - Sovereignty Policy:
./planning/SOVEREIGNTY-POLICY.md - Claude Code standing context:
./CLAUDE.md - Current task handoff:
./tasks/
Project Overview
Goal
Deploy BigBoy as a fully autonomous, idempotent AI inference server running
AlmaLinux 10.2 with GPU acceleration via NVIDIA RTX 5060 Ti. The system
serves local models via llama.cpp (llama-server, which includes its
own OpenAI-compatible API and web UI) and maintains reproducible deployment
logs for the RAG corpus. BigBoy also serves as the reference/testbed
machine for CE's future RHEL + llama.cpp client deployments — decisions
made here are expected to generalize, not just work once.
Why AlmaLinux 10?
- RHEL-compatible (10-year support lifecycle)
- Official NVIDIA precompiled GPU driver support
- DNF package manager with clear upgrade path
- Pragmatic middle ground: better stability than NixOS on this hardware (the original NixOS attempt hit unresolved RTX 5060 Ti driver friction), simpler than Debian for this workload
- RHEL-family chosen fleet-wide for AI-server workloads specifically, for compliance/audit reasons (EU AI Act relevance) — not just a fix for this one box
Why llama.cpp, not Ollama?
Decided deliberately, not a default: full control over CUDA build flags
and quantization, and llama-server's built-in web UI removes the need
for a separate Open WebUI layer entirely. Worth knowing: Ollama itself
wraps llama.cpp's ggml engine on NVIDIA/Linux under the hood, so it
wouldn't have avoided this dependency — it would only have added a layer
of indirection around it. See CLAUDE.md for the full standing rationale
(pinned-tag discipline, standard build flags, known sm_120/MXFP4 hazard).
Key Constraints
- GPU VRAM: 16GB (RTX 5060 Ti) — model selection accordingly
- Storage: 5 drives (1 NVMe + 4 SATA) with btrfs subvolume isolation on the SATA drives; NVMe root stays xfs (see Current Status)
- Network: home LAN (
192.168.1.24, static reservation) — originally deployed and tested on the workbench bench LAN (192.168.0.0/24), migrated once chat-interface testing required a GUI-capable device - Idempotency: every phase must be re-runnable
Current Status — read this before assuming anything is done
As of 2026-07-19: Phases 1 through 5 complete and verified on real
hardware. BigBoy is a working AI inference server — NVIDIA driver
loaded, llama.cpp built and pinned, Ministral 3 14B Instruct serving via
llama-server, chat confirmed working through both the built-in web UI
and a terminal client (aichat).
See ./planning/STATUS.md for the authoritative, current per-phase state.
Don't rely on this README for phase-by-phase status — it will drift out
of date faster than STATUS.md is maintained.
Deployed & Verified on Hardware
- Kickstart (
alma10-minimal-bigboy.ks) — AlmaLinux 10.2 minimal install, 2026-07-17 - Phase 2 (system configuration) — sudoers policy, SSH key, base checks
- Phase 3 (filesystem conversion) — all four SATA drives converted to real btrfs; root intentionally stays xfs (no supported in-place conversion path — a full reinstall is deferred ~a month out, after real-world testing, not a gap)
- Phase 4 (NVIDIA driver) — AlmaLinux's precompiled open-kmod path,
confirmed via
nvidia-smi - Phase 5 (llama.cpp) — built at pinned tag
b9968, serving Ministral 3 14B Instruct at Q4_K_M, systemd + nginx + firewall in place, chat confirmed working
Likely dropping
- Phase 8 (Open WebUI) —
llama-server's built-in UI already covers the chat-first use case; this is what Phase 5's success was actually verified against. Keeping this phase now needs an actual reason, not just inertia from the original plan.
Deferred, on purpose
- Phases 6–7, 9, 11–12, 14 — not yet started; see
STATUS.mdfor per-phase detail - Root filesystem's eventual xfs→btrfs conversion — full reinstall, deliberately deferred pending real-world testing
- An auditable testing/release system for llama.cpp (versioned releases,
btrfs snapshot safety net, promote/rollback, a Forgejo mirror that only
ever receives BigBoy-validated tags for client deployments) — fully
designed 2026-07-16, not yet built. See
CLAUDE.mdfor the deferred-work note andscripts/draft/if present. - The per-client agent playbook pattern — same status, see
CLAUDE.md - Borgmatic backup target (not yet decided)
- Phase 13's physical case move — network migration to the home LAN is done; the physical relocation into the new Modcase HYPER case is actively in progress as of 2026-07-19
A known, real gap — not yet closed
The systemd unit, nginx reverse-proxy config, and firewall rules that
make Phase 5's chat deployment actually work currently exist only on
BigBoy itself, configured directly during the live session — not yet
committed to this repo as tracked template files. Until that's fixed,
the deployment isn't fully reproducible from the repo alone. See
STATUS.md's "Immediate" blockers section.
Collaborators & Workflow
John A. Hoeven (@workbench)
- Architect & operator — makes deployment decisions, edits and pushes from workbench via VS Codium
Claude Code (@bigboy, or driven from @workbench)
- Executes scoped tasks against the actual repo — see
CLAUDE.mdfor standing context read automatically each session, andtasks/for the current scoped handoff - Does not decide scope — task docs name exactly what's in and out of bounds for a given session
Claude Desktop
- Plans and orchestrates: architecture discussion, design decisions,
drafts
CLAUDE.mdand task docs for Claude Code to execute against
Repo Structure
bigboy-alma-deploy/
├── CLAUDE.md # standing context, read by Claude Code every session
├── README.md # this file
├── alma10-minimal-bigboy.ks # canonical kickstart — this copy, not planning/'s
├── install-phases.txt # 15-phase (0-14) plan, high level
├── ansible/
│ ├── site.yml # main playbook
│ ├── inventory.ini
│ ├── group_vars/bigboy.yml # hardware facts: UUIDs, driver version, pinned llama.cpp tag
│ └── roles/
│ ├── phase-2-system-config/ # done, verified on hardware
│ └── phase-3-filesystems/ # done, verified on hardware
├── planning/
│ ├── ARCHITECTURE.md
│ ├── HARDWARE.md
│ ├── SOVEREIGNTY-POLICY.md
│ ├── STATUS.md # authoritative current status — check this first
│ ├── alma-*.md # Alma-specific reference docs (firewall, package mgmt, nvidia, etc.)
│ ├── build-journal-20260717.md # detailed build log: drive wipe, install, Phase 3 verification
│ └── alma10-minimal-bigboy.ks # reference snapshot from planning — root copy is canonical
├── tasks/
│ └── TASK-*.md # scoped Claude Code handoffs, one per unit of work
└── scripts/
├── package-recon.sh # confirms package availability before finalizing task docs
├── utility-recon.sh # checks for missing baseline CLI utilities
└── draft/ # draft auditable-release-pipeline scripts —
# unintegrated, see CLAUDE.md before touching
Getting Started
To review design
- Read
./planning/ARCHITECTURE.md— design rationale, 14-phase breakdown - Read
./planning/HARDWARE.md— component specs, UUIDs, network config - Read
./planning/STATUS.md— the real current state, phase by phase
To pick up work
- Check
./planning/STATUS.mdfor what's actually next - Look in
./tasks/for an existing scoped task doc covering it - If none exists, Claude Desktop drafts one before Claude Code starts — don't hand Claude Code a phase without a task doc scoping it
Next Steps (current priority)
- Confirm the tracked templates (
ansible/templates/llama-server.service.j2,llama.conf.j2,zram-generator.conf) and the firewall script (scripts/configure-llama-firewall.sh) all match what's actually configured live on BigBoy — they were written to match the live system, not the other way around, so worth a direct diff - Record the actual NVIDIA driver version in
group_vars/bigboy.yml(placeholder currently readsCHECK-WITH-nvidia-smi) - Decide Phase 8 (Open WebUI) — build or drop
- Complete the physical case move to Modcase HYPER, then proceed with Phases 6–7, 9, 11–12, 14
Notes
install-phases.txtdocuments 15 phases (0–14); this README's status section is a summary, not a replacement for it or forSTATUS.md.- RAG library cross-references from an earlier version of this project
(
nvidia-driver-almalinux.md,ollama-deployment.md,open-webui-deployment.mdunder a separate documents library) predate the llama.cpp decision — verify these still exist and are accurate, or need anllama-cpp-deployment.mdequivalent, before relying on them. - All artifacts are version-controlled and pushed to
https://git.jhoeven.net/giovannino/bigboy-alma-deploy.
Last updated: 2026-07-19 Current priority: Track the systemd/nginx/firewall configs that currently only exist on BigBoy itself; complete the case move; decide Phase 8