14 KiB
BigBoy Deployment Status
Real-time status of all deployment phases
Executive Summary
| Metric | Status |
|---|---|
| Overall Readiness | Foundation + core inference stack complete and verified on real hardware. Phases 6–14 remain, two deliberately deferred. |
| Phases Complete | 1 through 5 (Kickstart, System Config, Filesystems, NVIDIA Driver, llama.cpp) |
| Real Hardware Test | Done — chat confirmed working via web UI and a terminal client, 2026-07-19 |
| Target Go-Live | Base AI Server Deployment goal (chat working) — met. Full 14-phase scope continues. |
Phase Status Breakdown
Phase 1: Base Install (Kickstart) ✓
Status: ✓ COMPLETE — verified on real hardware, 2026-07-17
What was done:
- Unattended AlmaLinux 10.2 install via
alma10-minimal-bigboy.ks - Ten distinct issues found and fixed during install day (see
build-journal-20260717.mdfor the full account):%anacondasection incompatibility,bootloader --locationfix, a line-wrap syntax error,--noipv6correction, btrfs removed from RHEL10's kickstart (switched all partitions to xfs), a cross-repovim/vim-datamismatch,firewall-cmd's inability to run inside the install-time chroot (replaced with a direct zone-file write),systemctl start/restartbeing no-ops in that same chroot, andzerombradded for the genuinely-blank drives. - Password-auth SSH and a temporary NOPASSWD sudoers bootstrap deliberately left active — both closed out properly in Phase 2.
- Disk targeting via
/dev/disk/by-id, notsdX— proven necessary when drives landed on different letters than expected during the actual install, and correctly resolved anyway.
Phase 2: System Configuration ✓
Status: ✓ COMPLETE — verified on real hardware, 2026-07-19
What was done:
- Kickstart's temporary bootstrap sudoers grant replaced with a real, deliberately-scoped policy
- Workbench's SSH public key deployed (single key, by design — see
CLAUDE.mdfor the reasoning on why additional devices' keys wait) - Base system confirmed: hostname, timezone, chrony/NTP, all four expected repos enabled
Real bugs found and fixed during this run: missing ansible.posix
collection (not bundled with the Alpine ansible-core package), a
variable-naming mismatch (bigboy_admin_user vs. the actual
deploy_user) — see CLAUDE.md/group_vars/bigboy.yml for the current
state of both.
Phase 3: Filesystem Conversion ✓
Status: ✓ COMPLETE — verified on real hardware, 2026-07-19
What was done:
- The four SATA data drives converted from the kickstart's forced xfs to
real btrfs (
compress=zstd,noatime), mounted at final paths (/srv/backup,/srv/rag-library,/srv/ai-logs,/srv/prompt-library) - Stale kickstart-generated fstab entries removed
- Root filesystem intentionally NOT converted — stays xfs. No
supported in-place xfs-to-btrfs path exists; a full reinstall (clone/
reformat/restore via
debusb) is deferred until after real-world testing, not before ~a month out from 2026-07-17. This is a decision, documented inCLAUDE.md, not a gap.
Real bugs found and fixed during this run: a device-path bug
({{ item.byid }}1 instead of the correct {{ item.byid }}-part1
suffix), a mount-path mismatch specific to the ai-logs drive
(/srv/ai-logs-raw in the role vs. the kickstart's actual
/srv/ai-raw), and state: unmounted changed to state: absent so
stale fstab entries actually get removed rather than just the live
mount.
Phase 4: NVIDIA GPU Driver ✓
Status: ✓ COMPLETE — verified on real hardware, 2026-07-19
What was done:
- AlmaLinux's precompiled open-kmod path used (not manual DKMS):
almalinux-release-nvidia-driver→nvidia-driver-cuda nvidia-open-kmod, one reboot,nvidia-smiclean on the first attempt - The smoothest phase of the entire deployment — no issues encountered
Still needed: the actual installed driver version has not yet been
recorded in group_vars/bigboy.yml — run nvidia-smi and update the
placeholder there.
Phase 5: llama.cpp ✓
Status: ✓ COMPLETE — verified on real hardware, 2026-07-19
Not Ollama. Built from source, pinned to a specific, vetted tag —
decided deliberately, see CLAUDE.md for the full rationale.
What was done:
- llama.cpp built at pinned tag
b9968(checked against real commit dates and reviewed for fix/regression patterns before selection — standard flags only,-DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES=120) - Model: Ministral 3 14B Instruct, Q4_K_M, pulled via llama.cpp's own
Hugging Face integration (
-hfflag) llama-serverrunning, systemd unit in place, nginx reverse proxy fronting it, firewall opened for the proxy port on both the home LAN and bench LAN ranges, real API key set- Chat confirmed working — both through
llama-server's built-in web UI (tested from Surface over the home LAN) and throughaichat, a terminal client, for LAN segments without GUI access
Real problems found and fixed during this run:
cuda-toolkit(specificallynvcc) is a separate package from the driver runtime — not installed by Phase 4, had to be added separately; the plaincuda-toolkitmeta-package (not a version-pinned sub-package) was used to stay in sync with whatever driver version Phase 4 actually installednvcclands in a versioned path (/usr/local/cuda-13.3/bin), not onPATHby default — a symlink plus a/etc/profile.d/entry resolved this permanently- An OOM kill during the actual CUDA compile (
cc1plus/cudafe++confirmed viadmesg) — root cause was unbounded parallel compilation against memory-hungry CUDA source files with zero swap configured. Fixed with a 16GB NVMe swapfile (persistent via fstab) plus dropping to-j2. zram was separately decided as a complementary fast-swap tier (2026-07-19) and implemented the same day — both are now active simultaneously (zram0at priority 100,/swapfileat priority -2, confirmed viaswapon --show). Seeansible/templates/zram-generator.conffor the tracked config. - Two rounds of missing baseline utilities discovered mid-deployment
(
tar,which, then nine more via a proactiveutility-recon.shscan) — all now added to the kickstart's%packageslist for future installs.
This is the success criterion set on 2026-07-17 ("base AI Server Deployment... Mistral running, chat working, whatever quality, no refinement required") — met.
Phase 6: Build Suite
Status: 📋 READY TO BUILD — not yet started
Development tools / system utilities beyond what's already in the kickstart. Scope not yet finalized against current (llama.cpp-era) needs — the original scope assumed an Ollama-era package list; revisit before building.
Dependencies: Phase 4 (kernel headers matching running kernel) — done
Phase 7: Configuration
Status: 📋 READY TO BUILD — not yet started
Dotfiles, locale, timezone. Scope largely unchanged from original planning; source for dotfiles still TBD.
Dependencies: None blocking
Phase 8: Open WebUI
Status: ⚠️ LIKELY DROPPING — deliberate decision needed, not a default
llama-server's own built-in web UI already covers the chat-first use
case (this is what Phase 5's success was actually verified against).
Building a separate Open WebUI layer on top is no longer the default
plan — keeping this phase requires an actual reason to want a second UI
layer, not inertia from the original Ollama-era plan.
Phase 9: Security Hardening
Status: 📋 DESIGN READY — not yet started
Firewall (firewalld — already partially ahead of schedule, since Phase 5
opened the reverse-proxy port with the same LAN-restriction pattern used
elsewhere), SSH hardening (disable password auth, move to key-only —
this is the actual trigger condition already flagged in CLAUDE.md and
the original kickstart's own changelog), MAC pinning, SELinux review
(currently permissive per the kickstart; confirm AVC denial history
before flipping to enforcing).
Blocker: key-based login needs to be confirmed working over real, repeated use before disabling password auth — not just "it worked once."
Dependencies: Phase 5 (services running) — done
Phase 10: Borgmatic Backups
Status: 🚫 DEFERRED
Reason: Backup target not yet decided (NAS? USB? Aruba Object Storage, per the sovereignty-logging design discussion?)
Timeline: Decide backup target → implement Phase 10
Phase 11: Thermal Baseline Testing
Status: 📋 READY TO BUILD — not yet started
Original scope referenced llama2:70b/Ollama-specific testing — needs
updating to a llama.cpp-appropriate heavy-load test instead. Given the
case move to Modcase HYPER (see below), worth doing this after the
physical case transition, since cooling characteristics will differ from
open-frame testing.
Dependencies: Phase 5 (llama-server running) — done
Phase 12: Full System Validation
Status: 📋 READY TO BUILD — not yet started
Checklist needs updating away from Ollama/Open-WebUI-specific service
checks toward llama-server/nginx.
Dependencies: All prior phases — 1 through 5 done, 6-9/11 pending
Phase 13: Home LAN Migration
Status: 🔨 PARTIALLY UNDERWAY — ahead of original schedule
Original plan: deferred until case installation.
Actual state (2026-07-19): BigBoy is already reachable on the home
LAN (192.168.1.24, statically reserved on the FritzBox) — this
happened ahead of the physical case move, driven by needing to test the
chat interface from a device with a GUI, which the bench LAN doesn't
have. The physical relocation (case: now Modcase HYPER, not the
originally-planned Modcase EVO ITX-2 — printed and ready as of
2026-07-19) is still pending; BigBoy is being assembled into its new
case and permanent home over the next few days.
Modcase HYPER licensing note: personal-use licence purchased 2026-07-19. Do not print for resale or client delivery until the commercial tier is purchased.
Phase 14: Observation Period Runbook
Status: 📋 TEMPLATE READY — not yet started
Dependencies: All prior phases complete + operational — not yet at that point (6, 7, 9, 11, 12 still pending)
Cross-Cutting Concerns
Documentation
| Item | Status |
|---|---|
| Project docs (README, CLAUDE.md, STATUS.md) | 🔨 Being actively reconciled, 2026-07-19 |
| Deployment code (Kickstart, Ansible, inventory) | ✓ Current through Phase 5 |
| RAG library references | ⚠️ Predate the llama.cpp decision (Ollama/Open-WebUI docs) — verify or replace before relying on them |
| Systemd/nginx configs for llama-server | ❌ Only exist live on BigBoy — not yet tracked as reusable template files in this repo. Real gap: deployment isn't currently reproducible from the repo alone for this piece. |
Testing
| Item | Status |
|---|---|
| Real hardware test, Phases 1-5 | ✓ Done, 2026-07-17 through 2026-07-19 |
| Chat interface confirmed | ✓ Web UI + terminal client, 2026-07-19 |
| Idempotency verification | ✓ Confirmed for Phase 3 (re-run-safe checks in place) |
| Integration test (all phases) | ⏳ After Phases 6-9/11-12 built |
Git & Version Control
| Item | Status |
|---|---|
| Forgejo (giovannino/bigboy-alma-deploy) | ✓ Active, pushed, current through Phase 5 |
Critical Path to Go-Live (Full 14-Phase Scope)
Phase 1: Kickstart ✓
↓
Phase 2: System Config ✓
↓
Phase 3: Filesystems ✓ (root stays xfs, deliberately)
↓
Phase 4: GPU Driver ✓
↓
Phase 5: llama.cpp ✓ — chat confirmed working — BASE DEPLOYMENT GOAL MET
↓
Phase 6-7: Build Suite, Configuration 📋
↓
Phase 8: Open WebUI — likely dropping ⚠️
↓
Phase 9: Security Hardening 📋
↓
Phase 11: Thermal Test (post-case-move) 📋
↓
Phase 12: Validation 📋
↓
Phase 14: Observation (4 weeks manual) 📋
↓
✓ FULL GO-LIVE
Parallel: Phase 10 (backup) — deferred, target undecided
Parallel: Phase 13 (home LAN) — network migration done, physical case move in progress
Blockers & Decisions Needed
Immediate
- Real gap: firewall rich-rules script now tracked
(
scripts/configure-llama-firewall.sh); systemd/nginx/zram configs now tracked (ansible/templates/) — confirm all four are actually applied identically on BigBoy vs. what's committed, since the live system was configured first and the templates written to match it afterward - Action: record actual NVIDIA driver version in
group_vars/bigboy.yml
Medium Term
- Decision: Phase 8 (Open WebUI) — build or drop? (leaning drop)
- Blocker: Phase 10 backup target
- Decision: Phase 9's SELinux review — check AVC denial history before flipping enforcing
Long Term
- Root filesystem reformat to btrfs — deliberately deferred ~a month from 2026-07-17, after real-world testing
- The auditable llama.cpp release pipeline and per-client agent
playbook pattern — both designed, neither built (see
CLAUDE.md)
Success Criteria
✓ Base AI Server Deployment (2026-07-17 goal) — MET, 2026-07-19:
llama.cpp built, Ministral 3 14B Instruct running, chat confirmed working
through llama-server's web UI, whatever quality it delivers with zero
tuning.
Full 14-phase deployment success when:
- All phases run end-to-end, idempotently
llama-server/nginx stable- GPU verified healthy under real thermal load (post-case-move)
- Storage verified healthy (btrfs, no corruption)
- 4 weeks observation period completed with no critical issues
Last Updated: 2026-07-19 Next Update: After systemd/nginx configs are tracked, or Phase 6-9 implementation begins Prepared By: Claude (Desktop), reconciled against Claude Code's 2026-07-19 repo audit