Update build journal to complete phase 3-5

This commit is contained in:
John A. Hoeven 2026-07-31 00:42:01 +02:00
commit e5c1bae8de
Signed by: giovannino
GPG key ID: 306E507219506D4E

View file

@ -1,8 +1,9 @@
# BigBoy Build Journal — Alma Deployment # BigBoy Build Journal — Alma Deployment
**Period covered:** 2026-07-16 (Thursday) through 2026-07-17 (Friday) **Period covered:** 2026-07-16 (Thursday) through 2026-07-19 (Sunday)
**Status as of the latest entry:** AlmaLinux 10.2 installed and booted **Status as of the latest entry:** Phase 1 through 5 complete. NVIDIA
successfully. Phase 2 (Ansible) not yet run. driver working, llama.cpp built and serving Ministral 3 14B Instruct,
chat confirmed working via `llama-server`'s web UI and a terminal client.
This is a working log of decisions made and why — not a step-by-step This is a working log of decisions made and why — not a step-by-step
runbook. For the actual procedures, see `planning/` and the kickstart file runbook. For the actual procedures, see `planning/` and the kickstart file
@ -21,7 +22,7 @@ It had passed through a Forgejo migration and a rename (`bigboy-alma` →
outside it in an untracked parent directory. Diffed identical; the outside it in an untracked parent directory. Diffed identical; the
untracked copy was archived rather than deleted outright, in case untracked copy was archived rather than deleted outright, in case
anything in its history mattered later. anything in its history mattered later.
- **Two copies of the kickstart file** (repo root and `planning/`) — - **Two copies of the kickstart file** (repo root and `plannng/`) —
diffed down to a single trailing-newline difference. Root copy kept as diffed down to a single trailing-newline difference. Root copy kept as
canonical. canonical.
- **A stale `README.md`** — the version in the live repo had been reduced - **A stale `README.md`** — the version in the live repo had been reduced
@ -467,3 +468,180 @@ not all of it.
real btrfs with subvolumes, not just mount what's there real btrfs with subvolumes, not just mount what's there
Next entry picks up with Phase 2. Next entry picks up with Phase 2.
---
---
# Entry 3 — 2026-07-19 (Sunday): Phase 2 through 5, and a working chat server
**Status at end of this entry:** Phase 1 through 5 all complete on real
hardware. BigBoy is a functioning AI inference server — NVIDIA driver
loaded, llama.cpp built at pinned tag `b9968`, serving Ministral 3 14B
Instruct via `llama-server`, chat confirmed working both through the
built-in web UI and a terminal client. This is the "base AI Server
Deployment" success criterion set on 2026-07-17 — met.
## Repo reconciliation before touching anything
Before running Phase 2/3, the actual live `ansible/site.yml` was checked
against what had been drafted blind (without seeing the real file). It
contained the same root-SSH assumption already known to be wrong
(`ansible_user: root`, `ansible_become: false`) — and, worse, since
play-level `vars:` takes precedence over inventory-supplied variables in
Ansible's resolution order, this would have **silently overridden** any
fix made to `inventory.ini`, without any visible contradiction. Removed
entirely; connection details now live only in `inventory.ini`. The
static `roles:` list also referenced `phase-4-nvidia-driver` through
`phase-14-observation`, none of which exist as role directories —
Ansible resolves roles at parse time, before tag filtering applies, so
this would have failed the play outright regardless of `--tags`, not
just been an unbuilt-but-harmless stub.
## The actual first successful Ansible connection
`ansible_ssh_private_key_file` pointed at `id_rsa`, but the deployed key
was ed25519 — conventionally `id_ed25519`, a different filename, not
just a wrong path. Once corrected, and once BigBoy's actual current IP
(`.241`, not the stale `.240` on record) was reflected in `inventory.ini`,
`ansible ... -m ping` returned a clean `pong` — the first successful
Ansible connection to BigBoy in this project's history.
## Phase 2/3 — three real bugs, each caught by actually running it
1. **Missing `ansible.posix` collection.** `apk add ansible-core`
(deliberately chosen over the bundled `ansible` metapackage) gives the
engine but no collections — `ansible.posix.authorized_key` and
`ansible.posix.mount` both needed `ansible-galaxy collection install
ansible.posix` before anything using them would run.
2. **`bigboy_admin_user` undefined.** The role assumed this variable
name; the actual `group_vars/bigboy.yml` used `deploy_user` instead.
Same person, different name — fixed with a one-line alias
(`bigboy_admin_user: "{{ deploy_user }}"`) rather than renaming
references throughout the role.
3. **Phase 3's `mkfs.btrfs` targeted the wrong device path entirely**
`{{ item.byid }}1` instead of `{{ item.byid }}-part1`, a bare
concatenation error rather than the correct by-id partition suffix
convention (already established the first night, during the drive
wipe). `mkfs.btrfs` failed to even open the device — no data was
touched, since it errored before writing anything.
4. **A second, related bug specific to one drive**: the ai-logs role
entry used `raw_mount: /srv/ai-logs-raw`, but the kickstart actually
created `/srv/ai-raw` (shorter, dropping "logs" — inconsistent with
the other three drives' naming, which do match their final names).
The unmount task tried to act on a path that was never mounted,
correctly reported "nothing to do," and silently left the real
mountpoint completely untouched — only caught by manually running
`lsblk -f` after the "successful" run and noticing one drive still
mounted when the other three weren't.
5. **Stale fstab entries after conversion.** The unmount step used
`state: unmounted` rather than `state: absent` — unmounting without
removing the original kickstart-generated fstab line. Left in place,
this would have meant four fstab entries pointing at xfs UUIDs that
no longer exist after `mkfs.btrfs`, a real risk of hanging at next
boot waiting for devices that will never appear. Switched to
`state: absent`, which handles both the unmount and the stale-entry
cleanup in one step.
Once all four fixes landed, the re-run was clean: all four SATA drives
converted to btrfs with zstd compression, mounted at their correct final
paths, fstab confirmed accurate, root confirmed still untouched (xfs).
## Phase 4 — smooth, genuinely
AlmaLinux's precompiled open-kmod path worked exactly as the NVIDIA
documentation described: `almalinux-release-nvidia-driver` then
`nvidia-driver-cuda nvidia-open-kmod`, one reboot, `nvidia-smi` clean on
the first attempt. The one phase tonight with zero surprises.
## Phase 5 — the CUDA toolkit gap, actually hit this time
This was flagged as a real risk before Phase 4/5 even started ("has
Phase 4 actually run, and does `nvcc --version` resolve?") and it played
out exactly as anticipated: `nvidia-driver-cuda`/`nvidia-open-kmod`
install the driver runtime, not the CUDA *toolkit*`nvcc` is a
separate package entirely. `cmake`'s configure step failed cleanly with
"CUDA Toolkit not found." Fixed by installing the plain `cuda-toolkit`
meta-package (not a specific `cuda-toolkit-13-N` sub-version, which
would risk drifting out of sync with whatever driver version actually
got installed) from the same `almalinux-nvidia` repo the driver bootstrap
had already enabled. `nvcc` itself lands in a versioned path
(`/usr/local/cuda-13.3/bin`) not automatically on `PATH` — a symlink
(`/usr/local/cuda``/usr/local/cuda-13.3`) plus a `/etc/profile.d/`
entry resolved this permanently rather than needing a manual `export`
every session.
## The OOM kill — a real, if fixable, resource gap
With `nvcc` resolved, the actual CUDA compile ran into the kernel's
OOM killer partway through — `cc1plus` and `cudafe++` both confirmed as
victims via `dmesg`. Root cause: unbounded `-j` (all 6 cores at once)
against CUDA source files, which are known to be memory-hungry to
compile due to heavy template instantiation, on a 16GB box with **zero
swap configured** — a gap that traces back to Phase 2/3 never actually
implementing the `zram_percentage: 50` variable already sitting unused
in `group_vars/bigboy.yml` since the file's original Ollama-era drafting.
Fixed pragmatically for tonight with a 16GB NVMe swapfile (not the
zram variable — that's still an open reconciliation item, noted but not
resolved) plus dropping to `-j2`. Rebuild resumed from cached objects
rather than starting over, and completed cleanly.
## First successful chat, and a genuinely capable model choice reconfirmed
`llama-server` running Ministral 3 14B Instruct at Q4_K_M, served on
`127.0.0.1:8080`, responded correctly through the browser-based built-in
UI. This is the literal success criterion set on 2026-07-17 — met,
without any tuning, exactly as scoped.
## A TUI chat client, since the bench LAN has no GUI devices
Since BigBoy's actual network segment has no monitor/browser access,
`aichat` (a Rust-based terminal client with configurable OpenAI-compatible
endpoints — confirmed current, v0.30.0, actively maintained) was
installed to talk to the same running `llama-server` instance directly,
no separate model load. This surfaced two more gaps in the same minimal-
install pattern already seen with `vim`: **`tar` and `which`** were both
missing — neither is actually part of GNU coreutils, a common
misconception, and neither was explicitly requested in the kickstart's
trimmed `%packages` list. A proactive `utility-recon.sh` script was
written to check a broader baseline of commonly-expected CLI tools in
one pass, rather than continuing to discover gaps reactively one at a
time. Run against the live system, it found nine further gaps beyond
`tar`/`which`: `bc`, `bind-utils` (`dig`), `bzip2`, `lsof`, `nmap-ncat`
(`nc`), `rsync`, `strace`, `traceroute`, `tree` — a reasonable admin/dev
baseline (archive tools, network diagnostics, process inspection,
debugging, text tools) the kickstart's deliberately trimmed package list
simply never anticipated needing. All eleven (plus `vim-enhanced`,
confirmed installable cleanly once on the live network — the earlier
vim-data mismatch was purely an installer-time repo sync issue) are now
added to the kickstart for future installs from this file.
## State at the end of this entry
- Phases 1 through 5 all complete and verified on real hardware
- BigBoy is a working AI inference server: NVIDIA driver loaded,
llama.cpp built at pinned tag `b9968`, Ministral 3 14B Instruct serving
via `llama-server`, chat confirmed via both web UI and `aichat`
- A 16GB NVMe swapfile is active (`/swapfile`, persistent via fstab) —
not the originally-planned zram approach; that variable in
`group_vars/bigboy.yml` is now a stale, unreconciled leftover worth a
deliberate decision later
- The kickstart has been updated with `tar`, `which`, and `vim-enhanced`
added explicitly, and a general-purpose utility recon script exists
for checking future installs against a broader baseline
- Root remains xfs, untouched, as planned — its eventual conversion is
still deferred pending real-world testing, per the 2026-07-17 decision
- Not yet done, deliberately: systemd unit for `llama-server` (currently
running in foreground for testing), nginx reverse proxy, firewall
opening for the proxy port, `--api-key` set to a real value — all
still on the Phase 5 task doc's list, not yet executed tonight
- The auditable llama.cpp release pipeline, per-client agent playbook
pattern, and the whole sovereignty-logging design remain exactly as
deferred in earlier entries — nothing about tonight's success changes
their priority
Next entry picks up with hardening Phase 5 (systemd, nginx, firewall) —
or whatever else surfaces first.