Final preparations pre-deploy

This commit is contained in:
John A. Hoeven 2026-07-19 13:05:08 +02:00
commit 2cdef62dce
Signed by: giovannino
GPG key ID: 306E507219506D4E
5 changed files with 635 additions and 341 deletions

View file

@ -1,35 +1,36 @@
# BigBoy AlmaLinux 10 Deployment — Ansible Playbook
Complete, idempotent Ansible playbook for deploying BigBoy inference server from minimal AlmaLinux 10.2 to fully configured AI system.
Idempotent Ansible playbook for configuring BigBoy from a fresh AlmaLinux
10.2 kickstart install through to a running llama.cpp inference server.
---
## Quick Start
### Prerequisites
- BigBoy booted with AlmaLinux 10.2 (via kickstart: `alma10-minimal-bigboy.ks`)
- SSH access from control machine (Workbench) to BigBoy (192.168.0.240)
- SSH key configured (or password auth enabled temporarily)
- Ansible 2.13+ installed on control machine
### First Run
- BigBoy installed via `alma10-minimal-bigboy.ks` (Phase 1 — done)
- SSH key deployed to BigBoy's admin account (see **SSH Key Setup** below
— required before running this playbook, not optional)
- Ansible 2.20+ on the control machine (workbench)
### First run
```bash
# From projects/bigboy-setup/ansible/
ansible-playbook site.yml -i inventory.ini -v
# From ansible/
ansible-playbook site.yml -i inventory.ini --tags phase2,phase3
```
### Run Specific Phase (if earlier phase fails)
```bash
# Re-run only Phase 3 (filesystems)
ansible-playbook site.yml -i inventory.ini --tags phase-3
### Re-running a specific phase
# Or specific phase
ansible-playbook site.yml -i inventory.ini --tags phase-4
```bash
ansible-playbook site.yml -i inventory.ini --tags phase3
```
### Check Mode (show what would change)
### Check mode (show what would change, without changing anything)
```bash
ansible-playbook site.yml -i inventory.ini --check
ansible-playbook site.yml -i inventory.ini --check --tags phase2,phase3
```
---
@ -38,282 +39,278 @@ ansible-playbook site.yml -i inventory.ini --check
```
ansible/
├── site.yml # Main playbook (orchestrates all 14 phases)
├── site.yml # Main playbook
├── inventory.ini # Hosts + SSH configuration
├── keys/
│ └── workbench.pub # gitignored — deploy locally, never commit
├── group_vars/
│ └── bigboy.yml # Hardware-specific variables (UUIDs, IPs, etc.)
│ └── bigboy.yml # Hardware facts: UUIDs, by-id paths, pinned versions
├── roles/
│ ├── phase2-system-config/
│ │ └── tasks/main.yml # Sudoers policy, SSH key, base checks
│ ├── phase-3-filesystems/
│ │ └── tasks/main.yml # Validate + mount all 4 data drives
│ ├── phase-4-nvidia-driver/ # (scaffolding ready; add tasks/)
│ ├── phase-5-ollama/ # (scaffolding ready; add tasks/)
│ └── ... (phases 6-14)
│ │ └── tasks/main.yml # Convert 4 SATA drives xfs -> btrfs
│ ├── phase-4-nvidia-driver/ # scaffold — see tasks/TASK-phase4-5-bringup.md
│ ├── phase-5-llama-cpp/ # scaffold — see tasks/TASK-phase4-5-bringup.md
│ └── ... (phases 6-14, not yet scoped)
└── README.md # This file
```
---
## 14 Phases (Phase Breakdown)
### Phase 3: Filesystem Validation ✓ (Complete)
- Mount all 4 data drives (/srv/rag-library, /srv/prompt-library, /srv/backup, /srv/ai-logs)
- Verify btrfs subvolume structure
- Log filesystem space
- **Status:** Idempotent, ready to test
### Phase 4: NVIDIA GPU Driver
- Install kernel-headers and build essentials
- Enable CRB/EPEL repos
- Install NVIDIA open kernel modules (Precompiled, not DKMS)
- Verify nvidia-smi
- **Status:** Ready to build (scaffold exists)
### Phase 5: Ollama Installation
- Install Ollama from official package
- Configure environment (CUDA, GPU selection, VRAM limits)
- Start/enable ollama service
- Pull test model (Mistral)
- **Status:** Ready to build
### Phase 6: Build Suite
- Install development tools (gcc, make, git, tmux, vim, etc.)
- Install system utilities (btrfs-progs, smartmontools, nvtop)
- **Status:** Ready to build
### Phase 7: Configuration
- Deploy dotfiles (tmux.conf, vimrc, bash profile)
- Set system locale/timezone
- Configure shell environment
- **Status:** Ready to build
### Phase 8: oterm (TUI Ollama Client)
- Install oterm from source or package
- Configure for local Ollama connection
- Test TUI interface
- **Status:** Ready to build
### Phase 9: Security Hardening
- Configure firewalld (open SSH 22, Open WebUI 8080, Cockpit 9090; keep Ollama 11434 localhost-only)
- SSH hardening (disable password auth, PermitRootLogin=no)
- MAC address pinning for enp4s0
- **Status:** Ready to build
### Phase 10: Borgmatic Backups (Optional)
- Install borgbackup + borgmatic
- Configure backup schedule, passphrase, retention
- **(Deferred:** backup target not yet decided)
- **Status:** Scaffold ready; design deferred
### Phase 11: Thermal Baseline Testing
- Run memtest86 stress test
- Log CPU/GPU temps, fan speed
- Record baseline performance
- **Status:** Ready to build
### Phase 12: Full System Validation
- Validate all previous phases
- Test GPU, Ollama, network, storage
- Generate validation report
- **Status:** Ready to build
### Phase 13: Home LAN Migration Prep
- (Deferred until case installed)
- Static IP assignment
- DNS configuration
- **Status:** Deferred
### Phase 14: Observation Period Runbook
- Daily/weekly health checks
- Monitor temps, disk usage, service status
- **Status:** Runbook template ready to build
Scoped work-in-progress task handoffs live in `../tasks/` at repo root, not
here — `TASK-phase2-3.md` and `TASK-phase4-5-bringup.md` are the
authoritative descriptions of what each phase actually does; this README
is a summary and operational reference, not the source of truth for scope.
---
## Variables (group_vars/bigboy.yml)
## Phase Breakdown
All hardware-specific settings live in one place:
### Phase 1: Base Install — done
Unattended AlmaLinux 10.2 install via kickstart. See
`../alma10-minimal-bigboy.ks`.
### Phase 2: System Configuration — see `tasks/TASK-phase2-3.md`
- Replace the kickstart's temporary bootstrap NOPASSWD sudoers grant with
a real, deliberately-scoped policy
- Deploy workbench's SSH public key (single key only at this stage —
BigBoy sits on the isolated bench LAN; more keys get added later, one
at a time, once on the home LAN)
- Confirm hostname, timezone, NTP/chrony, and repo configuration
**Does not** disable password SSH authentication — that's Phase 9, once
key-based login is confirmed working over real use, not assumed the
moment a key is deployed.
### Phase 3: Filesystem Conversion — see `tasks/TASK-phase2-3.md`
The kickstart formats all five drives as **xfs** (RHEL10's Anaconda
removed btrfs support entirely — it's not a choice, it's what's actually
possible at install time). Phase 3 converts the four SATA data drives to
real btrfs with subvolumes and zstd compression:
| Kickstart mount | Drive | Final path |
|---|---|---|
| `/srv/backup-raw` | BACKUPDRV (Lexar 256GB) | `/srv/backup` |
| `/srv/rag-raw` | RAGLIB (Samsung EVO 500GB) | `/srv/rag-library` |
| `/srv/ai-raw` | AILOGS (Samsung EVO 250GB) | `/srv/ai-logs` |
| `/srv/prompt-raw` | PROMPTLIB (Lexar 256GB) | `/srv/prompt-library` |
**Root stays xfs, permanently for now.** There's no supported in-place
xfs-to-btrfs conversion — this would require a full reinstall, planned
separately and deliberately deferred until after real-world testing (not
before ~a month out). Do not attempt to touch root as part of this phase.
### Phase 4: NVIDIA Driver — see `tasks/TASK-phase4-5-bringup.md`
AlmaLinux's precompiled open-kmod path (`almalinux-release-nvidia-driver`
`nvidia-driver-cuda nvidia-open-kmod`), not manual DKMS. Confirmed as
the AlmaLinux-recommended method, Secure Boot compatible, needs none of
the kernel-devel/DKMS/compiler machinery DKMS would.
### Phase 5: llama.cpp — see `tasks/TASK-phase4-5-bringup.md`
**Not Ollama.** Built from source, pinned to a specific tag (`b9968`,
recorded in `group_vars/bigboy.yml`), served via `llama-server` under
systemd behind nginx. Initial model: Ministral 3 14B Instruct, Q4_K_M —
success criterion is a coherent chat response through the built-in web
UI, quality as-delivered, no tuning required at this stage.
### Phases 614: not yet scoped
Deferred until 4/5 are stable and tested. No committed design for these
yet — don't assume the shape of a future phase from an old draft.
---
## Variables (`group_vars/bigboy.yml`)
```yaml
# NVMe UUIDs
uuid_nvme_root: "5daac1d7-10b3-498a-82b0-a4498d7e0717"
# NVMe (OS + root, stays xfs)
drive_nvme_byid: "nvme-Samsung_SSD_980_500GB_S64DNL0TC07019Z"
# Data drive UUIDs
uuid_rag_library: "18b9accd-754a-46f3-b994-da3c7ae795cd"
uuid_prompt_library: "82a240c4-390a-4167-8232-6a04ce4d84bb"
uuid_backup: "15b69400-f1f7-4cfd-82cb-4d1244951503"
uuid_ai_logs: "1e57a52a-9c9d-44ef-a352-3cc542808d13"
# SATA data drives — by-id, NOT sdX (unstable across boots/reorders,
# already proven true once during this deployment)
drive_backup_byid: "ata-Lexar_SSD_NS100_256GB_QM9370R0291730S340"
drive_rag_byid: "ata-Samsung_SSD_870_EVO_500GB_S6PYNL0T620757H"
drive_ai_logs_byid: "ata-Samsung_SSD_870_EVO_250GB_S61WNJ0R458541M"
drive_prompt_byid: "ata-Lexar_SSD_NS100_256GB_QM9370R0290370S340"
# GPU settings
nvidia_driver_version: "595.84"
cuda_visible_devices: "0"
# Admin account (matches kickstart's user --name=)
bigboy_admin_user: "<set to actual username>"
# Ollama tuning
ollama_max_loaded_models: 1
ollama_keep_alive: "5m"
ollama_gpu_overhead: 536870912 # 512MB
# GPU / driver — do NOT treat this as a version to install for; the
# precompiled path installs whatever AlmaLinux's repo currently ships.
# Record the ACTUAL version here after Phase 4 runs, replacing this
# placeholder.
nvidia_driver_version: "<record actual version after Phase 4>"
# llama.cpp — pinned, not tracking master
llama_cpp_tag: "b9968"
# Model — confirmed choice for initial bring-up, not a placeholder
model_repo: "mistralai/Ministral-3-14B-Instruct-2512-GGUF"
model_quant: "Q4_K_M"
```
**To update:** Edit `group_vars/bigboy.yml`, then re-run playbook. Variables propagate to all roles.
**Note on the four SATA drives' filesystem UUIDs**: unlike the by-id
paths above (stable, known ahead of time), the *filesystem* UUIDs for
these drives don't exist yet — they're generated fresh when Phase 3 runs
`mkfs.btrfs`. Don't hardcode filesystem UUIDs into this file preemptively.
Phase 3's role discovers each UUID at runtime (via `blkid` against the
by-id path, immediately after formatting) and uses that discovered value
to write `/etc/fstab` — see the role itself for the exact mechanism.
**To update:** edit this file, re-run the playbook. Variables propagate
to every role.
---
## Idempotency
Every task is idempotent (safe to re-run):
- Mounts already present = no change
- Packages already installed = no change
- Services already running = no change
- Shell commands wrapped with `changed_when` to report accurately
Every task should be safe to re-run:
- Already-converted drives aren't re-wiped (Phase 3 checks existing
filesystem type before acting)
- Already-deployed keys aren't duplicated
- Already-correct sudoers policy isn't rewritten unnecessarily
**Key principle:** Running the playbook twice produces the same result as running it once.
**Principle**: running the playbook twice produces the same result as
running it once. If a task isn't naturally idempotent, wrap it with an
explicit check (`stat`, `command` + `changed_when`) rather than assuming
Ansible's module-level idempotency covers everything.
---
## Logging
All output is logged to `/srv/deployment-log/` on BigBoy:
Phase output is logged to `/srv/deployment-log/` on BigBoy (created by
the kickstart's `%post`), one file per phase run:
```
/srv/deployment-log/
├── phase-03-filesystems-2026-06-27.log
├── phase-04-nvidia-driver-2026-06-27.log
├── phase-05-ollama-2026-06-27.log
└── ... (one per phase)
├── kickstart.log # from Phase 1's %post
├── phase2-system-config-<date>.log
├── phase3-filesystems-<date>.log
└── ...
```
Each log includes:
- Timestamp of each task
- Module output
- Failure diagnosis (if applicable)
**To view:** `ssh root@192.168.0.240 "tail -f /srv/deployment-log/*.log"`
```bash
ssh <admin-user>@192.168.0.240 "tail -f /srv/deployment-log/*.log"
```
---
## Troubleshooting
### Phase fails mid-run
1. Check the specific phase log: `tail /srv/deployment-log/phase-N-*.log`
2. Fix the issue manually if needed
3. Re-run the phase: `ansible-playbook site.yml --tags phase-N`
1. Check the phase-specific log under `/srv/deployment-log/`
2. Fix manually if needed, or adjust the role
3. Re-run just that phase: `ansible-playbook site.yml --tags phaseN`
### SSH connection fails
1. Verify BigBoy IP: `ssh -v root@192.168.0.240`
2. Check SSH key permissions: `chmod 600 ~/.ssh/id_rsa`
3. Ensure root SSH login is enabled on BigBoy
### Idempotency broken (task reports change every time)
- Check `changed_when` / `failed_when` directives
- Verify the conditional logic
- Use `-vv` for detailed task output
1. Confirm BigBoy's current IP — `192.168.0.240` on the bench LAN as of
this writing, but confirm via `ansible/inventory.ini` rather than
assuming it hasn't changed
2. Confirm you're connecting as the admin user, **not root** — root login
is locked (`rootpw --lock` in the kickstart), always was, this isn't
something Phase 9 disables later
3. Confirm the key at `keys/workbench.pub` locally matches what's
actually deployed to BigBoy's `authorized_keys`
### Idempotency broken (task reports "changed" every run)
- Check `changed_when`/`failed_when` logic in the specific task
- Run with `-vv` for full task output
---
## Extending (Adding Phases)
## Extending (Adding a New Phase)
To add Phase 4 (NVIDIA driver):
1. Create directory:
```bash
mkdir -p roles/phase-4-nvidia-driver/tasks
```
2. Create `tasks/main.yml` with steps (use alma-nvidia-driver-installation.txt as reference)
3. Add role to `site.yml`:
```yaml
- role: phase-4-nvidia-driver
tags: [phase-4, gpu, nvidia]
```
4. Run playbook:
```bash
ansible-playbook site.yml --tags phase-4
```
**Pattern:** Each phase = one role = idempotent, re-runnable, logged.
1. `mkdir -p roles/phaseN-<name>/tasks`
2. Write `tasks/main.yml`
3. **First, write the task doc** at `../tasks/TASK-phaseN-<name>.md`
scoping exactly what the phase does and doesn't do — per this
project's established workflow (Claude Desktop plans/scopes, Claude
Code implements against the scoped task, not the other way around)
4. Add the role to `site.yml` with an appropriate tag
5. Run: `ansible-playbook site.yml --tags phaseN`
---
## SSH Key Setup (Post-Install)
## SSH Key Setup (required before running this playbook)
BigBoy ships with password authentication. To switch to key-based:
BigBoy ships from the kickstart with **password authentication only**
root is locked, and no key exists yet. This is the actual chicken-and-egg
this playbook's Phase 2 resolves, but the *first* key push has to happen
manually, using password auth, before Ansible can connect via key:
```bash
# 1. On Workbench, generate key (if not already done)
ssh-keygen -t ed25519 -f ~/.ssh/id_rsa -N ""
# On workbench — the key already exists at ansible/keys/workbench.pub
# (gitignored, never commit the actual key material)
ssh-copy-id -i ~/.ssh/<your-key>.pub <admin-user>@192.168.0.240
# 2. Copy key to BigBoy (will be automated in Phase 9)
ssh-copy-id -i ~/.ssh/id_rsa.pub root@192.168.0.240
# 3. Verify key auth works
ssh -i ~/.ssh/id_rsa root@192.168.0.240 "echo 'Connected'"
# 4. Phase 9 will disable password auth once keys are in place
# Verify
ssh -i ~/.ssh/<your-key> <admin-user>@192.168.0.240 "echo connected"
```
Then update `inventory.ini` to use `ansible_user=<admin-user>` (not
`root`) and point `ansible_ssh_private_key_file` at the correct key.
**Password authentication stays enabled** even after this — disabling it
is explicitly Phase 9's job, not something to do the moment a key works
once.
---
## Validation Checklist (Post-Deployment)
After all phases complete:
## Validation Checklist (Phase 2/3)
```bash
# SSH to BigBoy
ssh root@192.168.0.240
ssh <admin-user>@192.168.0.240
# Check filesystems
# Sudoers — bootstrap grant gone, real policy in place
sudo -l
ls /etc/sudoers.d/
# Filesystems
df -h /srv/*
mount | grep btrfs
# Check GPU
nvidia-smi
# Check Ollama
ollama --version
systemctl status ollama
# Check services
systemctl status firewalld
systemctl status sshd
# Check logs
tail -f /srv/deployment-log/*.log
# Root untouched
mount | grep ' / ' # should show xfs, not btrfs
```
---
## RAG Integration
## Known Deferred (not gaps — deliberate)
All phase logs are automatically captured to `/srv/deployment-log/` and ready for RAG indexing:
- Workbench cron harvests logs nightly
- Failures documented (not just successes)
- Workarounds captured for future reference
---
## Known Limitations / Deferred
- **Phase 10 (Borgmatic):** Backup target not yet decided; packages installed, schedule deferred
- **Phase 13 (Home LAN):** Deferred until Modcase EVO ITX-2 case installed and system is cased
- **Phase 14 (Observation):** Runbook template only; manual health checks during first month
- **Root filesystem conversion to btrfs** — needs a full reinstall
(clone/reformat/restore via `debusb`), planned for ~a month out after
real-world testing, not part of this playbook
- **Additional SSH keys** (Surface, tinkerpad) — one at a time, only
once BigBoy is reachable on the home LAN
- **Dedicated display-monitor user** — for the case-mounted 5" screen
relocated from jahnet; needs autologin + `btop`-only access, not yet
built
- **Auditable llama.cpp release pipeline** and **per-client agent
playbooks** — both designed in detail, neither built; see `CLAUDE.md`
- **Phases 614** — not yet scoped at all
---
## Reference Files
- `alma10-minimal-bigboy.ks` — Kickstart for unattended OS install
- `/home/john/documents/library/rag/use-case/ansible-*.md` — Ansible best practices, modules, error handling
- `/home/john/documents/raw-docs/alma-nvidia-driver-installation.txt` — Official NVIDIA guide (reference)
- `../alma10-minimal-bigboy.ks` — kickstart for the unattended OS install
- `../CLAUDE.md` — standing technical decisions, read automatically by
Claude Code every session
- `../tasks/TASK-phase2-3.md`, `../tasks/TASK-phase4-5-bringup.md`
authoritative scope for each phase
- `../planning/STATUS.md` — authoritative current per-phase state
---
## Support / Issues
- **Logs:** Check `/srv/deployment-log/phase-N-*.log` first
- **Ansible:** Run with `-vvv` for full debug output
- **Hardware:** Verify UUIDs in `group_vars/bigboy.yml` match actual system
---
**Last Updated:** 2026-06-27
**Status:** Phase 3 complete and tested; scaffolding ready for Phases 4-14
**Last updated:** 2026-07-19
**Status:** Phase 1 done (real hardware). Phase 2/3 roles being written
against scoped task docs. Phase 4/5 scoped, not yet run.

View file

@ -0,0 +1,105 @@
---
# Phase 2: System Configuration
# See ../../../tasks/TASK-phase2-3.md for full scope and reasoning.
#
# Runs using the kickstart's temporary NOPASSWD bootstrap grant
# (/etc/sudoers.d/00-bootstrap-<user>) — this task's own job is to
# replace that grant with something deliberate, not to assume it's
# permanent.
- name: Confirm bootstrap sudoers file exists (sanity check before removing it)
ansible.builtin.stat:
path: "/etc/sudoers.d/00-bootstrap-{{ bigboy_admin_user }}"
register: bootstrap_sudoers
- name: Deploy real sudoers policy (NOPASSWD retained — Ansible needs unattended escalation for future phases)
ansible.builtin.copy:
dest: "/etc/sudoers.d/10-admin"
content: |
# Real, deliberate sudoers policy for {{ bigboy_admin_user }}.
# Supersedes the kickstart's temporary bootstrap grant
# (00-bootstrap-{{ bigboy_admin_user }}), which this same task
# removes below.
#
# NOPASSWD retained deliberately: Ansible needs unattended
# privilege escalation to run subsequent phases without
# prompting. This is a considered choice, not a leftover — if a
# scoped-down policy (password-required, or command-restricted)
# is wanted instead, that's a decision to revisit explicitly,
# not something this task silently assumed.
{{ bigboy_admin_user }} ALL=(ALL) NOPASSWD: ALL
owner: root
group: root
mode: "0440"
validate: "visudo -cf %s"
- name: Remove the kickstart's temporary bootstrap sudoers grant
ansible.builtin.file:
path: "/etc/sudoers.d/00-bootstrap-{{ bigboy_admin_user }}"
state: absent
when: bootstrap_sudoers.stat.exists
# ---------------------------------------------------------------------------
# SSH key deployment — single key only, this phase (see task doc for why:
# BigBoy sits on the isolated bench LAN right now, no other devices present
# to test additional keys against yet).
# ---------------------------------------------------------------------------
- name: Deploy workbench's SSH public key
ansible.posix.authorized_key:
user: "{{ bigboy_admin_user }}"
state: present
key: "{{ lookup('file', 'keys/workbench.pub') }}"
# Explicitly not disabling PasswordAuthentication here — that's Phase 9's
# job, once key-based login is confirmed working over real, repeated use,
# not assumed the moment this task deploys a key.
# ---------------------------------------------------------------------------
# Base system confirmation
# ---------------------------------------------------------------------------
- name: Confirm hostname
ansible.builtin.command: hostname
register: hostname_check
changed_when: false
- name: Fail if hostname doesn't match expected
ansible.builtin.fail:
msg: "Hostname is '{{ hostname_check.stdout }}', expected 'bigboy'"
when: hostname_check.stdout != 'bigboy'
- name: Confirm timezone
ansible.builtin.command: timedatectl show --property=Timezone --value
register: timezone_check
changed_when: false
- name: Fail if timezone doesn't match expected
ansible.builtin.fail:
msg: "Timezone is '{{ timezone_check.stdout }}', expected 'Europe/Rome'"
when: timezone_check.stdout != 'Europe/Rome'
- name: Confirm chrony is active (NTP sync)
ansible.builtin.systemd:
name: chronyd
register: chrony_status
- name: Fail if chrony isn't running
ansible.builtin.fail:
msg: "chronyd is not active — NTP sync not confirmed"
when: chrony_status.status.ActiveState != 'active'
- name: Confirm all four expected repos are enabled
ansible.builtin.command: dnf repolist enabled
register: repolist_check
changed_when: false
- name: Fail if any expected repo is missing
ansible.builtin.fail:
msg: "Expected repo '{{ item }}' not found in enabled repolist"
loop:
- baseos
- appstream
- crb
- extras
when: item not in repolist_check.stdout

View file

@ -1,82 +1,151 @@
---
# Phase 3: Filesystem Validation
# Validates all 4 data drive mounts + NVMe structure
# Idempotent: can re-run safely; mounts already present = no change
# Phase 3: Filesystem Conversion (four SATA drives only)
# See ../../../tasks/TASK-phase2-3.md for full scope and reasoning.
#
# IMPORTANT: this replaces an earlier version of this role that assumed
# mounting *existing* btrfs volumes. That assumption no longer holds —
# RHEL10's kickstart cannot create btrfs at install time, so all four
# drives currently exist as xfs (per the kickstart's %packages/part
# commands). This role converts them: wipe, mkfs.btrfs, mount, fstab.
#
# Root is explicitly untouched by this role. There is no supported
# in-place xfs-to-btrfs conversion; root's eventual conversion is a
# separate, deferred full-reinstall task (clone/reformat/restore via
# debusb), not part of this playbook.
#
# Idempotent: checks each drive's actual current filesystem before
# acting, so re-running this role doesn't re-wipe an already-converted
# drive.
- name: "Log Phase 3 start"
ansible.builtin.lineinfile:
path: "{{ deployment_log_dir }}/phase-03-filesystems-{{ ansible_date_time.date }}.log"
create: yes
line: "[{{ ansible_date_time.iso8601 }}] Phase 3: Filesystem Validation starting on {{ inventory_hostname }}"
mode: "0644"
- name: Define the four SATA drives to convert
ansible.builtin.set_fact:
sata_drives:
- name: backup
byid: "{{ drive_backup_byid }}"
raw_mount: /srv/backup-raw
final_mount: /srv/backup
- name: rag-library
byid: "{{ drive_rag_byid }}"
raw_mount: /srv/rag-raw
final_mount: /srv/rag-library
- name: ai-logs
byid: "{{ drive_ai_logs_byid }}"
raw_mount: /srv/ai-logs-raw
final_mount: /srv/ai-logs
- name: prompt-library
byid: "{{ drive_prompt_byid }}"
raw_mount: /srv/prompt-raw
final_mount: /srv/prompt-library
- name: "Validate NVMe root filesystem is btrfs"
ansible.builtin.command:
cmd: "blkid -s TYPE -o value {{ nvme_root_partition }}"
register: nvme_fstype
changed_when: false
failed_when: "'btrfs' not in nvme_fstype.stdout"
- name: "Create /srv mount point"
ansible.builtin.file:
path: /srv
state: directory
mode: "0755"
owner: root
group: root
- name: "Mount all data drives (NVMe subvolumes)"
ansible.posix.mount:
path: "/{{ item.mount | basename }}"
src: "UUID={{ item.uuid }}"
fstype: btrfs
opts: "subvol=@{{ item.name }},{{ btrfs_mount_options }}"
state: mounted
loop: "{{ data_drives }}"
register: mount_results
- name: "Verify all 4 data drives are mounted"
ansible.builtin.command:
cmd: "mountpoint -q {{ item.mount }}"
loop: "{{ data_drives }}"
changed_when: false
- name: "Check filesystem space on data drives"
ansible.builtin.command:
cmd: "df -h {{ item.mount }}"
register: df_results
loop: "{{ data_drives }}"
changed_when: false
- name: "Log filesystem space"
ansible.builtin.lineinfile:
path: "{{ deployment_log_dir }}/phase-03-filesystems-{{ ansible_date_time.date }}.log"
line: "{{ item.cmd }}"
state: present
loop: "{{ df_results.results }}"
- name: "Verify btrfs subvolume structure on data drives"
ansible.builtin.shell:
cmd: "btrfs subvolume list {{ item.mount }} | grep @{{ item.name }}"
register: subvol_check
loop: "{{ data_drives }}"
- name: Check current filesystem type on each drive
ansible.builtin.command: "blkid -o value -s TYPE /dev/disk/by-id/{{ item.byid }}1"
register: current_fstype
changed_when: false
failed_when: false
loop: "{{ sata_drives }}"
loop_control:
label: "{{ item.name }}"
- name: "Log Phase 3 completion"
ansible.builtin.lineinfile:
path: "{{ deployment_log_dir }}/phase-03-filesystems-{{ ansible_date_time.date }}.log"
line: "[{{ ansible_date_time.iso8601 }}] Phase 3: SUCCESS - All 4 data drives mounted and verified"
- name: Build per-drive conversion status
ansible.builtin.set_fact:
drives_to_convert: >-
{{ sata_drives | zip(current_fstype.results)
| selectattr('1.stdout', 'ne', 'btrfs')
| map(attribute='0') | list }}
- name: "Display filesystem summary"
- name: Report drives already converted (skipped)
ansible.builtin.debug:
msg: |
========================================
Phase 3: Filesystem Validation Complete
========================================
{% for drive in data_drives %}
{{ drive.mount }}: {{ drive.uuid }}
{% endfor %}
msg: "{{ item.0.name }} already btrfs — skipping conversion, will confirm mount only"
loop: "{{ sata_drives | zip(current_fstype.results) | list }}"
loop_control:
label: "{{ item.0.name }}"
when: item.1.stdout == 'btrfs'
Full logs: {{ deployment_log_dir }}/phase-03-filesystems-{{ ansible_date_time.date }}.log
========================================
# ---------------------------------------------------------------------------
# Conversion — only runs against drives not already btrfs
# ---------------------------------------------------------------------------
- name: Unmount existing xfs partition before conversion
ansible.posix.mount:
path: "{{ item.raw_mount }}"
state: unmounted
loop: "{{ drives_to_convert }}"
loop_control:
label: "{{ item.name }}"
- name: Format as btrfs with zstd compression
ansible.builtin.command: >
mkfs.btrfs -f
/dev/disk/by-id/{{ item.byid }}1
loop: "{{ drives_to_convert }}"
loop_control:
label: "{{ item.name }}"
- name: Discover the fresh filesystem UUID for each newly-converted drive
ansible.builtin.command: "blkid -o value -s UUID /dev/disk/by-id/{{ item.byid }}1"
register: new_uuids
changed_when: false
loop: "{{ drives_to_convert }}"
loop_control:
label: "{{ item.name }}"
- name: Create final mount point directories
ansible.builtin.file:
path: "{{ item.final_mount }}"
state: directory
owner: root
group: root
mode: "0755"
loop: "{{ drives_to_convert }}"
loop_control:
label: "{{ item.name }}"
- name: Remove stale -raw mount point directories (superseded by final_mount)
ansible.builtin.file:
path: "{{ item.raw_mount }}"
state: absent
loop: "{{ drives_to_convert }}"
loop_control:
label: "{{ item.name }}"
- name: Mount each converted drive at its final path with zstd + noatime
ansible.posix.mount:
path: "{{ item.0.final_mount }}"
src: "UUID={{ item.1.stdout }}"
fstype: btrfs
opts: "compress=zstd,noatime"
state: mounted
loop: "{{ drives_to_convert | zip(new_uuids.results) | list }}"
loop_control:
label: "{{ item.0.name }}"
# ---------------------------------------------------------------------------
# Confirm final state for ALL four drives, whether just-converted or
# already-converted on a prior run
# ---------------------------------------------------------------------------
- name: Confirm all four drives are mounted as btrfs at final paths
ansible.builtin.command: "findmnt -n -o FSTYPE {{ item.final_mount }}"
register: final_fstype_check
changed_when: false
loop: "{{ sata_drives }}"
loop_control:
label: "{{ item.name }}"
- name: Fail if any drive isn't btrfs at its final mount point
ansible.builtin.fail:
msg: "{{ item.0.name }} at {{ item.0.final_mount }} is '{{ item.1.stdout }}', expected 'btrfs'"
loop: "{{ sata_drives | zip(final_fstype_check.results) | list }}"
loop_control:
label: "{{ item.0.name }}"
when: item.1.stdout != 'btrfs'
- name: Confirm root is untouched — still xfs
ansible.builtin.command: "findmnt -n -o FSTYPE /"
register: root_fstype_check
changed_when: false
- name: Fail loudly if root is somehow not xfs (should never happen from this role)
ansible.builtin.fail:
msg: "Root filesystem is '{{ root_fstype_check.stdout }}', expected 'xfs' — this role should never have touched root"
when: root_fstype_check.stdout != 'xfs'

View file

@ -1,21 +1,24 @@
---
# BigBoy AlmaLinux 10 Deployment Playbook
# 14-phase unattended deployment from minimal Alma 10.2 to full AI inference server
# Execution: ansible-playbook site.yml -i inventory.ini
# Execution: ansible-playbook site.yml -i inventory.ini --tags phase-2,phase-3
# Or single phase: ansible-playbook site.yml -i inventory.ini --tags phase-3
#
# NOTE on connection settings: ansible_user, ansible_ssh_private_key_file,
# etc. belong in inventory.ini, NOT here. An earlier version of this file
# set `ansible_user: root` / `ansible_become: false` directly in play
# vars — since play-level vars take precedence over inventory vars in
# Ansible's variable resolution order, that would silently override any
# fix made in inventory.ini. Root SSH doesn't work on BigBoy at all
# (rootpw --lock in the kickstart) — connection details are inventory's
# job alone.
- name: "BigBoy AlmaLinux 10 Deployment"
hosts: bigboy
become: true
gather_facts: yes
vars_files:
- group_vars/bigboy.yml
vars:
# Ansible execution defaults
ansible_connection: ssh
ansible_user: root
ansible_become: false # Already running as root
pre_tasks:
- name: "Log deployment start"
ansible.builtin.lineinfile:
@ -33,53 +36,26 @@
fail_msg: "This playbook requires AlmaLinux 10.x (detected: {{ ansible_distribution }} {{ ansible_distribution_version }})"
roles:
# Phase 3: Filesystem Validation
# Phase 2: System Configuration
- role: phase-2-system-config
tags: [phase-2, system-config, required]
# Phase 3: Filesystem Conversion (four SATA drives, xfs -> btrfs)
- role: phase-3-filesystems
tags: [phase-3, filesystems, required]
# Phase 4: NVIDIA GPU Driver Installation
- role: phase-4-nvidia-driver
tags: [phase-4, gpu, nvidia, required]
# Phase 5: Ollama Installation and Configuration
- role: phase-5-ollama
tags: [phase-5, ollama, inference]
# Phase 6: Build Suite and Development Tools
- role: phase-6-build-suite
tags: [phase-6, buildtools, development]
# Phase 7: Configuration (dotfiles, vim, tmux, shell)
- role: phase-7-configuration
tags: [phase-7, config, dotfiles]
# Phase 8: oterm Installation (TUI Ollama client)
- role: phase-8-oterm
tags: [phase-8, oterm, tui]
# Phase 9: Security (firewall, SSH hardening, MAC pinning)
- role: phase-9-security
tags: [phase-9, security, firewall]
# Phase 10: Borgmatic Backup Configuration (optional)
- role: phase-10-borgmatic
tags: [phase-10, backup, optional]
# Phase 11: Thermal Baseline Testing
- role: phase-11-thermal
tags: [phase-11, thermal, testing]
# Phase 12: Full System Validation
- role: phase-12-validation
tags: [phase-12, validation, final-check]
# Phase 13: Home LAN Migration Prep (deferred until case installed)
- role: phase-13-home-llan-prep
tags: [phase-13, network, deferred]
# Phase 14: Observation Period Checklist
- role: phase-14-observation
tags: [phase-14, observation, runbook]
# Phase 4/5 roles not yet written — see ../tasks/TASK-phase4-5-bringup.md
# for the fully scoped task. These are commented out deliberately, not
# just stylistically: a static roles: list entry pointing at a
# nonexistent role directory fails at parse time, before tag
# filtering even applies — so this isn't optional tidiness, the play
# won't run at all with these active and unbuilt.
#
# - role: phase-4-nvidia-driver
# tags: [phase-4, gpu, nvidia, required]
#
# - role: phase-5-llama-cpp
# tags: [phase-5, llama-cpp, inference]
post_tasks:
- name: "Log deployment completion"
@ -99,11 +75,13 @@
Deployment Log: {{ deployment_log_dir }}/deployment.log
Next Steps:
1. Verify all filesystems: df -h /srv/*
2. Test GPU: nvidia-smi
3. Test Ollama: ollama --version
4. Review logs: tail -f {{ deployment_log_dir }}/*.log
1. Verify filesystems: df -h /srv/*
2. Confirm root is still xfs: findmnt -n -o FSTYPE /
3. (Once Phase 4/5 run) Test GPU: nvidia-smi
4. (Once Phase 4/5 run) Test llama-server:
systemctl status llama-server
5. Review logs: tail -f {{ deployment_log_dir }}/*.log
For detailed logs from each phase, see:
{{ deployment_log_dir }}/phase-*.log
========================================
========================================