summaryrefslogtreecommitdiff
path: root/prompts/skills/f3s-storage
diff options
context:
space:
mode:
authorPaul Buetow <paul@buetow.org>2026-07-18 22:03:55 +0300
committerPaul Buetow <paul@buetow.org>2026-07-18 22:03:55 +0300
commitfbf9d4815ea25bd85be0f232ac8670bda1a64955 (patch)
tree5cdf0f6d2ce069f2682b7abeb53da660e27c1b39 /prompts/skills/f3s-storage
parentcee544948612841c8c7831b6de6776b288ce9d17 (diff)
skills: split f3s into six focused sibling skills (index pattern)
Execute plans/f3s-skill-split-plan.md. Carve the oversized f3s skill into six f3s-prefixed siblings so each loads on its own triggers, and slim f3s to a hub (38 reference files -> 10; 119 -> 73 lines): - f3s-storage ZFS/zrepl/CARP/NFS-stunnel, nfs-mount-monitor, backups - f3s-k3s cluster install, off-LAN access, ingress, etcd, r-node Rex - f3s-observability Prometheus/Alloy/Loki/Tempo, FreeBSD node_exporter - f3s-workloads Immich, Garage, Player, yChat, goprecords/uptimed - f3s-raspberry-pi pi0/pi1 NetBSD static site, pi2/pi3 Pi-hole/LAN DNS - f3s-dtail dserver deployment/ops (SSH 2222) f3s hub keeps the master Host-IP table, physical hosts, bhyve, WireGuard mesh, and off-LAN access as the canonical inward-pointing context, plus a Related skills block. Applies skill-maintenance best practices: fixes the three inlined SKILL.md duplications (Pi/webserver, DTail) by moving prose to one canonical home and keeping each new SKILL.md a slim index that points to its references rather than re-inlining them. All cross-skill links rewritten to ../../<skill>/references/ form (incl. inbound links from pkgrepo and rocky-vm-setup, and two pre-existing broken links); verified all 112 relative links resolve. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Diffstat (limited to 'prompts/skills/f3s-storage')
-rw-r--r--prompts/skills/f3s-storage/SKILL.md41
-rw-r--r--prompts/skills/f3s-storage/references/backups.md40
-rw-r--r--prompts/skills/f3s-storage/references/carp.md95
-rw-r--r--prompts/skills/f3s-storage/references/nfs-mount-monitor.md107
-rw-r--r--prompts/skills/f3s-storage/references/nfs.md204
-rw-r--r--prompts/skills/f3s-storage/references/troubleshooting.md198
-rw-r--r--prompts/skills/f3s-storage/references/usb-keys.md116
-rw-r--r--prompts/skills/f3s-storage/references/zfs.md90
-rw-r--r--prompts/skills/f3s-storage/references/zrepl.md237
9 files changed, 1128 insertions, 0 deletions
diff --git a/prompts/skills/f3s-storage/SKILL.md b/prompts/skills/f3s-storage/SKILL.md
new file mode 100644
index 0000000..82c0a5e
--- /dev/null
+++ b/prompts/skills/f3s-storage/SKILL.md
@@ -0,0 +1,41 @@
+---
+name: f3s-storage
+description: Reference skill for the f3s homelab storage layer, ZFS (`zdata`), zrepl replication, CARP storage VIP (f0/f1, `f3s-storage-ha` 192.168.1.138), NFS over stunnel, the nfs-mount-monitor watchdog, USB key material, local-path/backups, and storage troubleshooting (incl. thermal). Use when working on homelab storage, ZFS/zrepl, NFS mounts, CARP failover, or disk issues. Part of the f3s homelab skill family (hub, [`f3s`](../f3s/SKILL.md)).
+---
+
+# f3s Storage
+
+Persistent storage for k3s is served via **NFS over stunnel** from the FreeBSD hosts, backed by **ZFS** (`zdata` pool) with **CARP** for high availability and **zrepl** for continuous replication.
+
+Note: original plan was HAST, replaced by **zrepl** (ZFS send/receive) — more reliable, avoids the ZFS corruption during failover that HAST caused.
+
+## When to Use
+
+- Working on the homelab storage layer: ZFS pools/datasets, encryption, USB keys
+- zrepl replication (f0→f1 nfsdata, f3→f2 VM), CARP failover, NFS-over-stunnel
+- Diagnosing NFS mount problems, SUSPENDED pools, or thermal issues
+- For the physical hosts, WireGuard mesh, and host/IP inventory this depends on, see the [`f3s`](../f3s/SKILL.md) hub skill.
+
+## Reference Files
+
+- [ZFS Pools & Encryption](references/zfs.md) — `zdata` pool, physical disks, USB-stored keys mounted by `f3skeys` (not `/etc/fstab`), encrypted datasets, boot-time key loading
+- [USB Key Mounting](references/usb-keys.md) — `f3skeys`, `/usr/local/sbin/f3s-mount-keys`, and current `zfskeys_datasets` per f-host
+- [zrepl Replication](references/zrepl.md) — `f0 → f1` nfsdata, `f3 → f2` freebsd VM, sink configs, troubleshooting, DL-state recovery
+- [CARP HA VIP](references/carp.md) — VIP `192.168.1.138`, `carpcontrol.sh`, mgmt script, auto-failback, SUSPENDED-pool limitation
+- [NFS over stunnel](references/nfs.md) — NFS server, mutual-TLS stunnel, Rocky client config, `/etc/fstab`
+- [nfs-mount-monitor](references/nfs-mount-monitor.md) — systemd watchdog on r-nodes (mount/stat/write probes, fail counter, cordon-and-reboot escalation)
+- [Troubleshooting](references/troubleshooting.md) — NFS issues, ZFS pool SUSPENDED recovery, **thermal** troubleshooting (Beelink S12 Pro)
+- [Backups & Local-Path](references/backups.md) — S3 Glacier Deep Archive, when to use `local-path` instead of NFS
+
+## Storage Summary
+
+| Layer | Technology | Role |
+|-------|-----------|------|
+| Block | M.2+2.5" SSD (f0/f1) | Physical storage |
+| Filesystem | ZFS (`zdata/enc`) | Data integrity, AES-256-GCM encryption |
+| Replication | `zrepl` | Continuous ZFS replication f0→f1 (1min NFS, 10min VM) |
+| HA | CARP VIP 192.168.1.138 | Automatic failover for NFS/stunnel |
+| Network | NFS over stunnel | Encrypted shared storage, mutual TLS auth |
+| Local-path | k3s local-path provisioner | Node-local storage for SQLite/cache workloads |
+| LAN access | FreeBSD relayd on CARP VIP | TCP forwarding to k3s :80/:443 |
+| Backup | S3 Glacier Deep Archive | Off-site encrypted backup |
diff --git a/prompts/skills/f3s-storage/references/backups.md b/prompts/skills/f3s-storage/references/backups.md
new file mode 100644
index 0000000..99fd425
--- /dev/null
+++ b/prompts/skills/f3s-storage/references/backups.md
@@ -0,0 +1,40 @@
+# Backups and Local-Path Storage
+
+## AWS S3 Glacier Deep Archive Backups
+
+Encrypted incremental ZFS snapshots from `zdata` pool backed up daily to **AWS S3 Glacier Deep Archive** via cron. Scripts adapted from FreeBSD Home NAS setup. Also performs periodic zpool scrubbing.
+
+## Local-Path Storage for SQLite Workloads
+
+Some k3s workloads use `local-path` (k3s default storageClass) instead of NFS for
+their data volumes. This is appropriate when:
+
+- The application uses SQLite: NFS file-lock semantics cause `fcntl()` races on
+ pod restarts, and `Recreate` strategy only reduces (not eliminates) the risk.
+- Cache-heavy workloads: NFS over stunnel adds TLS round-trip latency to every
+ cache read. Navidrome's image/background cache init took ~19s over NFS; it
+ takes ~25ms from local disk.
+
+**Trade-off**: a local-path PV lives on one specific node. If that node is down,
+the pod reschedules elsewhere but finds no data volume — it starts with an empty DB,
+losing play history, scrobble queue, etc. For a home server this is acceptable.
+The deployment must pin the pod to the same node via `nodeSelector` so the local
+PV is always reachable.
+
+### Workloads using local-path
+
+| App | Node | Path on node |
+|-----|------|--------------|
+| navidrome `/data` (DB + cache) | r1 | `/var/lib/rancher/k3s/storage/pvc-*_services_navidrome-data-pvc` |
+
+### Migrating NFS hostPath → local-path
+
+1. Disable ArgoCD auto-sync: `kubectl patch application <app> -n cicd --type=json -p='[{"op":"replace","path":"/spec/syncPolicy","value":{}}]'`
+2. Scale deployment to 0: `kubectl scale deployment <app> -n services --replicas=0`
+3. Delete old PVC and static PV.
+4. Create new PVC with `storageClassName: local-path`.
+5. Create a migration pod pinned to the target node that mounts both the NFS hostPath
+ (source) and the new PVC (target); copy data with `cp -av /src/. /dst/`.
+6. Delete migration pod, apply updated deployment (with `nodeSelector`), scale back up.
+7. Re-enable ArgoCD auto-sync and push manifests to git; push to in-cluster git-server
+ (`git push r0 master`) so ArgoCD picks up the new storageClass spec.
diff --git a/prompts/skills/f3s-storage/references/carp.md b/prompts/skills/f3s-storage/references/carp.md
new file mode 100644
index 0000000..3f7b741
--- /dev/null
+++ b/prompts/skills/f3s-storage/references/carp.md
@@ -0,0 +1,95 @@
+# CARP: High-Availability VIP
+
+CARP (Common Address Redundancy Protocol) provides **VIP 192.168.1.138** that floats between f0 (primary) and f1 (standby). The VIP is what NFS clients and the FreeBSD `relayd` ingress connect to, so only the current MASTER serves traffic.
+
+## /etc/rc.conf configuration
+
+```sh
+# On f0 (default advskew=0, wins elections)
+ifconfig_re0_alias0="inet vhid 1 pass YOURPASSWORD alias 192.168.1.138/32"
+
+# On f1 (advskew=100, loses elections to f0)
+ifconfig_re0_alias0="inet vhid 1 advskew 100 pass YOURPASSWORD alias 192.168.1.138/32"
+```
+
+## Load CARP module
+
+```sh
+echo 'carp_load="YES"' | doas tee -a /boot/loader.conf
+# or immediately: doas kldload carp
+```
+
+## /etc/hosts for CARP VIP
+
+```
+192.168.1.138 f3s-storage-ha f3s-storage-ha.lan f3s-storage-ha.lan.buetow.org
+192.168.2.138 f3s-storage-ha.wg0 f3s-storage-ha.wg0.wan.buetow.org
+```
+
+## devd: CARP state change hook
+
+Add to `/etc/devd.conf` on f0 and f1:
+
+```
+notify 0 {
+ match "system" "CARP";
+ match "subsystem" "[0-9]+@[0-9a-z.]+";
+ match "type" "(MASTER|BACKUP)";
+ action "/usr/local/bin/carpcontrol.sh $subsystem $type";
+};
+```
+
+```sh
+doas service devd restart
+```
+
+## carpcontrol.sh — start/stop NFS+stunnel on failover
+
+Source of truth: `f3s/freebsd-hosts/carp/carpcontrol.sh`.
+
+Install on f0 and f1:
+
+```sh
+doas install -o root -g wheel -m 0555 carpcontrol.sh /usr/local/bin/carpcontrol.sh
+```
+
+The script must call `/usr/local/sbin/f3s-mount-keys` before any
+`zfs load-key` operation because `/keys` is not mounted by `/etc/fstab`; see
+[USB Key Mounting](usb-keys.md).
+
+## CARP management script (`/usr/local/bin/carp`)
+
+```sh
+doas carp # show current state
+doas carp master # force MASTER (e.g. reclaim after maintenance)
+doas carp backup # force BACKUP (trigger failover to f1)
+doas carp auto-failback disable # prevent auto-failback (for maintenance)
+doas carp auto-failback enable # re-enable auto-failback
+```
+
+## CARP failover limitation when ZFS is suspended
+
+If f0's ZFS pool is SUSPENDED but f0's OS is still running, f0 remains CARP MASTER
+(it keeps sending CARP advertisements). Attempts to manually demote f0 via:
+
+```sh
+doas carp backup # may return exit=0 but has no effect
+doas ifconfig re0 vhid 1 state backup # may return exit=1 silently
+doas ifconfig re0 vhid 1 advskew 254 # may return exit=1 silently
+```
+
+…can all silently fail because the kernel has too many stuck IO threads blocking
+the ifconfig ioctl path. The CARP VIP will **not** float to f1 in this case.
+**Only a hard power cycle of f0 reliably triggers CARP failover.** See
+[troubleshooting.md](troubleshooting.md) for the full SUSPENDED-pool recovery runbook.
+
+## Auto-failback from f1 to f0
+
+Script `/usr/local/bin/carp-auto-failback.sh` runs every minute via cron on f0. Checks: currently BACKUP? `/data/nfs` mounted? Marker file exists? Failback not blocked? If all conditions met, promotes f0 to MASTER.
+
+```sh
+echo "* * * * * /usr/local/bin/carp-auto-failback.sh" | doas crontab -
+doas touch /data/nfs/nfs.DO_NOT_REMOVE # marker file required for auto-failback
+```
+
+Logs to `/var/log/carp-auto-failback.log`.
diff --git a/prompts/skills/f3s-storage/references/nfs-mount-monitor.md b/prompts/skills/f3s-storage/references/nfs-mount-monitor.md
new file mode 100644
index 0000000..a9b71e7
--- /dev/null
+++ b/prompts/skills/f3s-storage/references/nfs-mount-monitor.md
@@ -0,0 +1,107 @@
+# NFS Auto-Repair: nfs-mount-monitor
+
+A systemd timer+service pair on r0/r1/r2 checks the NFS mount every 10 seconds and automatically repairs it if stale or missing.
+
+## Repo location
+
+```
+f3s/r-nodes/nfs-mount-monitor/
+ check-nfs-mount.sh # repair script → /usr/local/bin/
+ nfs-mount-monitor.service # one-shot service → /etc/systemd/system/
+ nfs-mount-monitor.timer # 10-second timer → /etc/systemd/system/
+f3s/r-nodes/Rexfile # Rex deploy task: nfs_mount_monitor
+```
+
+## Deploy
+
+```sh
+# From repo root — pushes to all three r-nodes and reloads systemd if anything changed
+rex -f f3s/r-nodes/Rexfile nfs_mount_monitor
+```
+
+## What it does
+
+Three probes run in sequence on every 10-second tick:
+
+1. **mountpoint probe** — detects completely missing mounts.
+2. **stat probe** (`timeout 2s stat`) — detects read hangs / stale cache misses.
+3. **write probe** (`timeout 5s sh -c "echo $$ > .healthcheck.<host> && rm -f ..."`) —
+ detects the "reads OK, writes hang" failure mode. Stunnel-wrapped NFSv4 can enter
+ a state where `stat` returns from cache but all writes block indefinitely; only this
+ probe catches it.
+
+If any probe fails, `fix_mount` runs:
+
+1. `mount -o remount -f` (cheapest, no disruption if mount is merely stale)
+2. Kill D-state processes pinning the mount (`kill_pinning_processes` — SIGKILLs
+ processes whose `wchan` starts with `nfs_` and whose cwd/fds point into the mountpoint)
+3. `umount -f` (force unmount)
+4. `umount -l` (lazy detach VFS node if `-f` failed)
+5. `systemctl restart stunnel` + 2s sleep (refresh the TLS transport)
+6. `mount -t nfs4 -o port=2323,soft,timeo=50,retrans=3` (explicit soft NFS mount — NOT
+ `mount $MOUNT_POINT` which reads fstab's `hard` flag and enters uninterruptible D-state
+ if the server is unreachable; SIGKILL cannot wake a D-state process on Linux;
+ `soft,timeo=50,retrans=3` returns ETIMEDOUT after ~15 s so the fail counter can
+ increment and eventually trigger the reboot escalation)
+
+A hard **60-second deadline** prevents `fix_mount` from outlasting its own timer interval.
+
+On successful repair, force-deletes pods on this node stuck in
+Unknown / Pending / ContainerCreating so the kubelet can reschedule them.
+
+**Consecutive-failure escalation**: each `fix_mount` failure increments a counter
+persisted to `/var/lib/nfs-mount-monitor/fail-count`. At `NFS_FAIL_THRESHOLD=5`
+consecutive failures (~50 s), the node cordons itself (`kubectl cordon`) and issues
+`systemctl reboot`. The cordon is stored in etcd and **persists across reboots** —
+after the underlying NFS issue is resolved, manually uncordon each affected node:
+```sh
+kubectl uncordon r0.lan.buetow.org
+kubectl uncordon r1.lan.buetow.org
+kubectl uncordon r2.lan.buetow.org
+```
+
+The counter is also exported to `/var/lib/node_exporter/textfile_collector/nfs_mount_monitor.prom`
+so Prometheus can alert on `nfs_mount_monitor_consecutive_failures` without parsing
+journal logs (warning ≥3, critical ≥5 — see
+`f3s/prometheus/manifests/nfs-mount-monitor-alerts.yaml`).
+
+Uses a lock file (`/var/run/nfs-mount-check.lock`) to prevent overlapping runs
+since the timer fires faster than the script's worst-case runtime. If the lock is
+older than **90 seconds** it was left by a run that was SIGKILLed before its EXIT
+trap could clean up (systemd kills with SIGKILL after its own timeout, bypassing
+`trap "rm -f $LOCK_FILE" EXIT`); the stale lock is removed and the run continues,
+preventing all health checks from being silently skipped forever.
+
+## Timer configuration
+
+| Parameter | Value | Reason |
+|-----------|-------|--------|
+| `OnBootSec` | 30s | Let network and NFS client start before first check |
+| `OnUnitActiveSec` | 10s | Check interval; each run is bounded by a 60-second deadline |
+| `AccuracySec` | 1s | Prevent systemd batching from delaying the 10 s interval |
+
+## Managing the monitor during an extended NFS outage
+
+During a prolonged NFS outage (e.g. while the storage host is being power-cycled or
+repaired), stop the timer on affected r-nodes to prevent the escalation counter from
+reaching the auto-reboot threshold prematurely:
+
+```sh
+# On each affected r-node (as root)
+systemctl stop nfs-mount-monitor.timer
+echo 0 > /var/lib/nfs-mount-monitor/fail-count # reset counter
+
+# After NFS is restored, restart and verify
+systemctl start nfs-mount-monitor.timer
+journalctl -u nfs-mount-monitor -f
+```
+
+Also reset the counter to 0 after uncordoning nodes (see escalation section above),
+because the old counter value would lower the effective threshold for the next outage.
+
+## Status and logs
+
+```sh
+systemctl status nfs-mount-monitor.timer
+journalctl -u nfs-mount-monitor -f
+```
diff --git a/prompts/skills/f3s-storage/references/nfs.md b/prompts/skills/f3s-storage/references/nfs.md
new file mode 100644
index 0000000..2c21114
--- /dev/null
+++ b/prompts/skills/f3s-storage/references/nfs.md
@@ -0,0 +1,204 @@
+# NFS over stunnel
+
+NFSv4 served from f0/f1 to the Rocky Linux k3s nodes (r0/r1/r2) over a
+TLS tunnel that terminates on the CARP VIP. NFS itself stays on localhost;
+stunnel handles transport encryption with mutual TLS.
+
+## NFS Server Configuration (f0 and f1)
+
+```sh
+doas sysrc nfs_server_enable=YES
+doas sysrc nfsv4_server_enable=YES
+doas sysrc nfsuserd_enable=YES
+doas sysrc nfsuserd_flags="-domain lan.buetow.org"
+doas sysrc mountd_enable=YES
+doas sysrc rpcbind_enable=YES
+doas sysrc nfs_reserved_port_only=NO # Required for NFS over stunnel (unprivileged ports)
+
+doas mkdir -p /data/nfs/k3svolumes
+doas chmod 755 /data/nfs/k3svolumes
+```
+
+> **FreeBSD 15.0 note**: FreeBSD 15.0 sets `nfs_reserved_port_only=YES` by default in `/etc/defaults/rc.conf`. The nfsd rc script (`/etc/rc.d/nfsd`) checks this variable and explicitly runs `sysctl vfs.nfsd.nfs_privport=1` at startup, overriding any value set in `/etc/sysctl.conf` or `/boot/loader.conf`. This blocks NFS clients connecting via stunnel (unprivileged ports). Fix on **each f-host**:
+> ```sh
+> # The ONLY correct fix — setting sysctl.conf does NOT work
+> doas sysrc nfs_reserved_port_only=NO
+> # Apply immediately without reboot
+> doas sysctl vfs.nfsd.nfs_privport=0
+> # Remount on each r-host
+> mount -a
+> ```
+
+`/etc/exports` (stunnel clients appear as localhost):
+
+```
+V4: /data/nfs -sec=sys
+/data/nfs -alldirs -maproot=root -network 127.0.0.1 -mask 255.255.255.255
+```
+
+Start services:
+
+```sh
+doas service rpcbind start
+doas service mountd start
+doas service nfsd start
+doas service nfsuserd start
+```
+
+## stunnel: Encrypted NFS over TLS
+
+stunnel binds to the CARP VIP (192.168.1.138), so only the CARP MASTER accepts connections. Uses mutual TLS with client certificate authentication.
+
+### Create CA and certificates (on f0)
+
+```sh
+doas mkdir -p /usr/local/etc/stunnel/ca
+cd /usr/local/etc/stunnel/ca
+doas openssl genrsa -out ca-key.pem 4096
+doas openssl req -new -x509 -days 3650 -key ca-key.pem -out ca-cert.pem \
+ -subj '/C=US/ST=State/L=City/O=F3S Storage/CN=F3S Stunnel CA'
+
+cd /usr/local/etc/stunnel
+doas openssl genrsa -out server-key.pem 4096
+doas openssl req -new -key server-key.pem -out server.csr \
+ -subj '/C=US/ST=State/L=City/O=F3S Storage/CN=f3s-storage-ha.lan'
+doas openssl x509 -req -days 3650 -in server.csr -CA ca/ca-cert.pem \
+ -CAkey ca/ca-key.pem -CAcreateserial -out server-cert.pem
+
+# Client certs for r0, r1, r2, earth
+for client in r0 r1 r2 earth; do
+ openssl genrsa -out ca/${client}-key.pem 4096
+ openssl req -new -key ca/${client}-key.pem -out ca/${client}.csr \
+ -subj "/C=US/ST=State/L=City/O=F3S Storage/CN=${client}.lan.buetow.org"
+ openssl x509 -req -days 3650 -in ca/${client}.csr -CA ca/ca-cert.pem \
+ -CAkey ca/ca-key.pem -CAcreateserial -out ca/${client}-cert.pem
+ cat ca/${client}-cert.pem ca/${client}-key.pem > ca/${client}-stunnel.pem
+done
+```
+
+### stunnel server config (`/usr/local/etc/stunnel/stunnel.conf`)
+
+```
+cert = /usr/local/etc/stunnel/server-cert.pem
+key = /usr/local/etc/stunnel/server-key.pem
+setuid = stunnel
+setgid = stunnel
+
+[nfs-tls]
+accept = 192.168.1.138:2323
+connect = 127.0.0.1:2049
+CAfile = /usr/local/etc/stunnel/ca/ca-cert.pem
+verify = 2
+requireCert = yes
+```
+
+```sh
+doas pkg install -y stunnel
+doas sysrc stunnel_enable=YES
+doas service stunnel start
+# Copy certs to f1 via tarball, configure identically
+```
+
+## NFS Client Configuration (Rocky Linux r0, r1, r2)
+
+```sh
+dnf install -y stunnel nfs-utils
+
+# Copy client cert and CA from f0
+scp f0:/usr/local/etc/stunnel/ca/r0-stunnel.pem /etc/stunnel/
+scp f0:/usr/local/etc/stunnel/ca/ca-cert.pem /etc/stunnel/
+```
+
+`/etc/stunnel/stunnel.conf` (r0 example):
+
+```
+cert = /etc/stunnel/r0-stunnel.pem
+CAfile = /etc/stunnel/ca-cert.pem
+client = yes
+verify = 2
+
+[nfs-ha]
+accept = 127.0.0.1:2323
+connect = 192.168.1.138:2323
+```
+
+```sh
+systemctl enable --now stunnel
+```
+
+### NFSv4 user mapping
+
+`/etc/idmapd.conf` on r0, r1, r2:
+
+```
+[General]
+Domain = lan.buetow.org
+```
+
+Fix inotify limit:
+
+```sh
+echo 'fs.inotify.max_user_instances = 512' > /etc/sysctl.d/99-inotify.conf
+sysctl -w fs.inotify.max_user_instances=512
+systemctl enable --now nfs-client.target nfs-idmapd
+```
+
+### Mount NFS
+
+```sh
+mkdir -p /data/nfs/k3svolumes
+mount -t nfs4 -o port=2323 127.0.0.1:/k3svolumes /data/nfs/k3svolumes
+```
+
+`/etc/fstab`:
+
+```
+127.0.0.1:/k3svolumes /data/nfs/k3svolumes nfs4 port=2323,_netdev,hard,timeo=600,retrans=3 0 0
+```
+
+NFS path structure on k3s nodes: `/data/nfs/k3svolumes/<app>/`
+
+The `nfs-mount-monitor` watchdog on each r-node detects and repairs stale or
+hung mounts automatically — see [nfs-mount-monitor.md](nfs-mount-monitor.md).
+
+## NFS Client Configuration (earth, roaming laptop)
+
+`earth` mounts two exports (`/earthdata`, `/k3svolumes`) the same way as the
+r-nodes — local stunnel client on `127.0.0.1:2323` → CARP VIP `192.168.1.138:2323`.
+Certs live in `/etc/stunnel/` (`earth-stunnel.pem`, `ca-cert.pem`). The mounts are
+`noauto` and mounted on demand.
+
+**Mount options differ from the r-nodes on purpose.** The r-nodes are on a fast,
+stable LAN and use `hard,timeo=600`. `earth` roams on WiFi, so it uses `soft` so it
+fails instead of hanging forever when off-network. The working `/etc/fstab` lines:
+
+```
+127.0.0.1:/earthdata /data/nfs/earthdata nfs4 port=2323,_netdev,soft,timeo=150,retrans=3,nofail,noauto 0 0
+127.0.0.1:/k3svolumes /data/nfs/k3svolumes nfs4 port=2323,_netdev,soft,timeo=150,retrans=3,nofail,noauto 0 0
+```
+
+> **Critical: do NOT use `timeo=10` on earth.** `timeo` is in **deciseconds**, so
+> `timeo=10` = **1.0s per RPC**. A large write is split into many `wsize` (128KB)
+> RPCs plus a final `COMMIT`; with `soft,timeo=10,retrans=2`, any single RPC or the
+> COMMIT that can't be acked within ~1s over jittery WiFi makes the soft mount abort
+> mid-transfer with EIO. That both **corrupts the partial file** and leaves a wedged
+> NFS session that retransmits constantly, saturating 2.4GHz WiFi airtime (slows ALL
+> WiFi traffic, e.g. unrelated scp to/from earth). `timeo=150` (15s) gives each RPC
+> and the COMMIT enough slack to ride out WiFi jitter while still failing eventually
+> if truly disconnected. Drop the deprecated `intr` option (no-op on modern Linux).
+
+**Superblock caching gotcha:** Linux NFS shares one superblock per `server:export`,
+so remounting with new options is ignored while any process still references the old
+mount (e.g. a shell with its cwd inside it). Either release all references first, or
+remount with `-o nosharecache` to force a fresh superblock with the new options.
+
+**The real fix is the mount option, not the transfer tool.** With `timeo=150` the
+mount never wedges, so transfers just complete and plain `cp`/`mv` are perfectly
+safe — they check return codes and report errors normally. The one time we saw a
+half-finalized destination file, it was *not* a `cp`/`mv` flaw: the in-flight `mv`
+was killed mid-finalize by a forced f0 reboot + stunnel restart that we only did
+*because* the mount had wedged (the `timeo=10` bug). Fix the hang and that whole
+chain disappears. `rsync` (temp-file-then-rename, `--remove-source-files` only after
+a verified copy) is worth using as **interruption insurance** on a flaky WiFi link —
+it can't leave a bad file under the real name if the link dies mid-transfer — but it
+is not a substitute for the `timeo` fix and `cp`/`mv` do not "cause corruption."
diff --git a/prompts/skills/f3s-storage/references/troubleshooting.md b/prompts/skills/f3s-storage/references/troubleshooting.md
new file mode 100644
index 0000000..7baaa63
--- /dev/null
+++ b/prompts/skills/f3s-storage/references/troubleshooting.md
@@ -0,0 +1,198 @@
+# Storage Troubleshooting
+
+NFS issues, ZFS pool SUSPENDED recovery, and thermal problems on the
+Beelink S12 Pro mini-PCs.
+
+## NFS Troubleshooting
+
+### All r-nodes show "access denied" when mounting NFS
+
+**Most likely cause**: `vfs.nfsd.nfs_privport=1` on the CARP MASTER. This happens after f-host reboots if `nfs_reserved_port_only` is not set to `NO` in rc.conf. The nfsd rc script (`/etc/rc.d/nfsd`) explicitly sets the sysctl based on this variable, overriding `/etc/sysctl.conf`. Fix: `doas sysrc nfs_reserved_port_only=NO` on both f0 and f1.
+
+### stunnel appears not running but port 2323 is bound
+
+`carpcontrol.sh` starts stunnel on CARP MASTER transition, but doesn't write a PID file. So `service stunnel status` reports "not running" even though stunnel is actually serving connections. Check with `doas sockstat -l | grep 2323`. If there's a stale stunnel process, kill it and restart: `doas kill <pid> && doas service stunnel start`.
+
+### Pods stuck in ContainerCreating/Unknown after NFS recovery
+
+After NFS is restored on the server side, the `nfs-mount-monitor` systemd timer on each r-node will auto-remount within ~10 seconds and force-delete stuck pods. If immediate recovery is needed: `mount /data/nfs/k3svolumes` on each r-node, then delete the stuck pods manually.
+
+**Note:** The monitor catches three failure modes: missing mountpoint, stat hang (reads unresponsive), and **silent write hang** (reads OK but writes block — the hardest case, e.g. stunnel-wrapped NFSv4 after a CARP failover). Watch the consecutive-failure counter via Prometheus (`nfs_mount_monitor_consecutive_failures`) — warning fires at ≥3, critical at ≥5. At 5 consecutive failures the node cordons itself and reboots.
+
+### Large file transfers from earth hang / WiFi slows to a crawl
+
+**Symptom**: copying a large file into earth's NFS mount hangs; kernel logs show
+`nfs: server 127.0.0.1 not responding, timed out` and `NFSv4: state recovery
+failed ... error = -116` (ESTALE). Unrelated WiFi traffic on earth (e.g. scp
+from f1) also slows dramatically at the same time.
+
+**Root cause**: earth's mount used `soft,timeo=10` (1.0s per RPC). Large writes
+abort mid-transfer, leaving a wedged NFS session that retransmits constantly and
+saturates 2.4GHz WiFi airtime — which throttles all of earth's WiFi traffic. This
+is a **client-side** problem; f0 itself is healthy (verified: local write to
+`zdata` ~1.6 GB/s, pool ONLINE, disk_wait ~1ms, cores 71–79°C).
+
+**Fix** (no f0 reboot needed — rebooting f0 only "worked" as a side effect of
+tearing down the wedged session, and it disrupts the k3s cluster via CARP failover):
+
+```sh
+# On earth: clear the wedged session
+sudo systemctl restart stunnel
+
+# Fix the mount options (see nfs.md → earth client config): use timeo=150, soft
+# Force new options past the cached superblock if a process still holds the mount:
+sudo umount -l /data/nfs/earthdata
+sudo mount -t nfs4 -o port=2323,_netdev,soft,timeo=150,retrans=3,nosharecache \
+ 127.0.0.1:/earthdata /data/nfs/earthdata
+```
+
+Then redo the transfer with `rsync --remove-source-files` (not `mv`), which avoids
+corrupt partials. Diagnostics that confirmed the transport (not f0) was the limit:
+raw ssh earth→f0 ~3 MB/s, NFS ~5 MB/s, while the radio negotiated 97/206 Mbit/s —
+the ceiling is single-stream TCP over jittery 2.4GHz WiFi + stunnel TLS, not f0 I/O.
+
+### Checklist for NFS outage on CARP MASTER (f0 or f1)
+
+```sh
+# 1. Check which host is CARP MASTER
+ssh paul@f0 'ifconfig re0 | grep carp'
+ssh paul@f1 'ifconfig re0 | grep carp'
+
+# 2. On the MASTER, verify:
+doas sysctl vfs.nfsd.nfs_privport # must be 0
+doas service nfsd status # must be running
+doas sockstat -l | grep 2323 # stunnel must be listening
+ls /data/nfs/nfs.DO_NOT_REMOVE # ZFS dataset must be mounted
+
+# 3. Fix if needed:
+doas sysrc nfs_reserved_port_only=NO # persist the fix
+doas sysctl vfs.nfsd.nfs_privport=0 # apply immediately
+doas service nfsd restart
+# For stunnel, kill stale process if needed, then:
+doas service stunnel start
+```
+
+## ZFS pool SUSPENDED recovery
+
+**Symptoms**: `doas zpool status zdata` shows `state: SUSPENDED`. All IO to the pool is
+halted — ZFS suspends itself to prevent corruption when IO errors exceed the threshold.
+Commands like `zpool clear`, `zpool scrub`, `zpool offline`, and even `ls /data/nfs/` hang
+indefinitely because they wait for kernel IO that will never complete.
+
+**Known cause (2026-05-15)**: Samsung 870 EVO 1TB on f0 (ada1) hit 107 read errors and
+105M+ write errors during normal operation. Subsequent investigation pointed at
+**thermal throttling** in the small Beelink S12 Pro enclosure — see the Thermal
+section below.
+
+**Recovery — hard power cycle only**:
+- Do NOT attempt `doas shutdown -r now` — if ZFS is suspended, the graceful shutdown hangs
+ at ZFS pool export and may stay stuck for 30–60+ minutes.
+- Do NOT attempt `doas zpool clear zdata` — it hangs because ada1 is unresponsive.
+- Do NOT attempt `doas ifconfig re0 vhid 1 state backup` or `doas carp backup` to fail
+ over to f1 first — these ifconfig ioctls can also be blocked when the kernel has too
+ many stuck IO threads. They may return exit=1 silently.
+- **Hard power cycle** (pull power or hold the power button) resolves the issue in ~9 s
+ (Rocky Linux VMs come up automatically, ZFS pool imports cleanly on next boot).
+
+**Post-recovery**:
+```sh
+# 1. Verify pool health
+doas zpool status zdata # should show ONLINE, 0 errors
+
+# 2. Check SMART for drive health
+doas smartctl -a /dev/ada1 | grep -iE '(temperature|reallocated|pending|uncorrectable|error)'
+
+# 3. Start a scrub to verify data integrity
+doas zpool scrub zdata
+doas zpool status zdata # monitor; "scrub repaired 0 in ..." means data intact
+
+# 4. Verify NFS is serving (stunnel listening on CARP VIP)
+doas sockstat -l | grep 2323
+```
+
+**After cluster recovery**:
+- Check for cordoned nodes: `kubectl get nodes` — if r0/r1/r2 show `SchedulingDisabled`,
+ uncordon them (see `nfs-mount-monitor.md` escalation section).
+- Reset fail counters on all r-nodes: `echo 0 > /var/lib/nfs-mount-monitor/fail-count`
+
+## Thermal Troubleshooting
+
+The 2026-05-16 f0 incident — and the 2026-05-15 ZFS SUSPENDED above — both trace
+back to **thermal problems in the Beelink S12 Pro enclosure**, not to any
+software-side cause. The mitigations and side-investigations (zrepl interval,
+autotrim, encryption overhead) are not what fixed it; reseating the drive and
+improving cooling did.
+
+### Symptoms of thermal throttling on f-hosts
+
+- SSD I/O slowness (writes dropping from MB/s to KB/s)
+- ZFS txg sync times jumping from <100 ms to many seconds
+- rsync / zrepl jobs going into D-state (waiting on ZFS I/O)
+- SMART reporting elevated drive temperature
+
+### How to check temperatures
+
+- **coretemp (real per-core die temps)**: `kldload coretemp; sysctl dev.cpu | grep temperature`
+ - Persist via `/boot/loader.conf` (`coretemp_load="YES"`)
+- **hw.acpi.thermal.tz0**: often a constant lie (e.g. always 27.9 °C) — do NOT rely on it
+- **SSD temperature**: `smartctl -a /dev/adaN` (requires `smartmontools`; may not be installed)
+- **Disk I/O performance**: `gstat -bp -I 1s -d` (FreeBSD `gstat`, not Linux `iostat`)
+
+### Beelink S12 Pro specifics
+
+The S12 Pro has an **active cooler** (fan + copper heat pipe), not purely passive.
+Intel N100 TJmax is **105 °C**; hardware throttling begins around 100 °C.
+
+Observed ranges (coretemp per-core die temps via `sysctl dev.cpu | grep temperature`):
+
+| Scenario | Expected | Action |
+|----------|----------|--------|
+| True idle (no VMs, no ZFS activity) | 35–50 °C | — |
+| Moderate sustained load (k3s + bhyve + ZFS) | 55–75 °C | Normal for f0–f2 |
+| Heavy load (Prime95 / scrub + rsync + repl) | 75–96 °C | Occasional throttling is OK short-term |
+| **>80 °C at moderate load** | — | **Investigate**: dust, airflow, thermal paste |
+| **>90 °C sustained** | — | **Stop I/O workloads, inspect hardware** |
+| >100 °C | — | Hardware thermal throttle active; shut down |
+
+Real-world observations (2026-05-17, post-reboot, k3s + bhyve running):
+- f3 (light load, no k3s): 43–44 °C — reference baseline
+- f0 (after drive reseat + cleaning): 64–66 °C — healthy
+- f1 (full k3s + ZFS): 76–77 °C — elevated but within range
+- f2 (full k3s + ZFS): 79–80 °C — near concern threshold; physical check recommended
+
+NVMe sits close to the CPU in the small chassis — both heat each other.
+The enclosure gets hot to the touch before temps fully register in software.
+
+### Cause and resolution (2026-05-16 f0)
+
+The cascade was thermal-only:
+
+1. Hot enclosure (NVMe physically very hot) → SSD/SATA thermal throttling
+2. Throttled disk → ZFS txg syncs balloon from <100 ms to multi-second
+3. rsync / zrepl block on ZFS → D-state, hung pods on r-nodes
+
+**Root cause**: hot enclosure / inadequate cooling. **Resolution**: shut down,
+reseat the drive, clean dust and improve airflow; the disk recovered immediately
+and ZFS txg sync times returned to normal.
+
+### Remediation steps
+
+1. SSH in and check temps: `kldload coretemp && sysctl dev.cpu | grep temperature`
+2. If >80 °C at moderate load: investigate airflow and dust first
+3. If >90 °C sustained: stop heavy I/O workloads (`service zrepl stop`, cancel scrubs)
+4. Physical: shut down, reseat NVMe, clean dust from vents, improve airflow
+5. Persist coretemp: ensure `/boot/loader.conf` has `coretemp_load="YES"`
+
+### Temperature monitoring
+
+```sh
+# FreeBSD: load coretemp for CPU package temperature
+doas kldload coretemp
+sysctl -a | grep temperature # hw.acpi.thermal.*: and dev.cpu.*:
+# Persist across reboots
+echo 'coretemp_load="YES"' | doas tee -a /boot/loader.conf
+
+# SSD temperature (install smartmontools if absent)
+doas pkg install -y smartmontools
+doas smartctl -a /dev/ada1 | grep -i temperature # "194 Temperature_Celsius"
+```
diff --git a/prompts/skills/f3s-storage/references/usb-keys.md b/prompts/skills/f3s-storage/references/usb-keys.md
new file mode 100644
index 0000000..9871daa
--- /dev/null
+++ b/prompts/skills/f3s-storage/references/usb-keys.md
@@ -0,0 +1,116 @@
+# USB Key Mounting for ZFS Encryption
+
+The f-hosts keep raw ZFS encryption keys on per-host UFS USB sticks mounted at
+`/keys`. All four sticks are labeled `F3S_KEYS` and hold all 8 key files as
+cross-host backups.
+
+Do **not** mount `/keys` from `/etc/fstab`. A missing or corrupt key stick must
+not block the FreeBSD base OS from booting.
+
+## Managed Files
+
+Source files live in the conf repo:
+
+```text
+f3s/freebsd-hosts/keys/
+ f3s-mount-keys
+ f3s-load-zfs-keys
+ f3skeys.rc
+```
+
+Installed paths on each f-host:
+
+```text
+/usr/local/sbin/f3s-mount-keys
+/usr/local/sbin/f3s-load-zfs-keys
+/etc/rc.d/f3skeys
+```
+
+`f3skeys` runs before FreeBSD's built-in `zfskeys` service. If the USB stick is
+missing or `fsck_ufs -p` fails, the helper logs the problem and exits
+successfully so boot continues. Encrypted datasets stay locked until the stick
+is repaired and `/usr/local/sbin/f3s-load-zfs-keys` is run manually.
+
+## Setup
+
+Format a new key stick:
+
+```sh
+doas newfs -L F3S_KEYS /dev/da0
+doas mkdir -p /keys
+doas mount -t ufs -o ro /dev/ufs/F3S_KEYS /keys
+```
+
+Label an existing stick without rebuilding it:
+
+```sh
+doas umount /keys
+doas tunefs -L F3S_KEYS /dev/da0
+```
+
+Keep the old `/etc/fstab` line commented on all f-hosts:
+
+```fstab
+# /dev/da0 /keys ufs rw 0 2
+```
+
+Enable boot loading:
+
+```sh
+doas sysrc f3skeys_enable=YES
+doas sysrc zfskeys_enable=YES
+```
+
+Current `zfskeys_datasets` values:
+
+```sh
+# f0
+doas sysrc zfskeys_datasets="zdata/enc zdata/enc/nfsdata zroot/bhyve zroot/garage"
+
+# f1
+doas sysrc zfskeys_datasets="zdata/enc zroot/bhyve zroot/garage zdata/sink/f0/zdata/enc/nfsdata"
+
+# f2
+doas sysrc zfskeys_datasets="zdata/enc zroot/bhyve zroot/garage zroot/sink/f3/zroot/bhyve/freebsd"
+
+# f3
+doas sysrc zfskeys_datasets="zroot/bhyve"
+```
+
+Replicated sinks with raw encryption need explicit file keylocations:
+
+```sh
+# f1
+doas zfs set keylocation=file:///keys/f0.lan.buetow.org:zdata.key \
+ zdata/sink/f0/zdata/enc/nfsdata
+
+# f2
+doas zfs set keylocation=file:///keys/f3.lan.buetow.org:bhyve.key \
+ zroot/sink/f3/zroot/bhyve/freebsd
+```
+
+Manual recovery after boot:
+
+```sh
+doas /usr/local/sbin/f3s-mount-keys --strict
+doas /usr/local/sbin/f3s-load-zfs-keys
+```
+
+## Verification
+
+```sh
+mount | grep ' /keys '
+sysrc -n f3skeys_enable
+sysrc -n zfskeys_enable
+sysrc -n zfskeys_datasets
+doas /usr/local/sbin/f3s-load-zfs-keys
+zfs list -H -o name,encryption,keylocation,keystatus,mounted |
+ awk '$2 != "off" { print }'
+```
+
+Full reboot validation was run on f0, f1, f2, and f3 on 2026-05-30 after this
+change.