diff --git a/engine/lint-report.md b/engine/lint-report.md index 3b81de2..d86d940 100644 --- a/engine/lint-report.md +++ b/engine/lint-report.md @@ -1,6 +1,6 @@ # Vault Lint Report -Generated: 2026-08-15T12:00:04Z | Docs scanned: 115 | Elapsed: 0.0s +Generated: 2026-08-15T18:00:04Z | Docs scanned: 115 | Elapsed: 0.0s ## Summary diff --git a/vault/.obsidian/workspace.json b/vault/.obsidian/workspace.json index 38f0409..e4ad7fb 100644 --- a/vault/.obsidian/workspace.json +++ b/vault/.obsidian/workspace.json @@ -199,6 +199,9 @@ }, "active": "17bd4a6166f789d0", "lastOpenFiles": [ + "projects/fleet-storage-memory-upgrade.md.tmp.2730138.81f9975eb3a6", + "runbooks/navi-lift-to-media.md.tmp.2730138.bee0e1488a4b", + "runbooks/navi-lift-to-media.md.tmp.2730138.cb383db7c987", "runbooks/navi-lift-to-media.md.tmp.2730138.5755c5b630d8", "runbooks/navi-lift-to-media.md.tmp.2730138.74d5fc239a82", "projects/navi-recon-separation.md.tmp.2730138.1a9ee67ba5ca", @@ -208,9 +211,6 @@ "docs/hardware/environment.md.tmp.2730138.58738aa3e22c", "docs/hardware/environment.md.tmp.2730138.e8527bb48e1d", "docs/hardware/environment.md.tmp.2730138.0fcdc6bee9f6", - "runbooks/edge2-access-reference.md.tmp.3701844.be7e6cfc6226", - "runbooks/peertube-sitemap-redis-oom.md.tmp.3701844.88863a688de8", - "projects/fleet-storage-memory-upgrade.md.tmp.3701844.7763e1bbbff5", "runbooks/edge2-boot-recovery.md", "runbooks/corescope-ingest-stall-oom.md", "projects/navi-recon-separation.md", diff --git a/vault/projects/fleet-storage-memory-upgrade.md b/vault/projects/fleet-storage-memory-upgrade.md index 0b7f8c3..18748a0 100644 --- a/vault/projects/fleet-storage-memory-upgrade.md +++ b/vault/projects/fleet-storage-memory-upgrade.md @@ -20,16 +20,22 @@ Placing a batch of acquired drives and memory across the fleet, and the storage- ## Parts on hand -| Part | Qty | Form factor | Where it fits | Gain | -|---|---|---|---|---| -| 32 GB DDR5 SODIMM | 2 | SODIMM | **media only** | 32 → 64 GB, its board maximum | -| 2 TB NVMe | 1 | M.2 2280 | media's freed M.2 | +2 TB fast local | -| 1 TB NVMe | 2 | M.2 2280 | utility, cloud, or toc | replaces a 512 GB, needs migration | -| 16 GB DDR4 SODIMM | 4 | SODIMM | **nowhere** | cold spares only | +Status as of 2026-08-15. -The DDR4 has no home. data, utility and cloud each already run 2 × 16 GB DDR4-3200; data and utility are hard-capped at 32 GB, and cloud would need 32 GB sticks to gain anything. toc's two empty slots take **full-size DIMMs**, not SODIMM. Keep them as field spares for the three ThinkCentres. +| Part | Qty | Form factor | Disposition | +|---|---|---|---| +| 32 GB DDR5 SODIMM | 2 | SODIMM | **INSTALLED** in media — 32 → 64 GB, board maximum reached | +| 2 TB NVMe (WD Green SN350) | 1 | M.2 2280 | **INSTALLED** in media as VG `tank` → `/mnt/nvme2tb` | +| 1 TB NVMe | 1 | M.2 2280 | **used elsewhere** — moved into a laptop | +| 1 TB NVMe | 1 | M.2 2280 | **spare, earmarked for utility** (no date set) | +| 16 GB DDR4 SODIMM | 2 | SODIMM | **used elsewhere** — moved into the same laptop | +| 16 GB DDR4 SODIMM | 2 | SODIMM | spare | +| 512 GB Intel NVMe | 1 | M.2 2280 | spare — pulled from media, 0% wear, 9,791 hrs, holds a 2022 BitLocker Windows install | +| 16 GB DDR5-4800 ADATA SODIMM | 2 | SODIMM | spare — displaced from media | -Freed by the media work: one 512 GB Intel NVMe (0% wear, 9,791 hrs) and two 16 GB DDR5-4800 ADATA SODIMMs. +**Why utility is the target for the remaining 1 TB:** it has the fullest thin pool in the fleet (45% vs cloud's 23%) and the fastest-wearing drive — a 512 GB WD SN740 at **12% life used with 85 TB written in only 3,459 hours**, roughly 166 full drive-writes. Its single M.2 holds the boot drive, so this is a clone-and-swap, not an add. Its empty 2.5" SATA bay cannot take an M.2 drive. + +The DDR4 SODIMMs had no home in the fleet — data, utility and cloud all already run 2 × 16 GB and the first two are hard-capped at 32 GB, while toc's empty slots need full-size DIMMs. Two went to the laptop instead. --- diff --git a/vault/runbooks/navi-lift-to-media.md b/vault/runbooks/navi-lift-to-media.md index bd05d4d..b84efed 100644 --- a/vault/runbooks/navi-lift-to-media.md +++ b/vault/runbooks/navi-lift-to-media.md @@ -15,7 +15,9 @@ updated: 2026-08-15 --- # Separate navi onto its own VM on media -**Status: Phase 0 complete. Nothing from Phase 1 onward has been run.** +**Status: COMPLETE through Phase 5 (2026-08-15). Phase 6 deliberately NOT run — rollback stays open.** + +navi now runs on **navi-vm (VMID 1131) on media**, tailnet `100.64.0.27`, LAN `192.168.1.132`. Cutover outage was **2 min 10 s** (20:56:47–20:58:57 UTC), navi only — `recon-vm` was never stopped and recon/kiwix stayed up throughout. Verified 9/9 on the functional suite plus headless confirmation of map render, search, place details, DEM elevation, and the USFS/BLM/Public Lands layers. Splits [[navi]] out of `recon-vm` (VMID 1130 on data) into its own VM on media, taking **all live data onto local NVMe**. `recon-vm` keeps running throughout — it is never stopped. Background on the entanglement is [[navi-recon-separation]]. @@ -183,7 +185,7 @@ Only after navi has been healthy on media for **at least a week**: - `systemctl disable --now virtiofsd-nav` on data - `rm -rf /opt/recon` on navi-vm - Drop 1130's memory from 24 GB to ~8 GB -- Optionally remove the 658 GB DEM from pi-nas once the local copy is proven +- **Never remove anything from pi-nas** — it is the source tier, not a cache (Matt, 2026-08-15). The 658 GB DEM stays there alongside the build material it was derived from. **Keep the source data and VM 1130 intact until then.** That is the rollback path. @@ -199,6 +201,20 @@ Only after navi has been healthy on media for **at least a week**: --- +## What actually bit us (2026-08-15 run) + +**`systemctl start` is not `restart`.** Phase 3 ran twice; the first attempt started the navi services *before* the second DEM reference in `navi-offroute.env` was found and fixed. The second run used `start`, which is a no-op on a running unit, so `navi-offroute` kept the stale environment and threw `FileNotFoundError: '/mnt/nas/nav/planet-dem.pmtiles'` on every terrain request while the on-disk config looked correct. **Verify the live process environment, not the config file:** `tr '\0' '\n' < /proc/$(systemctl show -p MainPID --value )/environ | grep `. + +**Two files reference the DEM, not one** — `navi-geo.env` *and* `navi-offroute.env`. + +**Do not rsync onto a running database.** The first Phase 5 delta was taken with navi-vm's postgres still running, overwriting `pgdata` underneath it. Recovered by stopping postgres on **both** sides and re-syncing (17 files, 66 KB — the first pass had been essentially complete). Quiesce the destination as well as the source. + +**qmrestore clones the MAC and IP.** The restored VM carried recon-vm's `BC:24:11:07:B0:F9` and its cloud-init IP. Both were changed with the NIC `link_down=1`, using the QEMU guest agent over virtio-serial, before it ever saw the network. + +**Tailscale must be re-registered before nginx will start.** nginx has an upstream on `central.echo6.mesh`, which only resolves via MagicDNS. With tailscaled disabled during identity cleanup, `nginx -t` failed with `host not found in upstream`. (That upstream is fine, incidentally — CT 104 kept the name "central" but runs Conduit.) + +**A size guard cost a re-run.** `658 GiB` is `705,591,064,874` decimal bytes; a `> 700000000000` sanity check rejected a byte-perfect copy. + ## Known traps **`virtiofsd` dies with the VM.** Stopping VM 1130 leaves `virtiofsd-{nav,kiwix,library}.service` inactive, and it then **refuses to start** with `Failed to connect to /run/virtiofsd-*.sock`. Restart all three first. Confirmed the hard way 2026-08-15. navi-vm has no virtiofs, so it never inherits this.