auto: docs sync 2026-08-16T00:00:04+00:00

Files changed: engine/lint-report.md vault/.obsidian/workspace.json vault/projects/fleet-storage-memory-upgrade.md vault/runbooks/navi-lift-to-media.md
This commit is contained in:
echo6-autocommit 2026-08-16 00:00:04 +00:00
commit 5496a8d428
4 changed files with 38 additions and 16 deletions

View file

@ -15,7 +15,9 @@ updated: 2026-08-15
---
# Separate navi onto its own VM on media
**Status: Phase 0 complete. Nothing from Phase 1 onward has been run.**
**Status: COMPLETE through Phase 5 (2026-08-15). Phase 6 deliberately NOT run — rollback stays open.**
navi now runs on **navi-vm (VMID 1131) on media**, tailnet `100.64.0.27`, LAN `192.168.1.132`. Cutover outage was **2 min 10 s** (20:56:4720:58:57 UTC), navi only — `recon-vm` was never stopped and recon/kiwix stayed up throughout. Verified 9/9 on the functional suite plus headless confirmation of map render, search, place details, DEM elevation, and the USFS/BLM/Public Lands layers.
Splits [[navi]] out of `recon-vm` (VMID 1130 on data) into its own VM on media, taking **all live data onto local NVMe**. `recon-vm` keeps running throughout — it is never stopped. Background on the entanglement is [[navi-recon-separation]].
@ -183,7 +185,7 @@ Only after navi has been healthy on media for **at least a week**:
- `systemctl disable --now virtiofsd-nav` on data
- `rm -rf /opt/recon` on navi-vm
- Drop 1130's memory from 24 GB to ~8 GB
- Optionally remove the 658 GB DEM from pi-nas once the local copy is proven
- **Never remove anything from pi-nas** — it is the source tier, not a cache (Matt, 2026-08-15). The 658 GB DEM stays there alongside the build material it was derived from.
**Keep the source data and VM 1130 intact until then.** That is the rollback path.
@ -199,6 +201,20 @@ Only after navi has been healthy on media for **at least a week**:
---
## What actually bit us (2026-08-15 run)
**`systemctl start` is not `restart`.** Phase 3 ran twice; the first attempt started the navi services *before* the second DEM reference in `navi-offroute.env` was found and fixed. The second run used `start`, which is a no-op on a running unit, so `navi-offroute` kept the stale environment and threw `FileNotFoundError: '/mnt/nas/nav/planet-dem.pmtiles'` on every terrain request while the on-disk config looked correct. **Verify the live process environment, not the config file:** `tr '\0' '\n' < /proc/$(systemctl show -p MainPID --value <svc>)/environ | grep <VAR>`.
**Two files reference the DEM, not one** — `navi-geo.env` *and* `navi-offroute.env`.
**Do not rsync onto a running database.** The first Phase 5 delta was taken with navi-vm's postgres still running, overwriting `pgdata` underneath it. Recovered by stopping postgres on **both** sides and re-syncing (17 files, 66 KB — the first pass had been essentially complete). Quiesce the destination as well as the source.
**qmrestore clones the MAC and IP.** The restored VM carried recon-vm's `BC:24:11:07:B0:F9` and its cloud-init IP. Both were changed with the NIC `link_down=1`, using the QEMU guest agent over virtio-serial, before it ever saw the network.
**Tailscale must be re-registered before nginx will start.** nginx has an upstream on `central.echo6.mesh`, which only resolves via MagicDNS. With tailscaled disabled during identity cleanup, `nginx -t` failed with `host not found in upstream`. (That upstream is fine, incidentally — CT 104 kept the name "central" but runs Conduit.)
**A size guard cost a re-run.** `658 GiB` is `705,591,064,874` decimal bytes; a `> 700000000000` sanity check rejected a byte-perfect copy.
## Known traps
**`virtiofsd` dies with the VM.** Stopping VM 1130 leaves `virtiofsd-{nav,kiwix,library}.service` inactive, and it then **refuses to start** with `Failed to connect to /run/virtiofsd-*.sock`. Restart all three first. Confirmed the hard way 2026-08-15. navi-vm has no virtiofs, so it never inherits this.