--- title: Separate navi onto its own VM on media type: runbook tags: - media - proxmox aliases: [] related: - [[navi-recon-separation]] - [[fleet-storage-memory-upgrade]] - [[navi]] - [[deployment]] - [[environment]] updated: 2026-08-15 --- # Separate navi onto its own VM on media **Status: Phase 0 complete. Nothing from Phase 1 onward has been run.** Splits [[navi]] out of `recon-vm` (VMID 1130 on data) into its own VM on media, taking **all live data onto local NVMe**. `recon-vm` keeps running throughout — it is never stopped. Background on the entanglement is [[navi-recon-separation]]. ## What moves and why Two datasets are live for [[navi]], on two different slow tiers: | Data | Size | Currently on | Problem | |---|---|---|---| | `nav` | 625 GB | data's SATA SSD (virtiofs) | SATA, and the volume is 93% full | | `planet-dem.pmtiles` | **658 GB** | **pi-nas over NFS** | 7200 rpm spinning disk across 1 GbE | The DEM is the worse of the two. It is read on **every elevation query** (`NAVI_DEM_PMTILES` in `navi-geo.env`) *and* served **directly to browsers** for terrain (`alias /mnt/nas/nav/planet-dem.pmtiles` in `sites-available/dem-origin`). There is already a `dem-cache.conf` in nginx built specifically to paper over the NFS latency. **The remaining ~540 GB on pi-nas is build material and stays there** — `contour-rebuild` 206 GB, `worldcover` 116 GB, `hillshade-work` 94 GB, `planet-latest.osm.pbf` 87 GB, plus dumps and staging. Nothing reads those at runtime. **[[recon]] keeps `kiwix` (138 GB) and `library` (59 GB)** on data. Once `nav` leaves, data's SATA SSD drops from **93% to about 21%**. ## Sizing on media's 1.82 TB | Volume | Virtual | Actual | |---|---|---| | `scsi0` OS (restored) | 180 GB | ~179 GB | | `scsi1` `/data/nav` | 750 GB | 625 GB | | `scsi2` `/data/dem` | 700 GB | 658 GB | | **Total** | **1.63 TB (90%)** | **1.46 TB (80%)** | qcow2 is sparse, so actual consumption is what matters day to day — but the fully-allocated worst case still fits, which is the point of sizing it this way. **Keep contour rebuilds and other build work on pi-nas**; there is not room for them here. Memory: navi-vm at 32 GB puts media at 49 of 62 GB allocated, leaving room for the viewshed budget from [[navi-recon-separation]]. --- ## Phase 0 — done - **106 GB archive** at `pinas-backup:backup/vzdump-qemu-1130-2026_08_15-00_07_11.vma.zst`, verified. Rollback point. - **Baseline 9/9** on the navi suite; map renders headless. - `discard=on` staged on 1130's `scsi0` — applies on its next restart, which this runbook does **not** trigger. `navi_test.sh` is `BASE=` parameterised and checks page shell, health, geocode, place details, landclass, Valhalla height, Valhalla route and both pmtiles range reads. **Run it plus a headless screenshot after every phase.** --- ## Phase 1 — build navi-vm on media (no outage) Restore the archive as VMID 1131. This is a lift — every navi service arrives already configured. ```bash ssh ts-media qmrestore pinas-backup:backup/vzdump-qemu-1130-2026_08_15-00_07_11.vma.zst 1131 \ --storage media-nvme ``` **Do not start it yet** — it is byte-identical to the running VM. ```bash qm set 1131 --delete args # drop virtiofs devices + memory-backend-memfd qm set 1131 --name navi-vm --memory 32768 qm set 1131 --ipconfig0 ip=192.168.1.132/24,gw=192.168.1.1 qm set 1131 --net0 virtio,bridge=vmbr0,link_down=1 qm set 1131 --scsi1 media-nvme:750,format=qcow2,discard=on,ssd=1 qm set 1131 --scsi2 media-nvme:700,format=qcow2,discard=on,ssd=1 qm start 1131 ``` From the console, with the link still down: ```bash tailscale logout; systemctl stop tailscaled rm -f /var/lib/tailscale/tailscaled.state rm -f /etc/machine-id && systemd-machine-id-setup hostnamectl set-hostname navi-vm sed -i '/virtiofs/d;/mnt\/nas/d' /etc/fstab # old share + pi-nas mounts mkfs.ext4 -L navdata /dev/sdb && mkfs.ext4 -L demdata /dev/sdc mkdir -p /data/nav /data/dem cat >> /etc/fstab <<'EOF' LABEL=navdata /data/nav ext4 defaults,noatime,nofail 0 2 LABEL=demdata /data/dem ext4 defaults,noatime,nofail 0 2 /data/nav /mnt/nav none bind,nofail 0 0 EOF mount -a ``` The bind keeps `/mnt/nav` valid so **no navi service config changes** — except the two DEM references, handled in Phase 3. Then bring the NIC up (`qm set 1131 --net0 virtio,bridge=vmbr0`) and re-register Tailscale as a **new** node. ## Phase 2 — copy 1.28 TB (no outage) `recon-vm` keeps serving throughout. ```bash # nav, from data's SATA SSD rsync -aHAX --info=progress2 zvx@100.64.0.24:/mnt/nav/ /data/nav/ # DEM, straight from pi-nas rather than through recon-vm mkdir -p /mnt/pinas && mount -t nfs -o vers=3,ro 192.168.1.245:/export/data /mnt/pinas rsync -aHAX --info=progress2 /mnt/pinas/nav/planet-dem.pmtiles /data/dem/ umount /mnt/pinas ``` Expect **4–6 hours total**. The DEM leg is bounded by pi-nas's spinning disk, not the network. ## Phase 3 — repoint navi at local storage On **navi-vm**, the only config edits in the whole migration: ```bash sed -i 's#/mnt/nas/nav/planet-dem.pmtiles#/data/dem/planet-dem.pmtiles#' \ /etc/navi-backend/navi-geo.env sed -i 's#/mnt/nas/nav/planet-dem.pmtiles#/data/dem/planet-dem.pmtiles#' \ /etc/nginx/sites-available/dem-origin grep -rn "/mnt/nas" /etc/navi-backend/ /etc/nginx/sites-enabled/ # expect nothing nginx -t && systemctl restart nginx navi-geo navi-offroute ``` Strip recon and kiwix from navi-vm — they came along with the clone: ```bash systemctl disable --now recon recon-watchdog kiwix ``` Leave `/opt/recon` on disk for now; deleting it is a Phase 6 task. ## Phase 4 — verify before exposing Test the new VM **directly**, before touching the front door: ```bash BASE=http://:8440 ./navi_test.sh ``` Do not proceed on anything less than **9/9**. Valhalla needs ~1 minute after boot to load 32,396 tiles; 502s before that are expected. ## Phase 5 — cutover (navi outage only, ~2 minutes) `recon-vm` stays **running**. Only navi's [[services]] stop. ```bash # on recon-vm — quiesce navi writers so the delta is consistent systemctl stop navi-geo navi-landclass navi-offroute navi-places \ navi-contacts navi-admin navi-config photon argus-resolver docker stop nominatim-v5 valhalla systemctl stop postgresql@16-main # final delta (nav only; the DEM is static) rsync -aHAX --delete zvx@100.64.0.24:/mnt/nav/ /data/nav/ ``` Start navi's stack on navi-vm, then repoint [[caddy]] on utility CT 101 — **four `reverse_proxy` lines** in the `navi.echo6.co` block (`@authed_api`, `@public_api`, `@valhalla`, and the fallthrough) — at navi-vm on 8440. ```bash pct exec 101 -- caddy validate --config /etc/caddy/Caddyfile pct exec 101 -- systemctl reload caddy ``` [[dns]] and [[authentik]] are untouched. ## Phase 6 — reclaim (DEFERRED — do not run) Only after navi has been healthy on media for **at least a week**: - `rm -rf /mnt/data/nav` on data → SATA SSD 93% → ~21% - `systemctl disable --now virtiofsd-nav` on data - `rm -rf /opt/recon` on navi-vm - Drop 1130's memory from 24 GB to ~8 GB - Optionally remove the 658 GB DEM from pi-nas once the local copy is proven **Keep the source data and VM 1130 intact until then.** That is the rollback path. --- ## Rollback | Stage | Recovery | |---|---| | Phases 1–4 | Nothing changed — navi still served from recon-vm | | After Phase 5 | Revert the four Caddy lines, restart navi's services on 1130 | | Catastrophic | `qmrestore` the Phase 0 archive | --- ## Known traps **`virtiofsd` dies with the VM.** Stopping VM 1130 leaves `virtiofsd-{nav,kiwix,library}.service` inactive, and it then **refuses to start** with `Failed to connect to /run/virtiofsd-*.sock`. Restart all three first. Confirmed the hard way 2026-08-15. navi-vm has no virtiofs, so it never inherits this. **nginx fails on boot.** `satellite-cache.conf` resolves an upstream hostname at parse time; if DNS is not ready, nginx dies and stays dead. It did exactly that from 2026-07-28 to 08-15 — **18 days of navi 502s, unnoticed**. Start it manually after any boot, and fix it properly with a `resolver` directive plus the upstream in a variable. **Do not restart recon-vm.** recon and navi share it until Phase 5 completes, so any restart takes both down.