auto: docs sync 2026-08-15T12:00:04+00:00

Files changed: engine/.embcache.json engine/changelog.md engine/lint-report.md vault/.obsidian/workspace.json vault/docs/hardware/environment.md vault/projects/fleet-storage-memory-upgrade.md vault/projects/navi-recon-separation.md vault/runbooks/corescope-ingest-stall-oom.md vault/runbooks/edge2-access-reference.md vault/runbooks/edge2-boot-recovery.md vault/runbooks/navi-lift-to-media.md vault/runbooks/peertube-sitemap-redis-oom.md
This commit is contained in:
echo6-autocommit 2026-08-15 12:00:04 +00:00
commit 8f95c6e5c3
12 changed files with 166 additions and 152 deletions

View file

@ -1,185 +1,191 @@
---
title: Lift navi off recon-vm onto media
title: Separate navi onto its own VM on media
type: runbook
tags:
- media
- proxmox
- recon
aliases: []
related: []
related:
- [[navi-recon-separation]]
- [[fleet-storage-memory-upgrade]]
- [[navi]]
- [[deployment]]
- [[environment]]
updated: 2026-08-15
---
# Lift navi off recon-vm onto media
# Separate navi onto its own VM on media
**Status: planned, not executed.** Nothing below has been run.
**Status: Phase 0 complete. Nothing from Phase 1 onward has been run.**
[[navi]] and [[recon]] share one VM (`recon-vm`, VMID 1130 on data), so neither can move without the other — see [[navi-recon-separation]]. This lifts the whole VM to media as a clone, then strips each copy down to one platform. No services are rebuilt.
Splits [[navi]] out of `recon-vm` (VMID 1130 on data) into its own VM on media, taking **all live data onto local NVMe**. `recon-vm` keeps running throughout — it is never stopped. Background on the entanglement is [[navi-recon-separation]].
Target is media, which now has 64 GB RAM and a 1.82 TB NVMe volume — see [[fleet-storage-memory-upgrade]].
## What moves and why
Two datasets are live for [[navi]], on two different slow tiers:
| Data | Size | Currently on | Problem |
|---|---|---|---|
| `nav` | 625 GB | data's SATA SSD (virtiofs) | SATA, and the volume is 93% full |
| `planet-dem.pmtiles` | **658 GB** | **pi-nas over NFS** | 7200 rpm spinning disk across 1 GbE |
The DEM is the worse of the two. It is read on **every elevation query** (`NAVI_DEM_PMTILES` in `navi-geo.env`) *and* served **directly to browsers** for terrain (`alias /mnt/nas/nav/planet-dem.pmtiles` in `sites-available/dem-origin`). There is already a `dem-cache.conf` in nginx built specifically to paper over the NFS latency.
**The remaining ~540 GB on pi-nas is build material and stays there** — `contour-rebuild` 206 GB, `worldcover` 116 GB, `hillshade-work` 94 GB, `planet-latest.osm.pbf` 87 GB, plus dumps and staging. Nothing reads those at runtime.
**[[recon]] keeps `kiwix` (138 GB) and `library` (59 GB)** on data. Once `nav` leaves, data's SATA SSD drops from **93% to about 21%**.
## Sizing on media's 1.82 TB
| Volume | Virtual | Actual |
|---|---|---|
| `scsi0` OS (restored) | 180 GB | ~179 GB |
| `scsi1` `/data/nav` | 750 GB | 625 GB |
| `scsi2` `/data/dem` | 700 GB | 658 GB |
| **Total** | **1.63 TB (90%)** | **1.46 TB (80%)** |
qcow2 is sparse, so actual consumption is what matters day to day — but the fully-allocated worst case still fits, which is the point of sizing it this way. **Keep contour rebuilds and other build work on pi-nas**; there is not room for them here.
Memory: navi-vm at 32 GB puts media at 49 of 62 GB allocated, leaving room for the viewshed budget from [[navi-recon-separation]].
---
## Why lift rather than rebuild
## Phase 0 — done
navi is ~20 interlocking services: seven `navi-*` gunicorns, nginx:8440, Valhalla and Nominatim in Docker, Photon (Java), and PostgreSQL 16 whose `data_directory` sits at `/mnt/nav/overture/pgdata`. Standing that up fresh means reproducing every version, path and config. Cloning carries all of it intact; you then delete what doesn't belong.
- **106 GB archive** at `pinas-backup:backup/vzdump-qemu-1130-2026_08_15-00_07_11.vma.zst`, verified. Rollback point.
- **Baseline 9/9** on the navi suite; map renders headless.
- `discard=on` staged on 1130's `scsi0` — applies on its next restart, which this runbook does **not** trigger.
Cost: the clone inherits recon's files too, which you remove afterwards, and it inherits the source VM's identity, which **must** be fixed before it ever touches the network.
`navi_test.sh` is `BASE=` parameterised and checks page shell, health, geocode, place details, landclass, Valhalla height, Valhalla route and both pmtiles range reads. **Run it plus a headless screenshot after every phase.**
---
## Facts this plan relies on
## Phase 1 — build navi-vm on media (no outage)
| Item | Value |
|---|---|
| Source | `recon-vm` VMID 1130 on data, 4 cores, 24 GB, 180 GB disk |
| New VMID | **1131** (next free; 1130 is the highest in use) |
| navi data | `/mnt/data/nav`**625 GB** |
| Front door | Caddy on utility CT 101 → `100.64.0.24:8440` (nginx in recon-vm) |
| Auth | Authentik forward_auth at the Caddy layer — unchanged by this |
| Space needed on media | ~180 GB system + 625 GB data = **~805 GB** of 1.82 TB |
**recon does not call navi.** Grepping `/opt/recon` for navi ports (84xx, 8002, 8010, 2322) and hostnames returns nothing. The split creates no cross-host dependency.
**Stays on data:** `recon.service` (8420), `recon-watchdog`, `kiwix-serve` (8430), and the `library` + `kiwix` shares.
---
## The three traps
**1. Identity collision.** The clone is byte-identical: same static IP `192.168.1.130`, same hostname, same `machine-id`, same Tailscale node key. Booting it on the network unfixed will fight the original for its IP and its tailnet identity. **Fix identity before first network boot** (step 3).
**2. The virtiofs `args:` will not start on media.** `recon-vm` carries raw QEMU args wiring `vhost-user-fs-pci` devices to `/run/virtiofsd-{nav,kiwix,library}.sock`, plus a `memory-backend-memfd` object that virtiofs requires. Those sockets do not exist on media, so the clone fails to start as-is.
Rather than recreate virtiofsd on media, **give the new VM a real virtual data disk**. That makes it self-contained — no host-side sockets, no raw args — and therefore movable with [[pve-guest-park-and-adopt]], which the current VM is not.
**3. PostgreSQL moves with the data.** `pgdata` lives inside `/mnt/nav`, so it travels with the 625 GB. Nothing separate to dump and restore, but the copy must be consistent — stop postgres before the final sync (step 5).
---
## 1. Back up the source
```bash
ssh ts-data
vzdump 1130 --mode snapshot --storage pinas-backup --compress zstd
```
`--mode snapshot` runs live — no downtime. This is both the rollback point and the clone source. Note the archive name.
## 2. Restore as VMID 1131 on media
Restore the archive as VMID 1131. This is a lift — every navi service arrives already configured.
```bash
ssh ts-media
qmrestore pinas-backup:backup/vzdump-qemu-1130-<ts>.vma.zst 1131 --storage media-nvme
qmrestore pinas-backup:backup/vzdump-qemu-1130-2026_08_15-00_07_11.vma.zst 1131 \
--storage media-nvme
```
`media-nvme` is the 1.82 TB volume. **Do not start it yet.**
## 3. Fix identity before it ever boots on the network
**Do not start it yet** — it is byte-identical to the running VM.
```bash
qm set 1131 --name navi-vm
qm set 1131 --delete args # drop virtiofs devices + memory-backend-memfd
qm set 1131 --name navi-vm --memory 32768
qm set 1131 --ipconfig0 ip=192.168.1.132/24,gw=192.168.1.1
qm set 1131 --memory 32768 # navi + viewshed headroom
qm set 1131 --delete args # drop virtiofs wiring
```
Removing `args` also removes the `memory-backend-memfd` object, which is only needed for virtiofs.
Then boot **with the NIC detached**, and from the console clear the inherited identity:
```bash
qm set 1131 --net0 virtio,bridge=vmbr0,link_down=1
qm set 1131 --scsi1 media-nvme:750,format=qcow2,discard=on,ssd=1
qm set 1131 --scsi2 media-nvme:700,format=qcow2,discard=on,ssd=1
qm start 1131
# on the console:
tailscale logout && systemctl stop tailscaled
```
From the console, with the link still down:
```bash
tailscale logout; systemctl stop tailscaled
rm -f /var/lib/tailscale/tailscaled.state
hostnamectl set-hostname navi-vm
rm -f /etc/machine-id && systemd-machine-id-setup
hostnamectl set-hostname navi-vm
sed -i '/virtiofs/d;/mnt\/nas/d' /etc/fstab # old share + pi-nas mounts
mkfs.ext4 -L navdata /dev/sdb && mkfs.ext4 -L demdata /dev/sdc
mkdir -p /data/nav /data/dem
cat >> /etc/fstab <<'EOF'
LABEL=navdata /data/nav ext4 defaults,noatime,nofail 0 2
LABEL=demdata /data/dem ext4 defaults,noatime,nofail 0 2
/data/nav /mnt/nav none bind,nofail 0 0
EOF
mount -a
```
Only then reattach the NIC (`qm set 1131 --net0 virtio,bridge=vmbr0`) and re-register Tailscale as a new node.
The bind keeps `/mnt/nav` valid so **no navi service config changes** — except the two DEM references, handled in Phase 3.
## 4. Attach and prepare the data disk
Then bring the NIC up (`qm set 1131 --net0 virtio,bridge=vmbr0`) and re-register Tailscale as a **new** node.
## Phase 2 — copy 1.28 TB (no outage)
`recon-vm` keeps serving throughout.
```bash
qm set 1131 --scsi1 media-nvme:700,format=qcow2
# inside the guest:
mkfs.ext4 -L nav /dev/sdb
mkdir -p /mnt/nav
echo "LABEL=nav /mnt/nav ext4 defaults,noatime,nofail 0 2" >> /etc/fstab
mount /mnt/nav
# nav, from data's SATA SSD
rsync -aHAX --info=progress2 zvx@100.64.0.24:/mnt/nav/ /data/nav/
# DEM, straight from pi-nas rather than through recon-vm
mkdir -p /mnt/pinas && mount -t nfs -o vers=3,ro 192.168.1.245:/export/data /mnt/pinas
rsync -aHAX --info=progress2 /mnt/pinas/nav/planet-dem.pmtiles /data/dem/
umount /mnt/pinas
```
700 GB gives the 625 GB room to grow. Also remove the three virtiofs lines from the guest's `/etc/fstab`, which will otherwise fail on boot.
Expect **46 hours total**. The DEM leg is bounded by pi-nas's spinning disk, not the network.
## 5. Copy the 625 GB
## Phase 3 — repoint navi at local storage
First pass live, while navi keeps serving from data:
On **navi-vm**, the only config edits in the whole migration:
```bash
rsync -aHAX --info=progress2 zvx@100.64.0.24:/mnt/nav/ /mnt/nav/
sed -i 's#/mnt/nas/nav/planet-dem.pmtiles#/data/dem/planet-dem.pmtiles#' \
/etc/navi-backend/navi-geo.env
sed -i 's#/mnt/nas/nav/planet-dem.pmtiles#/data/dem/planet-dem.pmtiles#' \
/etc/nginx/sites-available/dem-origin
grep -rn "/mnt/nas" /etc/navi-backend/ /etc/nginx/sites-enabled/ # expect nothing
nginx -t && systemctl restart nginx navi-geo navi-offroute
```
Expect 23 hours over 1 GbE. Then stop the writers on the source and take a final delta:
Strip recon and kiwix from navi-vm — they came along with the clone:
```bash
# on recon-vm
systemctl disable --now recon recon-watchdog kiwix
```
Leave `/opt/recon` on disk for now; deleting it is a Phase 6 task.
## Phase 4 — verify before exposing
Test the new VM **directly**, before touching the front door:
```bash
BASE=http://<navi-vm-tailnet-ip>:8440 ./navi_test.sh
```
Do not proceed on anything less than **9/9**. Valhalla needs ~1 minute after boot to load 32,396 tiles; 502s before that are expected.
## Phase 5 — cutover (navi outage only, ~2 minutes)
`recon-vm` stays **running**. Only navi's [[services]] stop.
```bash
# on recon-vm — quiesce navi writers so the delta is consistent
systemctl stop navi-geo navi-landclass navi-offroute navi-places \
navi-contacts navi-admin navi-config photon argus-resolver
docker stop nominatim-v5 valhalla
systemctl stop postgresql@16-main
# on navi-vm
rsync -aHAX --delete zvx@100.64.0.24:/mnt/nav/ /mnt/nav/
# final delta (nav only; the DEM is static)
rsync -aHAX --delete zvx@100.64.0.24:/mnt/nav/ /data/nav/
```
The postgres stop is what makes `pgdata` consistent. Do not skip it.
## 6. Strip each side
On **navi-vm** (1131) — remove recon and kiwix:
Start navi's stack on navi-vm, then repoint [[caddy]] on utility CT 101 — **four `reverse_proxy` lines** in the `navi.echo6.co` block (`@authed_api`, `@public_api`, `@valhalla`, and the fallthrough) — at navi-vm on 8440.
```bash
systemctl disable --now recon recon-watchdog kiwix
rm -rf /opt/recon
```
On **recon-vm** (1130) — remove navi and the geo stack:
```bash
systemctl disable --now navi-geo navi-landclass navi-offroute navi-places \
navi-contacts navi-admin navi-config photon argus-resolver postgresql@16-main
docker rm -f nominatim-v5 valhalla
```
Leave `/mnt/data/nav` on data untouched until step 9.
## 7. Cutover — one line
In the Caddyfile on utility CT 101, inside the `navi.echo6.co` block, repoint every `reverse_proxy` from `100.64.0.24:8440` to the new VM's tailnet address on 8440. There are four occurrences (`@authed_api`, `@public_api`, `@valhalla`, and the fallthrough `handle`).
```bash
ssh ts-utility
pct exec 101 -- caddy validate --config /etc/caddy/Caddyfile
pct exec 101 -- systemctl reload caddy
```
DNS and Authentik are untouched — the hostname and auth flow do not change.
[[dns]] and [[authentik]] are untouched.
## 8. Verify
## Phase 6 — reclaim (DEFERRED — do not run)
Confirm navi answers through the front door, not just that the VM runs — check the map loads, a geocode returns, and a route computes. Follow [[headless-browser-page-verification]] for the UI. Then confirm recon is still healthy on 8420 and kiwix on 8430.
Only after navi has been healthy on media for **at least a week**:
## 9. Reclaim — a week later, not the same day
- `rm -rf /mnt/data/nav` on data → SATA SSD 93% → ~21%
- `systemctl disable --now virtiofsd-nav` on data
- `rm -rf /opt/recon` on navi-vm
- Drop 1130's memory from 24 GB to ~8 GB
- Optionally remove the 658 GB DEM from pi-nas once the local copy is proven
Once navi has been healthy on media for several days:
```bash
# on data
rm -rf /mnt/data/nav
systemctl disable --now virtiofsd-nav
qm set 1130 --memory 8192 # recon alone needs far less than 24 GB
```
That drops `/mnt/data` from 93% to roughly 21% and frees ~16 GB of RAM on data — which matters, since data is hard-capped at 32 GB.
**Keep the source data and VM 1130 intact until then.** That is the rollback path.
---
@ -187,8 +193,16 @@ That drops `/mnt/data` from 93% to roughly 21% and frees ~16 GB of RAM on data
| Stage | Recovery |
|---|---|
| Any point before step 7 | Nothing changed for users — navi still serves from data |
| After cutover | Revert the four Caddy lines to `100.64.0.24:8440`, restart navi's services on recon-vm |
| Catastrophic | `qmrestore` the step 1 archive from `pinas-backup` |
| Phases 14 | Nothing changed — navi still served from recon-vm |
| After Phase 5 | Revert the four Caddy lines, restart navi's services on 1130 |
| Catastrophic | `qmrestore` the Phase 0 archive |
The rollback path stays open until step 9 deletes the source data. Do not run step 9 early.
---
## Known traps
**`virtiofsd` dies with the VM.** Stopping VM 1130 leaves `virtiofsd-{nav,kiwix,library}.service` inactive, and it then **refuses to start** with `Failed to connect to /run/virtiofsd-*.sock`. Restart all three first. Confirmed the hard way 2026-08-15. navi-vm has no virtiofs, so it never inherits this.
**nginx fails on boot.** `satellite-cache.conf` resolves an upstream hostname at parse time; if DNS is not ready, nginx dies and stays dead. It did exactly that from 2026-07-28 to 08-15 — **18 days of navi 502s, unnoticed**. Start it manually after any boot, and fix it properly with a `resolver` directive plus the upstream in a variable.
**Do not restart recon-vm.** recon and navi share it until Phase 5 completes, so any restart takes both down.