echo6-docs/vault/projects/fleet-patch-audit.md
echo6-autocommit 54652f31ea auto: docs sync 2026-06-22T12:00:09+00:00
Files changed: engine/changelog.md engine/lint-report.md vault/.obsidian/workspace.json vault/projects/fleet-patch-audit.md vault/projects/fleet-platform-baseline.md
2026-06-22 12:00:09 +00:00

761 lines
87 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
title: Fleet Patch Audit — 2026-06-19
type: project
tags:
- proxmox
- ai
related: []
updated: 2026-06-22
status: complete
---
# Fleet Patch Audit — 2026-06-19
Read-only audit snapshot as of 2026-06-19. **Nothing has been applied — this is a planning document to build the patch plan from.**
**Topology note:** the old Contabo VPS has been rebuilt as **edge1 (mail-only)**; **edge2 is now the front door for everything else**. edge1 is excluded from this audit (mid-rebuild/maintenance). **Headscale:** edge2 CT107 is the main fleet tailnet (34 nodes, `vpn.echo6.co`, self-hosted Headscale 0.28.0); utility CT106 is a separate IdahoMesh sub-tailnet (`vpn.idahomesh.com`, 3 nodes, low-risk). No services route through old-Contabo. **Mailcow CT108:** destroyed 2026-06-20 (`pct destroy 108 --purge`); backup preserved durably on pi-nas (`…/contabo-prewipe-2026-06/mailcow/`, sha256-verified); live mail on edge1 (MX/A for mail.echo6.co → 5.189.158.149).
---
## Campaign complete (2026-06-22)
**Phases 13 are fully done. The entire fleet is on the current platform.**
- **Phase 1** (guest/VM security apt) — COMPLETE 2026-06-20. 26 guests patched, ~600+ security packages cleared, zero data loss.
- **Phase 2** (app/container updates) — COMPLETE 2026-06-21. All app upgrades done: Authentik 2025.12.4→2026.5.3 (sequential), Forgejo 14→15, Headscale 0.28→0.29.1 (both instances), Immich 2.5.6→2.7.5, Nextcloud AIO→NC 33.0.5, media stack (Jellyfin/SABnzbd/arr), cortex AI stack (Ollama/TEI/Qdrant/Open-WebUI), and the low-urgency batch.
- **Phase 3** (platform/reboot windows) — COMPLETE 2026-06-22. All 5 PVE nodes on 9.2.3/kernel 7.0.12-1-pve (including toc+cortex); pi-nas on OMV 8.4/kernel 6.18; cortex NVIDIA driver 580.167.08 + DKMS + nvidia-container-toolkit 1.19.1; GPU passthrough (vfio) survived the 7.0 kernel; cluster 5/5 quorate.
**Intentionally deferred / out of scope (not failures):**
- **RabbitMQ** — left on 3.12 by decision; OTS does not support 4.x and exposure is localhost-bound. Revisit only if/when OTS officially supports RabbitMQ 4.x.
- **Nominatim v5** — full re-import project, spun off to [[nominatim-v5-reimport]].
- **Optional cosmetic cleanups:** navidrome cert renewal (`navidrome.echo6.co` expired unmanaged cert); MediaMTX deprecated config param rename (`protocols``rtspTransports`, `encryption``rtspEncryption`); retained `/root` rollback artifacts (DB dumps, binary backups on their CTs); vestigial utility exit-node route (0.0.0.0/0, harmless); `rpi-eeprom` still held on pi-nas (SPI bootloader, intentionally untouched).
---
## Prioritized Backlog
### Tier 1 — Security-Urgent Guest OS
These containers have the highest raw security-update counts and have not been patched recently (or never). Address before any platform work.
| Host | Guest | Upgradable / Security | Notes |
|------|-------|-----------------------|-------|
| utility | CT119 mesh-territory | 179 / 91 sec | Never patched |
| utility | CT108 meshai | 109 / 79 | — |
| cloud | CT120 immich guest-OS | 191 / 101 | — |
| cloud | CT121 nextcloud guest-OS | 98 / 75 | — |
| media | CT110 peertube | 81 / 37 | — |
| utility | CT109 opentakserver | 34 / 31 | — |
| utility | CT104 central | 50 / 38 | Includes PostgreSQL 16.13 → 16.14 |
### Tier 2 — App / Container Updates
Updates where the application or its Docker images have drifted from current upstream, ordered roughly by operational risk.
| Scope | Guest | Item | Notes |
|-------|-------|------|-------|
| edge2 | CT105 | authentik 2025.12.4 → 2026.5.3 | **#1 security item** — 7 CVEs + 5 GHSAs in gap; sequential upgrade (min: 2025.12.6) |
| edge2 | CT107 | headscale 0.28.0 → 0.29.1 | Also a 2nd headscale on utility CT106 |
| edge2 | CT103 | forgejo 14.0.5 → 15.0.3 | **14.x EOL 2026-04-30** — migrate branch, not just patch |
| edge2 | CT106 | Synapse 1.155.0 / Element / MAS | Image drift + pending OS apt security updates |
| edge2 | CT104 | livesync couchdb:3.4 | Docker image drift |
| edge2 | CT108 | ~~mailcow (18 containers)~~ | ✅ **Decommissioned 2026-06-20** — superseded by edge1; no longer an update target |
| cloud | CT120 | immich — server/ml/valkey:9/postgres(14-vectorchord) | 4 images drifted |
| cloud | CT121 | nextcloud AIO — mastercontainer + NC app 32.0.4 | Mastercontainer behind; 12-container stack |
| cortex | VM150 | ollama / tei(1.7) / qdrant / open-webui / obsidian | 5 AI containers drifted |
| media | VM105 | arr stack — 8 containers | jellyfin/sonarr/radarr/prowlarr/sabnzbd/lidarr/navidrome/jellyseerr all :latest |
### Tier 3 — Platform / Reboot Windows
Reboot-bearing updates. Coordinate maintenance windows carefully. cortex and toc are PROTECTED hosts.
| Scope | Item | Detail |
|-------|------|--------|
| data, utility, cloud, media, toc (PVE 9 nodes) | PVE 9.1.1 → 9.2 | Reboot required |
| data, utility, cloud, media, toc | QEMU 10 → 11 | Reboot required |
| data, utility, cloud, media, toc | LXC 6 → 7 | Reboot required |
| data, utility, cloud, media, toc | Kernel 6.17.2 → 6.17.13 | Reboot required |
| edge2 | PVE 8.4.19 | Already fully patched — no action needed |
| cortex VM150 (PROTECTED) | NVIDIA driver 580.159 → 580.167 + DKMS | Reboot required |
| cortex VM150 (PROTECTED) | nvidia-container-toolkit 1.18 → 1.19 | — |
| pi-nas | Kernel 6.12.62 → 6.18.34+rpt-rpi-2712 | ✅ DONE 2026-06-22 — operator present; `/boot` backup at `/root/boot-backup-pre6.18-20260622.tar.gz`; booted clean; rpi-eeprom held |
| pi-nas | OMV 8.1 → 8.4 | ✅ DONE 2026-06-21 — 8.4.0-3 |
### Cross-Cutting (All / Most Hosts)
| Item | Detail |
|------|--------|
| Tailscale | 1.94 → 1.98 nearly everywhere |
| Docker CE | → 29.6 wherever Docker is installed |
---
## Non-Update Flags
Issues noted that are not package/image updates but warrant attention.
| Host / Guest | Flag | Detail |
|---|---|---|
| data | Disk 92% full | ~73 GB / 938 GB free; address before patching |
| utility CT118 archivist | rpcbind on 0.0.0.0:111 | No Tailscale client or firewall on this CT; exposed port |
| media VM105 jellyseerr | Non-stable image | Running preview-OIDC tag, not a stable release |
| data VM1130 nominatim | Stale image (14 months) | nominatim:4.5, pinned; confirm intentional |
| edge2 CT106 matrix / CT107 headscale | No Tailscale client | Ingress via Caddy; verify internal routing before patching |
---
## Application Currency (2026-06-19)
Running application version vs latest stable upstream, per app — the "is everything up to date?" view (distinct from OS-package/image-drift above). Read-only. **Caveat:** version deltas were measured directly from the running apps and are reliable; the specific CVE/advisory IDs below were surfaced by the audit agents and are **post-knowledge-cutoff — verify against primary advisories before acting on them.** GitHub's anonymous API rate-limited a few "latest" lookups (noted inline).
### Security-relevant — prioritize
| App | Where | Running → Latest | Behind | Security note (verify) |
|-----|-------|------------------|--------|------------------------|
| **Authentik** | edge2 CT105 | 2025.12.4 → **2026.5.3** | ~6 mo / 5 majors | **#1** — claimed 7 CVEs + 5 GHSAs across two May-2026 waves. Min-disruption: 2025.12.6 (same branch, backported fixes); full: 2026.5.3. SSO — upgrade sequentially. |
| **Immich** | cloud CT120 | 2.5.6 → **2.7.5** | 9 rel | shared-link ACL bypass (2.6.0) + stored XSS panorama viewer (2.7.0) |
| **Valkey** | cloud CT120 | 9.0.2 → **9.1.0** | 3 patch | 6 CVEs, "SECURITY" urgency (use-after-free, DoS, RESP injection) |
| **RabbitMQ** | utility CT109 | 3.12.1 → **4.2.8** | EOL major | 3.x abandoned upstream; 4 advisories 2026-06-18, no 3.x backports |
| **Nominatim** | data/recon-vm | 4.5.0 → **5.3.2** | ~19 mo, major | v4→v5 data-model break; needs image swap + full re-import (pair w/ Photon 1.2.0) |
| **Ollama** | cortex | 0.16.1 → **0.30.10** | 14 minor | history of SSRF / path-traversal CVEs |
| **Nextcloud** | cloud CT121 | 32.0.4 → **32.0.11**; AIO ~v12.5 → **13.2.1** | 7 patch / ~8 AIO | resource-exhaustion fix (32.0.10); AIO image ~5 mo stale |
| **Forgejo** | edge2 CT103 | 14.0.5 → **15.0.3** | EOL branch | 14.x EOL 2026-04-30; no further backports — migrate to 15.0.3 (LTS) |
| **SABnzbd** | media VM105 | 4.5.5 → **5.0.4** | major | 5.x security hardening over 4.x |
| **Jellyfin** | media VM105 | 10.11.6 → **10.11.11** | 5 patch | 10.11.x patch cycle has carried CVE fixes |
| **Jellyseerr** | media VM105 | **preview-OIDC dev** → 3.3.0 | unreleased | running unpinned dev code — no stable security posture |
| **PeerTube** | media CT110 | 8.0.2 → **8.2.1** | 2 patch | 8.x patch releases include security fixes |
| **Open-WebUI** | cortex | 0.8.1 → **0.9.6** | ~1.5 minor | touches auth/session surface |
| **MediaMTX** | utility CT109 | 1.13.0 → **1.19.1** | 6 minor | RTSP/WebRTC streaming attack surface |
| **Headscale** ×2 | utility CT106 + edge2 CT107 | 0.28.0 → **0.29.1** | 1 minor | upgrade-guide-required; two separate instances; no security flag noted |
### Operational / host-level (not a simple app bump)
| Item | Where | Finding |
|------|-------|---------|
| **Mailcow** | edge2 CT108 | ✅ **DECOMMISSIONED 2026-06-20** — destroyed (`pct destroy 108 --purge`); backup on pi-nas, live mail on edge1. |
| **Host kernel** | **edge2 host** | Agent flagged DirtyFrag (CVE-2026-43284/-43500) + copy.fail (CVE-2026-31431, claimed CISA KEV) as host-kernel LPE. **Tension:** the host audit showed edge2 fully patched (0 upgradable) on its repo — **verify** whether these need a kernel newer than the no-subscription repo provides. |
### Current / already past the fix (no action)
Vaultwarden 1.36.0 (edge2 CT102 — has the SSO-takeover/org-access CVE fixes) · PDM 1.1.4 (edge2 CT100 — past the RCE PSA) · WordPress 7.0 core + all plugins/themes (edge2 CT101) · Synapse 1.155.0 / Element / MAS (edge2 CT106 — current, only minor `:latest` digest drift) · obsidian-remote v1.12.7 (cortex) · PostgreSQL 16.14 (recon-vm).
### Lower urgency
Mumble 1.5.517→1.5.901 · Caddy 2.10.2/2.11.3→2.11.4 · Qdrant 1.16.3→1.18.2 · TEI 1.7.4→1.9.3 · Valhalla 3.6.3→3.7.0 · Photon 1.1.0→1.2.0 · kiwix 3.7.0→3.8.2 · CouchDB 3.4.3→3.5.2 (livesync) · Navidrome 0.60.3→0.62.0 · Sonarr/Radarr/Prowlarr/Lidarr 12 versions · NATS 2.14.0→2.14.2 · PostgreSQL 16.12/16.13→16.14 · meshmonitor (~1 mo, exact ver undeterminable) · searxng (rolling, ~4.5 mo) + valkey-8 sidecar 8.1.5→8.1.8 · mautrix-signal v0.2603.0.
### Internal echo6 apps (no upstream to track)
central-*, meshai, archivist, meshwars, recon / recon-watchdog, navi-* — running; version = current git head.
---
## Proposed Patch Approach (NOT executed — for planning)
Lowest-risk changes first; everything reboot-bearing deferred to scheduled windows.
**Principles**
- **Security-pocket apt only** in the first pass (openssl, openssh, gnutls, krb5, libc6, samba, nghttp2, etc.). No PVE 9.2 / QEMU / LXC / kernel, no Docker image pulls, no OMV/NVIDIA, no reboots.
- **Protected hosts (cortex, toc) never go in a bulk pass** — handled individually in their own window. Note: toc hosts cortex (VM150), so a toc reboot drops cortex — the two must be coordinated together.
- **edge1 (mail)** excluded while it's mid-rebuild.
- `needrestart` will bounce affected daemons after glibc/openssl upgrades — seconds of blip per guest, no data risk.
**Phased plan**
| Phase | Scope | Reboot? | Notes |
|------|-------|---------|-------|
| **1 — Guest/VM security apt** | utility CT100,101,102,103,104,106,107,108,109,112,118,119 · cloud CT120,121 · media VM105,CT110,CT111 · data VM1130 · edge2 CT100107 | No | Lowest blast radius. Worst-first: CT119, CT108, immich/nextcloud guest-OS, peertube. |
| **2 — Hypervisor host OS security** | data, utility, cloud, media host OSes (**not toc**) | No | One node at a time; restarts hypervisor-side daemons (smbd etc.). edge2 host already patched. |
| **3 — App / container updates (Tier 2)** | per-app, native updater each | Per-app | **COMPLETE 2026-06-21** — all app upgrades done (security-critical, low-urgency batch, and cortex AI stack). See Phase 2 Execution Log. |
| **4 — Reboot windows (Tier 3)** | PVE 9.2 + QEMU 11 + LXC 7 + kernel on the 5 PVE-9 nodes; pi-nas kernel + OMV; cortex NVIDIA/DKMS | **Yes** | **COMPLETE (2026-06-22)** — all 5 PVE nodes on 9.2.3/kernel 7.0.12; pi-nas on OMV 8.4/kernel 6.18; cortex NVIDIA 580.167/DKMS/toolkit 1.19.1; GPU passthrough verified. See Phase 3 Execution Log. |
| **Cross-cutting** | Tailscale 1.94→1.98 fleet-wide | No | Can ride along Phase 1/2. |
**Special handling — do NOT bulk-patch these; use the native updater**
- **mailcow** (edge2 CT108) → ✅ decommissioned 2026-06-20, no action
- **nextcloud AIO** (cloud CT121) → update mastercontainer, then in-UI update button (port 8080)
- **immich** (cloud CT120) → `docker compose pull && up -d` from its compose dir
- **authentik** (edge2 CT105) → sequential version upgrades with migrations; 2025.12.4 → 2026.2.0 cannot skip releases
- **pi-nas OMV** → OMV's own update path, not raw apt
- **cortex / toc** → manual, protected, own window
---
## Open Decisions
### Resolved (2026-06-19/20)
1. **RabbitMQ 3.12→4.x****ACCEPTED** as a decoupled sub-task (Erlang ≥26 → 3.13.x → enable feature-flags → 4.x); the OpenTAKServer bump does NOT cover it. In scope for this campaign.
2. **Break-glass****VALIDATED**: all 7 LAN SSH paths work independent of the tailnet; every PVE node (incl. edge2) has local `pam`/`pve` realms; Authentik `akadmin`, Forgejo `matt`, Nextcloud `admin` local logins confirmed. ⚠️ **One gap → PDM (edge2 CT100)** has only `openid:authentik` in `domains.cfg`; confirm `root@pam` login at `https://100.64.0.28:8443` (or add a `pam:` stanza) **before** the Authentik upgrade.
3. **Phase scope****ALL PHASES**; no-reboot work (Phase 12) first, platform/reboot (Phase 3) scheduled separately.
4. **Daemon-restart tolerance****DOWNTIME ACCEPTED**; needrestart blips OK on the stateful guests (still take logical DB dumps for safety, but no holding restarts for availability).
5. **edge2 host kernel (DirtyFrag/copy.fail)****NO ACTION / NO REBOOT**: running `6.8.12-30-pve`, the newest its repos offer; no update available, no reboot-required flag. Not actionable without a repo/branch change. edge2 needs no Phase 3 reboot.
6. **CVE verification****NOT GATING**: being behind on versions is sufficient justification; post-cutoff CVE IDs are not chased or relied on for ordering.
7. **Maintenance windows****NO CONSTRAINT** (single user); disruptive steps may run anytime, no scheduling/announcement needed.
8. **Nominatim 4→5****DEFERRED / OUT OF SCOPE**: spun off to [[nominatim-v5-reimport]] as its own project.
9. **Headscale location****RESOLVED**: fleet control plane = edge2 CT107 (`vpn.echo6.co`, 34 nodes — high-risk); IdahoMesh sub-tailnet = utility CT106 (`vpn.idahomesh.com`, 3 nodes — low-risk). No Contabo routing.
10. **mailcow CT108****RESOLVED**: superseded by edge1; safe to decommission (pending backup-off-CT verify below).
11. **Phase 2 host-OS sequencing****RESOLVED**: folded into Phase 3 reboot window.
12. **In-scope app campaign** — Forgejo 14→15, Matrix/Synapse, Authentik chain, RabbitMQ, media stack, immich, nextcloud, headscale, cortex AI. (Nominatim excluded per #8.)
### Pre-flight gates — status (2026-06-20)
- **PDM break-glass** — ✅ **DONE**: local `admin@pam` (Administrator) created on PDM CT100, login verified via API, Authentik realm untouched; cred in `credentials`, config backup at CT100 `/root/access.bak-2026-06-20`.
- **CT118 archivist rpcbind** — ✅ **DONE**: nftables rule restricts port 111 to source `192.168.1.240` (NFS server) only; NFS mount healthy, ruleset persisted. Now 100% LAN-internal.
- **mailcow CT108** — ✅ **DONE**: backup durable on pi-nas (`…/contabo-prewipe-2026-06/mailcow/`, sha256-verified), live mail confirmed on edge1; **CT108 destroyed (`pct destroy 108 --purge`, 2026-06-20)** — config + disk image purged. edge2 now hosts CT100107.
- **data disk (92%)** — ⏏️ **DE-SCOPED**: not a gate. OS package updates are roll-back-able (reinstall prior version); the migration-heavy apps that need real rollback (Authentik, Forgejo, Nextcloud, PeerTube) live on cloud/edge2, not `data`. Optional cleanup only (~6 GB obviously-safe: stale ISO, zimit temp) if ever wanted.
**Rollback model (corrected):** OS packages → reinstall the prior version (no VM snapshot needed). App DB-migration upgrades (Authentik/Forgejo/Nextcloud/PeerTube) → restore a quiesced DB dump (cheap), since reinstalling the old binary won't unwind a migrated schema.
---
## Phase 1 — Execution Log (2026-06-20)
**Status: COMPLETE.** 26 guests patched (security-pocket only), every service validated back up, ~600+ security packages cleared, zero data loss. No host reboots (guests only). Method: snapshot-or-DB-dump → security-only apt (`apt-get install --only-upgrade` from the security pocket) → `pct reboot` → service-by-service validation.
### Guests patched
- **utility:** CT100, CT101, CT102, CT103, CT104, CT106, CT107, CT108, CT109, CT112, CT118, CT119
- **cloud:** CT120, CT121
- **media:** VM105, CT110, CT111
- **data:** VM1130
- **edge2:** CT100, CT101, CT102, CT103, CT104, CT105, CT106, CT107
(CT107 = headscale — see Critical Finding 1 below.)
### DB dumps / rollback points captured (reusable in Phase 2)
| Guest | Location | Notes |
|---|---|---|
| authentik CT105 | `/root/authentik-db-20260620.sql` (126 MB) | — |
| matrix CT106 | `/root/matrix-db-20260620.sql` (60 MB) | — |
| forgejo CT103 | `/root/forgejo-db-20260620.sql` | — |
| wordpress CT101 | dump taken on CT101 | — |
| peertube CT110 | DB dump | — |
| opentakserver CT109 | PG dump + LVM snapshot | — |
| central CT104 | PG dump + snapshot | — |
| immich CT120 | `/root/immich-predates-20260620.sql` | — |
| headscale CT107 | `/root/headscale-db-20260620.sqlite` | — |
**Note:** many edge2 and cloud CTs' storage does NOT support `pct snapshot` — DB dump or package-reinstall is the rollback model for those.
### Incidental fixes made during Phase 1
- **(a) CT110 peertube — immutable `/etc/resolv.conf` blocked reboot.** The file had `chattr +i` set (intentional NordVPN DNS protection). Cleared the immutable flag to allow the reboot, then verified the flag was restored and DNS remained healthy after boot.
- **(b) CT111 mcc — DNS hijacked to unreachable MagicDNS.** Tailscale `accept-dns` was redirecting DNS to a MagicDNS address that was not reachable from this CT. Disabled `tailscale accept-dns`, set `1.1.1.1` / `8.8.8.8` persistently.
- **(c) PostgreSQL on central CT104 moved 16.13→16.14** as part of the security-pocket apt pass.
- **(d) Fleet-wide stale `/etc/hosts` fix** — see Critical Finding 2.
### ✅ Critical Finding 1 — CT107 headscale reboot-survival — FIXED 2026-06-20
On `pct reboot 107`, the fleet headscale coordinator crash-looped and required approximately 15 minutes of manual recovery. Two root causes:
- **(a) Docker compose bridge `headscale_default` came up `linkdown`** after the unprivileged-LXC reboot.
- **(b) Bootstrap chicken-and-egg:** Docker port-binds headscale to `100.64.0.38` (tailscale0), but tailscale0 needs headscale (the coordinator) to come up first.
Recovery required: `compose down``docker restart` → manually add `100.64.0.38/32` to tailscale0 → `compose up` → re-auth CT107's own tailscale node with a fresh preauthkey → temporarily DNAT `vpn.echo6.co` through CT107's internal IP to bootstrap, then revert Caddy config and clean up iptables.
**Resolution (2026-06-20):** Root cause was a triple boot deadlock: (a) headscale/headplane Docker ports were bound to the tailscale IP `100.64.0.38` (which only exists after tailscaled connects to headscale — circular); (b) the `headscale-stack.service` systemd unit waited on `tailscale-online.target` (also needs headscale up); (c) `only_start_if_oidc_is_available: true` blocked startup on Authentik reachability.
Fix applied (all originals backed up `.bak-20260620`):
1. Rebound both ports `100.64.0.38:8084/3100``10.10.10.25:8084/3100` (CT's static internal IP, always up at boot) in `/opt/headscale/docker-compose.yml`.
2. `headscale-stack.service``After=docker.service` only (dropped `tailscale-online.target` dependency).
3. `only_start_if_oidc_is_available: false` in `/opt/headscale/config.yaml`.
4. Repointed edge2 `/etc/caddy/Caddyfile` `vpn.echo6.co` upstream (and `/admin*` headplane) to `10.10.10.25`.
**PROVEN by a controlled `pct reboot 107`: headscale self-healed in ~45 seconds with ZERO manual intervention**`/health` 200, RestartCount 0, clean logs, 32 fleet nodes online.
Known minor leftover (cosmetic, non-blocking): CT107's OWN tailscale client node (100.64.0.38) stays in NoState because its tailscaled ControlURL points at its own IP (`http://100.64.0.38:8084`) — a self-bootstrap chicken-and-egg. The headscale SERVICE is fully healthy and the fleet is coordinated; CT107 is managed via the edge2 host OOB path, so its own tailnet membership is optional. Candidate follow-up: point CT107 tailscaled ControlURL at `https://vpn.echo6.co` (now always reachable via Caddy→10.10.10.25) so it self-registers cleanly.
### 🔴 Critical Finding 2 — fleet-wide stale /etc/hosts broke coordinator connectivity
7 fleet nodes — data, cloud, media, utility hosts, plus caddy CT101, cobalt CT112, and peertube CT110 — had a stale `5.189.158.149 vpn.echo6.co` line in `/etc/hosts` left over from before the 2026-06-19 headscale migration to edge2. This pinned `vpn.echo6.co` to edge1 (now mail-only), so tailscaled hit Mailcow's TLS cert and could never reach the real coordinator — affected nodes showed OFFLINE in headscale while coasting on persistent WireGuard tunnels (still SSH-reachable, masking the problem).
**Fixed 2026-06-20:** removed the stale line and ran `tailscale up` on all 7; all confirmed ONLINE in the coordinator. `/etc/hosts.bak-20260620` backups left on each host.
**IMPACT on Phase 3:** this class of bug means a reboot of an affected host while its control connection was stale would have failed to rejoin the tailnet — the primary lockout risk. Now that all nodes have valid control connections, reboots should re-register cleanly. However: **verify each node is ONLINE in headscale before AND after any Phase-3 reboot.**
Headscale stale-node cleanup: deleted dead nodes `mailcow` (destroyed CT108) and a stale peertube duplicate, both on 2026-06-20.
---
## Phase 2 — Execution Log (2026-06-21)
**Status: COMPLETE (2026-06-21).** All app upgrades done — high-priority, security-critical, low-urgency batch, and cortex AI stack. Only Phase 3 (platform/reboots) and the deferred Nominatim project remain.
### Completed upgrades
All targets below were validated healthy post-upgrade; rollback DB dumps retained on the respective CTs under `/root`.
**Wave 1 — cloud apps**
| App | Host | From → To | Notes |
|-----|------|-----------|-------|
| Immich | cloud CT120 | 2.5.6 → 2.7.5 | XSS + ACL-bypass fixes; stack includes Valkey 9 and vectorchord PG |
| PeerTube | media CT110 | 8.0.2 → 8.2.1 | — |
| Nextcloud AIO | cloud CT121 | NC 32.0.4 → **33.0.5** (major); mastercontainer updated | Major NC version; AIO orchestrated update path |
**Wave 2 — media stack (media VM105)**
| App | From → To | Notes |
|-----|-----------|-------|
| Jellyfin | 10.11.6 → 10.11.11 | — |
| Sonarr | → 4.0.17 | — |
| Radarr | → 6.2.1 | — |
| Prowlarr | → 2.4.0 | — |
| Navidrome | → 0.62.0 | — |
| SABnzbd | 4.5.5 → **5.0.4** (major) | Major upgrade; clean migration |
| Lidarr | DEFERRED | Image maintainer hasn't shipped v3 — needs image swap to linuxserver to go v3; left on current |
| Jellyseerr | INTENTIONALLY HELD on `preview-OIDC` tag | Stable lacks OIDC support; switching would break SSO login |
**Authentik (edge2 CT105) 2025.12.4 → 2026.5.3**
Done as 3 sequential hops: 12.4 → 12.6 → 2026.2.4 → 2026.5.3. `pg_dump` before each hop; migrations clean each time. **SSO verified end-to-end by Matt logging into navi.** Closes the May-2026 CVE waves. (ak app version 5.2.15.)
**Forgejo (edge2 CT103) 14.0.5 → 15.0.3**
EOL-branch migration; v15 schema migrations applied cleanly; web 200, API reports 15.0.3. Pre-15 DB dump at `/root/forgejo-db-pre15-20260621.sql`.
**Headscale 0.28 → 0.29.1 (both instances)**
- **CT106** (utility, IdahoMesh 3-node mesh) — upgraded first as rehearsal; native systemd binary.
- **CT107** (edge2, fleet coordinator, Docker) — 32/32 nodes reconnected post-upgrade; health 200; boot-survival fix preserved (ports remain on 10.10.10.25). 0.29 breaking changes documented in [[meshtastic-headscale-runbook]] (key changes: `randomize_client_port` removed — was a hard blocker; ephemeral key config nested; minimum Tailscale client 1.80.0; bare ACL `*` now tailnet-only).
**OpenTAKServer (utility CT109) 1.7.10 → 1.7.12**
Clean; all 9 TAK services healthy.
### Key decision — RabbitMQ LEFT on 3.12.1 (EOL)
Empirically confirmed during the OTS update: updating OTS to 1.7.12 does **not** touch RabbitMQ — OTS 1.7.12 runs fine against RabbitMQ 3.12.1/Erlang OTP 25. Forcing RabbitMQ → 4.x is the wrong move: OTS is built and tested against 3.12, and 4.x has breaking changes (queue-mirroring removal, feature-flag requirements) that would likely break OTS, which does not appear to support 4.x. AMQP/MQTT ports are localhost-bound, so exposure is low. The EOL 3.12 is a low-priority latent risk to revisit only if/when OTS officially supports RabbitMQ 4.x. **NOT a current action item.**
### Incidental fixes and side-work during Phase 2
- **CT107 boot-survival fix** (applied earlier in the effort, during Phase 1 resolution) — rebound headscale/headplane ports to 10.10.10.25, dropped `tailscale-online.target` dependency, disabled `only_start_if_oidc_is_available` gate, repointed edge2 Caddy; proven by reboot self-heal in ~45 s. Also corrected CT107's own tailscale node ControlURL to `vpn.echo6.co` so it self-registers cleanly.
- **Utility node incident (resolved):** a batch delete of 9 LVM-thin snapshots triggered an SSD TRIM/discard storm that spiked I/O and load transiently; compounded by CT103 argus running hot (transcription + docker-compose build churn). Matt migrated argus to the cloud node, resolving the issue; utility load returned to normal. **LESSON: delete thin-pool snapshots one at a time — not in a batch — to avoid the discard storm.**
- **Nextcloud:** granted `matt@echo6.co` the NC admin role. (user_oidc has no group-claim sync, so this is durable across SSO logins.)
- **Radarr:** set up a `\\192.168.1.160\manual` SMB drop folder on the same NFS export as the library (atomic-move imports) for manual movie filing.
- **Snapshot hygiene:** all rollback snapshots cleaned up after validation — Phase 1 `presec-*`, Phase 2 `prewave2-*`, OTS `pre-ots-*` snapshots all removed.
### Completed — low-urgency batch and cortex AI stack
**Low-urgency batch — DONE:**
- Caddy (utility CT101): 2.10.2 → 2.11.4
- CouchDB (edge2 CT104 livesync): 3.4 → 3.5.2 — all 5 DBs intact
- valkey-8 sidecar (utility CT102 searxng): bumped to latest 8.x
- NATS (utility CT104 central): 2.14.0 → 2.14.2 — all 12 JetStream streams intact
- MediaMTX (CT109): 1.13.0 → 1.19.1
- Mumble (CT109): already latest in Ubuntu repo (1.5.517 — no upstream action possible without going off-distro)
**cortex AI stack — DONE** (protected host; containers only; NO reboot, NO driver touch):
- qdrant: 1.16.3 → 1.18.2
- tei: 1.7 → 1.9 (bge-m3 on GPU)
- ollama: 0.16.1 → 0.30.10 (vault-tagger 100% GPU)
- open-webui: 0.8.1 → 0.9.6
- Both `.ref` vault-engine deps (ollama vault-tagger + tei bge-m3) confirmed working on GPU.
### Minor follow-ups noted (non-urgent)
- **Caddy expired cert:** Caddy flagged an EXPIRED unmanaged cert for `navidrome.echo6.co` (pre-existing condition, not caused by the upgrade) — fix if that hostname matters.
- **MediaMTX deprecated config params:** 1.19 uses deprecated param names (`protocols``rtspTransports`, `encryption``rtspEncryption`) — works now (warnings only); rename before a future MediaMTX release removes them.
- **Rollback artifact cleanup:** retained rollback artifacts to clean once comfortable: `/root` DB dumps on their respective CTs (authentik hop1/2/3, forgejo, headscale107, immich, peertube, OTS), binary backups (caddy.bak, nats-server.bak, mediamtx.bak on their CTs), and openwebui DB backup on cortex (`/home/zvx/openwebui-webui.db.bak-20260621`).
- **Vestigial utility exit-node route:** stale 0.0.0.0/0 exit-node route on utility — optional cleanup (harmless post-0.29 Headscale).
### Separate deferred projects
- Nominatim v5 re-import — see [[nominatim-v5-reimport]]
---
## Phase 3 — Execution Log (2026-06-21/22)
**Status: COMPLETE (2026-06-22)** — all five PVE nodes on 9.2.3/kernel 7.0.12, pi-nas on OMV 8.4/kernel 6.18, cortex NVIDIA driver + DKMS + container-toolkit updated, GPU passthrough verified. Phase 3 has no open items.
### Pre-flight findings
- Cluster `echo6-cluster` 5/5 quorate, quorum 3, NO HA configured (clean guest stop/start, no fencing).
- BLOCKER found + fixed: data, cloud, media had NO Proxmox APT repo configured at all — added `pve-no-subscription` (trixie, matching utility/toc); all 5 then saw the 9.2 stack.
- Note: PVE 9.2 ships **kernel 7.0 as the new default** (proxmox-default-kernel) — nodes boot 7.0.12-1-pve, not 6.17.13 as the audit predicted.
### Cluster nodes — ALL upgraded to PVE 9.2.3 / kernel 7.0.12-1-pve (QEMU 11, LXC 7)
One at a time; cluster stayed 5/5 quorate throughout.
- **media** ✅ — done first (canary). Incidental: CT110 peertube failed to auto-start (recurring `/etc/resolv.conf` immutable-flag vs LXC pre-start-hook conflict) → **permanently fixed**: the flag was an obsolete workaround (NordVPN `set dns` no longer overwrites resolv.conf), cleared it + enabled NordVPN auto-connect, so future reboots won't trip it.
- **data** ✅ — recon-vm (VM1130) healthy; virtiofsd-{kiwix,library,nav} auto-recovered this time; nginx needed the one expected restart (pre-existing mesh-DNS startup race).
- **utility** ✅ — all 11 CTs; central JetStream intact (12 streams); caddy proxying, mesh headscale (CT106, 3 nodes), mesh-bridge (CT107) both tailnets, OTS stack all healthy. CT102 searxng needed a manual `pct start` (transient auto-start miss, no persistent fault).
- **cloud** ✅ — done last (per operator). immich (4 containers + API 200), nextcloud (12 AIO containers, occ healthy, v33.0.5), argus running. (Note: argus CT103 on cloud has argus-capture missing / argus-transcribe masked — operator's in-progress argus→cloud migration, not from the upgrade.)
### pi-nas ✅ FULLY DONE (OMV 8.4 + kernel 6.18)
**OMV (2026-06-21):** OMV 8.1.0-2 → 8.4.0-3 + Debian userspace (Docker, OpenSSL, salt, tailscale). All 5 NFS exports (arr/immich/nextcloud/peertube/data) serving; immich+nextcloud mounts confirmed OK.
**Kernel jump (2026-06-22):** 6.12.62 → 6.18.34+rpt-rpi-2712 — operator physically present. Pre-jump: full `/boot` backup at `/root/boot-backup-pre6.18-20260622.tar.gz` (154 MB) + old kernel left in place as fallback. Booted clean in ~15 s; all 5 NFS exports serving; immich+nextcloud mounts recovered; SD card healthy. Only package remaining HELD: `rpi-eeprom` (SPI bootloader — intentionally untouched; optional to update later).
**Important architecture note recorded:** pi-nas (2.8 TB RAID1) is the NFS storage backend for immich's photo library (644 GB), nextcloud files, peertube, and the *arr library — so a pi-nas reboot stalls those services' storage.
### toc + cortex ✅ DONE (2026-06-22, run from matt-desktop)
- **cortex:** NVIDIA driver 580.159.03 → 580.167.08 + DKMS (built for running kernel) + nvidia-container-toolkit 1.19.1; all AI containers healthy on GPU — vault-tagger generates (Ollama), TEI bge-m3 embeds (200 OK from localhost:8090), Qdrant healthy.
- **toc:** PVE 9.1 → 9.2.3 / kernel 7.0.12-1-pve. **GPU passthrough (vfio) survived the 7.0 kernel** — the key risk, verified.
- Cluster 5/5 quorate with toc rejoined; cortex VM150 auto-started; no leftover snapshot.
**Phase 3 has no open items.**
### Wrap-up (2026-06-22)
**Rollback-artifact sweep done.** Approximately 10.4 GB of campaign DB dumps and binary/config backups removed fleet-wide. Largest single item: a 7.79 GB central Postgres dump on utility CT104. **Retained intentionally:** `/root/boot-backup-pre6.18-20260622.tar.gz` on pi-nas (kernel rollback artifact; 154 MB) and the durable mailcow backup on pi-nas (`…/contabo-prewipe-2026-06/mailcow/`).
**Accepted final state — operator-acknowledged decisions (2026-06-22).** These are closed decisions, not TODOs.
- **Guest OS = security-only.** LXC containers and VMs received security-pocket apt updates only (the intentional Phase 1 scope). Non-security package drift (e.g. Docker CE versions, miscellaneous libs) was deliberately not swept with a full `apt full-upgrade`. The PVE hosts, pi-nas, and cortex did receive full upgrades. Operator accepted this state. A full guest `apt full-upgrade` ("Phase 1.5") remains an option if ever wanted.
- **Apps capped by external factors (decisions, not failures):** RabbitMQ 3.12 left by decision — OTS depends on it and 4.x would break it; Lidarr v2 — the lidarr-on-steroids image maintainer has not shipped v3, would require an image swap; Jellyseerr on `preview-OIDC` — kept because stable 3.3.0 lacks OIDC/SSO support; Mumble 1.5.517 — newest version in the Ubuntu 24.04 repo, upstream 1.5.901 is not available without going off-distro.
- **Deferred project:** Nominatim v5 re-import — spun off to [[nominatim-v5-reimport]].
- **Optional cosmetic follow-ups (non-urgent, non-blocking):** navidrome.echo6.co expired unmanaged cert; MediaMTX deprecated config param names (`protocols`/`encryption`); `rpi-eeprom` held on pi-nas; vestigial utility exit-node route (0.0.0.0/0).
- **Kernel summary:** all 5 PVE hosts on kernel 7.0.12-1-pve; LXC containers share the host kernel (7.0); VM guests (recon-vm, arr VM105) have their own Ubuntu kernels (security-patched during Phase 1, not necessarily absolute-latest upstream); pi-nas on 6.18.34+rpt-rpi-2712; cortex kernel updated as part of the toc+cortex window.
---
## Execution Runbook (Meticulous)
Step-by-step plan to bring every application and package current. Per-app target versions live in **Application Currency** above; this section is the *how/when/order*. Execution model: scoped Sonnet task per target → Opus verifies output → report to Matt; **approval gate before every load-bearing or reboot step.**
### Standing guardrails (apply to EVERY step)
- **Protected hosts — cortex & toc — never in a bulk pass.** They get a dedicated, manual window. Note the coupling: **toc reboot drops cortex (VM150), which is the Claude Code host** — expect to lose the control session during that window; drive it from elsewhere or accept the outage.
- **One target at a time.** Verify health before moving to the next.
- **Snapshot/backup before each mutating step:** `pct snapshot` / `qm snapshot` (or `vzdump`) the guest first; DB dump before any DB-bearing upgrade; run the app's own backup where it has one (Nextcloud AIO, mailcow).
- **No host reboots outside the explicit Phase 3 windows.** Security upgrades that set `reboot-required` are applied but activation deferred to Phase 3. Host OS-security apt for the four cluster nodes is NOT applied in Phase 1 — it rides the Phase 3 full-upgrade window.
- **Rollback for DB-bearing apps = quiesced logical dump + old binary together.** For authentik, forgejo, nextcloud, peertube, and immich the rollback unit is a `pg_dump` / AIO borg / peertube dump taken with the app stopped and the old binary still in place. "Redeploy previous image tag" is NOT valid after forward-only migrations — Django/TypeORM migrations do not reverse, and a `pct snapshot` of a hot separate-volume DB may restore torn. Record before/after version in the tracking table.
- Every result must survive a reboot (standing infra policy).
### Load-bearing dependencies (drive the ordering)
- **Authentik (edge2 CT105) = SSO.** Upgrading it briefly breaks login to everything behind it. Confirm **break-glass admin** access first; do in a low-traffic window; verify dependent-service logins after.
- **Caddy (utility CT101) = home-services ingress.** A restart blips all home services — fold its update into a deliberate moment, not mid-day.
- **Headscale ×2:** edge2 CT107 = main fleet tailnet (34 nodes, `vpn.echo6.co`, ControlURL `https://vpn.echo6.co` → 184.174.35.153) — **HIGH lockout risk**. utility CT106 = IdahoMesh sub-tailnet (`vpn.idahomesh.com`, 3 nodes) — low risk. Upgrade CT106 first as rehearsal. Edge2 CT107 upgrade MUST be driven from an out-of-band path (edge2 host console / direct SSH to 184.174.35.153, NOT over `vpn.echo6.co`); extend node key-expiry first; dump DB off-tailnet before touching it.
- **Nominatim 4→5 is NOT a routine update** — it's a re-import project (see Phase 2C).
### Phase 0 — Pre-flight (no app changes yet)
**STEP 1 (FIRST — gates everything mutating): Verify break-glass and out-of-band access.** Confirm a non-Tailscale, non-Caddy path to every node: Proxmox/Contabo console for edge2, LAN `192.168.1.x` SSH for home nodes, PVE noVNC/`pct enter` for each CT. Confirm PVE/Proxmox web auth is PAM/local, NOT behind Authentik. Prove it: log into Contabo/PVE console; `pct enter` 105/106/107; confirm LAN SSH; locate break-glass creds in `.ref/credentials`. Verify Authentik akadmin (private browser session) + each protected app has a local-admin fallback (Forgejo `admin user`, Nextcloud `occ`, PDM/Vaultwarden local). Every node + CT shell must be reachable WITHOUT Authentik/headscale/Caddy. Losing the management path mid-run is the single highest-consequence failure mode.
**STEP 2: Measure snapshot headroom (per node).** Run `pvesm status` + `df -h` on data/utility/cloud/media/toc/edge2. Record the storage backend per node (LVM-thin/ZFS/dir-qcow2) and its over-provision failure mode. Hard numeric rule: never start a snapshot below ~20% free; require ≥2× the largest guest's expected write-delta. Cloud node CT120/CT121 free space must be measured now — not assumed. Identify where vzdump lands; if it's `data`, freeing `data` is a fleet-wide rollback prerequisite.
**STEP 3: Free space on `data`** (92% full, ~73 GB free) — sized to the largest planned snapshot (VM1130 recon-vm / nominatim+overture+padus). Prune stale vzdump/snapshots/ISOs; `docker image prune` on recon-vm (do NOT remove the pinned `nominatim:4.5` image). Confirm nothing deleted is the only copy of a backup or live NAS data. Verify: usage <~85% on snapshot-backing storage; test snapshot succeeds then remove. **Blocks data host Phase 3 and canary use.**
**STEP 4: Verify off-host restorable backups** (distinct from per-step snapshots) for every stateful guest: central PG/NATS (CT104), opentakserver PG + RabbitMQ definitions (CT109), forgejo PG + repos (CT103), matrix Synapse PG (CT106), nextcloud AIO borg (CT121), edge2 livesync CouchDB (CT104). At least one copy must live off the node being changed.
**STEP 5: Capture pre-change baseline HEALTH snapshot.** Record per-target up/down, container states (`docker ps`/`pct`/`qm status`), key endpoint 200 checks, and `apt list --upgradable`/security counts. Note known oddities: jellyseerr on preview-OIDC dev tag, mailcow CT108 already STOPPED, pinned-stale nominatim, CT118 archivist rpcbind on `0.0.0.0:111`.
**STEP 6: Decision gates.** Close genuinely-open decisions before Phase 1/2/3 can start: daemon-restart tolerance sign-off (Open Dec #7); verify post-cutoff CVE claims against primary advisories with owner+URL per claim; resolve edge2 host-kernel CVE question fully (pin `uname -r`, check repo kernel availability, decide reboot yes/no — not left conditional); define and announce maintenance windows in America/Boise naming user-facing blips.
**STEP 7: edge2 CT108 mailcow — ✅ DONE.** CT108 destroyed 2026-06-20 (`pct destroy 108 --purge`); config + disk image purged. Backup verified durable on pi-nas (`…/contabo-prewipe-2026-06/mailcow/`, sha256-verified); edge1 confirmed serving mail (MX/A → 5.189.158.149). edge2 now hosts CT100107.
### Phase 1 — Guest OS security packages (no reboot; hosts folded into Phase 3)
Guests only in Phase 1 (low blast radius). **The four cluster hosts' OS-security apt rides their Phase 3 full-upgrade window** — do NOT patch data/utility/cloud/media host OS in Phase 1 (avoids a mixed-state libc/openssl across the entire Phase 2 app campaign; both phases get cleaned up in one reboot anyway). edge2 host is already fully patched — no action. toc excluded (Phase 3, with cortex).
Per guest target: `apt-get update` → snapshot → apply **security-pocket** upgrades → needrestart policy (see below) → verify service health and security-upgradable count hits 0. Kernel/libc security that flags `reboot-required` → applied, reboot activation deferred to Phase 3.
**needrestart policy:**
- **Stateful guests — `NEEDRESTART_MODE=l` (list-only), then deliberate bounce:** CT104 central (PG16/NATS/JetStream), CT109 opentakserver, CT110 peertube, VM1130 recon-vm. Stage the packages; take a logical DB dump with the app quiesced; bounce each service deliberately in controlled order. A JetStream restart mid-write loses messages on at-most-once pipelines; a RabbitMQ/redis bounce mid-job is not a clean "blip."
- **Lockout-critical daemons — `NEEDRESTART_MODE=l` + deliberate controlled bounce:** Authentik CT105, headscale instance(s), Caddy CT101/CT111, sshd on any host you're connected through. Run `needrestart -r l` to see what would restart; bounce deliberately so a libc/openssl bump can't auto-bounce auth/tailnet/SSH daemons out from under the operator.
- **All other guests:** `needrestart -r a` (automatic) is acceptable — seconds of blip, no data risk.
**Ordering:**
1. **Canary first** (utility CT112 cobalt idle, or CT102 searxng) — prove the full procedure on a low-stakes guest before touching never-patched targets.
2. **Worst-first (after canary proven):** utility CT119 mesh-territory (91 sec, never patched), CT108 meshai (79 sec), cloud CT120/CT121 guest-OS (101/75 sec), media CT110 peertube (37 sec), utility CT109 opentakserver (31 sec), CT104 central (38 sec, incl. PG 16.13→16.14 + NATS).
3. **Remaining guests:** all other utility/cloud/media/edge2 CTs + data VM1130 (recon-vm). Gate data VM1130 step on data disk remediation verified DONE.
4. **edge2 CT106 matrix / CT107 headscale (NO Tailscale client):** drive via `pct exec` from the edge2 HOST (`root@184.174.35.153`) ONLY — never over service ingress. Treat as own singleton steps, not part of edge2 batch. `needrestart -r l`; do NOT restart headscaled/synapse unless required.
5. **Control-plane singletons LAST:** home-ingress Caddy CT101 and headscale instance(s) — each with DB backup + node-stays-joined verification, never inside a batch.
6. **cortex VM150 (PROTECTED):** security apt manually in the toc+cortex window (Phase 3 coordination); not in bulk.
7. **Cross-cutting:** Tailscale 1.94→1.98 and Docker CE→29.6 ride along Phase 1 under the same snapshots; EXCLUDE edge2 CT106/CT107 from Tailscale bump (no TS client).
8. **pi-nas:** Debian security pocket with `linux-image-*` and `openmediavault*` packages held; kernel upgrade rides Phase 3.
### Phase 2 — Application updates
**Hard isolation rule:** {Headscale, Authentik, Caddy, Forgejo} are NEVER in the same maintenance window — at least one of {tailnet, SSO, ingress, git} must always be a known-good recovery path.
**2A — Load-bearing SSO and control-plane apps (in this order):**
- **Authentik (edge2 CT105) — FIRST, sequential hops, pg_dump per hop.**
- **Pre-flight before any hop:** (1) Prove break-glass: log in as akadmin in a private browser session; confirm every protected app has a working local-admin fallback (Forgejo `admin user`, Nextcloud `occ`, PDM/Vaultwarden local, PVE PAM). (2) Audit/rename duplicate group names — the Django migration fails loudly on dupes. (3) If local `/media` storage is used, stop the service, `mv ./media ./data/media`, rewrite compose volume, before starting any new version.
- **Hop sequence (no skipping):** 2025.12.4 → **2025.12.6** (backported security) → **2026.2.x****2026.5.3**. One hop at a time: `pct snapshot` + `pg_dump` (labelled per hop) → pull new tag → `compose up -d` → migrations run → verified-login gate (all dependent SSO logins tested) → abort+restore on FIRST migration error. Verify all dependent SSO logins BEFORE upgrading any of the dependents.
- Rollback: `compose down` → restore prior tag + dump; or `pct rollback` for the full CT.
- **Caddy (utility CT101 + media CT111) — early in Phase 2, before home-service app verification.** Pair CT101/CT111 in one deliberate low-traffic window (separate from Authentik, Forgejo, and headscale windows). Validate `caddy validate`; spot-check Caddy↔Authentik path; document a direct Caddy-bypassing Authentik admin URL as break-glass. Verify all home services reachable through the new ingress before proceeding with other app upgrades.
**2B — Self-managed updaters** (snapshot → run native updater → verify):
- **OpenTAKServer (utility CT109) — app upgrade first, RabbitMQ SEPARATE:**
- Run OTS upgrade script (upgrades OTS + webUI + schema only). Snapshot first. Verify OTS web/API up; CoT works; TAK clients reconnect.
- **RabbitMQ 3.12.1→4.2.8 is a SEPARATE sub-task (decoupled, after OTS app step).** The OTS updater does NOT migrate RabbitMQ, Erlang, or feature flags — the EOL RabbitMQ 3.12 would silently remain. Decoupled upgrade: `pct snapshot` + `rabbitmqctl export_definitions` → upgrade Erlang to ≥26 → 3.12→3.13.x → `rabbitmqctl enable_feature_flag all` (confirm all enabled, no classic-queue-mirroring config remains) → only then 4.x. A naive 3.12→4.x jump refuses to boot. Gate "OTS done" on RabbitMQ 4.x up + TAK clients reconnecting before accepting the bundled MediaMTX/Mumble bumps.
- Rollback: restore snapshot + definitions export.
- **Immich (cloud CT120)** — prune old images + `docker image prune` before snapshot; disable Immich background regeneration until free space is confirmed; `docker compose pull && up -d`. Do NOT interrupt first-boot 2.5→2.7 DB migration. Also clears Valkey 9.0.2→9.1.0. Rollback = restore `pct snapshot` + `pg_dump` taken with the app stopped (not a hot-DB snapshot alone). Verify web, mobile sync, ML.
- **Nextcloud AIO (cloud CT121)** — AIO borg backup → stop → update mastercontainer → trigger child update via AIO UI (:8080). A stale AIO may need two cycles. Rollback unit = AIO borg backup (mastercontainer downgrade is not supported — borg is the only path back). Verify `occ status` 32.0.11; 12 containers green; integrity clean.
**2C — Versioned apps** (snapshot → bump/upgrade → verify):
- **Forgejo (edge2 CT103) — branch migration 14.x→15.0.3 (separate window from Authentik).**
- Pre-steps: `git clone --mirror` echo6-docs (and other critical repos) to an OFF-edge2 location; pause the `echo6-docs-autocommit` cron so recovery docs survive an edge2/Forgejo failure; run `forgejo doctor check --all [--fix]` to repair stopwatch/tracked_time inconsistencies BEFORE backup+bump. Read 15.0 breaking changes + forward-only FK migration notes.
- `pct snapshot 103` + repos volume + `pg_dump` + app.ini → bump tag 14→15 → pull/up → migrations → verify.
- Rollback: restore 14.0.5 image + volume + dump.
- **Headscale 0.28→0.29.1 — utility CT106 FIRST (rehearsal), edge2 CT107 LAST.**
- CT106 (IdahoMesh, 3 nodes, low risk): standard 0.28→0.29 upgrade guide; `pct snapshot` + DB backup; stop → upgrade binary → config migrate → start. Verify `nodes list` all joined; verify headplane compatibility.
- CT107 (main fleet, 34 nodes, HIGH lockout risk — edge2 front door): extend/disable node key-expiry on all 34 nodes BEFORE starting; `pct snapshot 107` + dump headscale DB off-tailnet (NOT via vpn.echo6.co); keep a second authenticated SSH session open to 184.174.35.153. Drive entirely from edge2 host console / direct SSH to `root@184.174.35.153` — NOT over the tailnet being upgraded. Verify boot chain: edge2 host → CT107 autostart → headscale up → `vpn.echo6.co` resolves → Caddy proxies. Rollback: restore snapshot via console (not via tailnet).
- NOT in the same window as Caddy, Authentik, or Forgejo.
- **Media stack (media VM105):**
- `qm snapshot 105` before batch. SABnzbd 4.5→5.0: pause/empty queue; audit custom post-proc scripts (`empty_postproc` removed in 5.x, scripts now run on failed jobs); pin 5.0.4 tag; pull/up; verify queue intact. Lidarr 2.x→3.x: pin v3; pull/up; DB migration on start; verify library+indexers+download client. Sonarr/Radarr/Prowlarr (1-2 ver): routine bumps after SABnzbd/Lidarr. Jellyfin 10.11.6→10.11.11: pin tag; pull/up. Navidrome: routine bump.
- **Jellyseerr preview-OIDC→stable 3.3.0: DECISION GATE before retag.** preview-OIDC → stable 3.3.0 is a schema DOWNGRADE; TypeORM stable may refuse to start against a dev-migrated DB. Confirm the correct destination artifact (jellyseerr stable vs "Seerr" successor) and, if OIDC is in active use, confirm 3.3.0 parity. Test stable image against a COPY of the DB first. If parity unclear → STOP and report.
- Verify each UI; snapshot before, per-app snapshots for majors.
- **PeerTube (media CT110)** — quiesce import/transcode jobs; `pct snapshot 110` + `pg_dump peertube_prod`; run PeerTube upgrade script 8.0.2→8.2.1; restart; verify. Rollback = `pct rollback` + restore dump.
- **cortex AI stack (PROTECTED — manual, own window):** ollama 0.16.1→0.30.10 (snapshot models volume — the only rollback after pulling new models; upgrade; verify existing `ollama run` without re-pull; verify vault-tagger at localhost:11434 and TEI at localhost:8090 still work); open-webui 0.8.1→0.9.6; qdrant 1.16→1.18; tei 1.7→1.9. Pull images, verify.
**2D — Lower-urgency batch:** caddy (minor), couchdb livesync (3.4→3.5.2), valhalla, photon, kiwix, NATS 2.14.0→2.14.2, valkey-8 sidecar (8.1.5→8.1.8) — snapshot + update as convenient.
Vaultwarden, PDM, WordPress, Synapse/Element/MAS, obsidian, recon-vm Postgres → **already current, skip.**
**2E — Separate project (not in this campaign):**
- **Nominatim 4.5→5.3 + Photon 1.1→1.2 (recon-vm)** — requires full OSM re-import into a SEPARATE DB/instance; keep `nominatim:4.5` image + data until 5.x validated; transient import disk (flat-nodes tens of GB) cannot live on the 92%-full `data` pool; Photon 1.2 AFTER nominatim 5.x validated, never concurrently. Schedule as its own effort.
### Phase 3 — Platform & reboot windows (coordinated, approval-gated)
> ~~🔴 BLOCKED — edge2 reboot was NOT permitted until the CT107 headscale boot-survival fix was in place.~~ **✅ CLEARED 2026-06-20 — the CT107 headscale boot-survival fix is applied and proven** (see Critical Finding 1 in the Phase 1 Execution Log). The triple boot deadlock (tailscale-IP port bind + tailscale-online.target wait + OIDC gate) has been resolved; headscale now self-heals in ~45 seconds after a CT107 or edge2 reboot with zero manual intervention. The Phase-3 edge2 reboot blocker is lifted. The remaining gates below (corosync quorum + Finding-2 headscale-ONLINE check per node) still apply.
> **🔴 GATE for every Phase-3 host reboot (Critical Finding 2):** Before AND after rebooting any node, confirm that node is ONLINE in the headscale coordinator (`headscale nodes list`). Stale `/etc/hosts` entries previously masked coordinator disconnects behind coasting WireGuard tunnels — the stale entries have been removed fleet-wide, but verify ONLINE status at each reboot step to catch any regression. This gate is IN ADDITION TO the corosync-quorum gate below.
**Corosync cluster rules (apply to EVERY node reboot in this section):**
- The five PVE nodes (data, utility, cloud, media, toc) are ONE `echo6-cluster`, quorum = 3 of 5. Losing quorum makes `/etc/pve` read-only fleet-wide and freezes all VM/CT start/stop/config on every surviving node.
- Before EACH node reboot: `pvecm status` must show `Quorate: Yes` and `Expected votes: 5`. Also run `ha-manager status` — if any guest is HA-managed, the reboot triggers fencing/auto-migration rather than a clean local stop/start.
- Reboot ONE node at a time. Wait for full 5/5 rejoin before the next node.
- Disable HA migration and NO live migration during the mixed-version window (QEMU 10/11 boundary — running VMs keep QEMU 10 machine type until cold-started; cross-boundary migration fails).
**Per-node procedure:** `apt update && apt full-upgrade` (so proxmox-ve/qemu/lxc metapackages pull) → this also clears the held host OS-security packages → reboot → verify `pveversion` + all guests return with `onboot=1`. Confirm `pvecm status` shows 5/5 before proceeding.
**Pre-window for each node:** vzdump all guests to an EXTERNAL target (not the local pool, never `data`) so a guest that won't start under LXC 7 can be restored elsewhere. Read PVE 9.2 + LXC 7 release notes for unprivileged-CT/cgroupv2/apparmor breaking changes before the first node.
**Canary node order (lowest blast-radius with headroom first):** media (single arr VM + peertube, non-auth/non-DB-critical) → cloud (immich/nextcloud, after media proven) → data (after disk remediation verified; its near-full pool can fail the mandatory pre-reboot snapshot — do NOT use as canary until remediated) → **utility LAST of the four** (carries Caddy ingress + mesh/central stack — deliberately-scheduled full home-ingress + mesh-coordination outage, announce in advance) → **toc+cortex LAST AND ALONE** (see below).
**Post-reboot re-verification gate (per node):** after each node boots into 9.2/LXC7, re-run the Phase-2 health checks for every guest that received a major app upgrade on that node (CT120 immich, CT121 nextcloud, CT110 peertube, VM105 arr, CT109 OTS). The substrate changed under them.
**toc + cortex — DONE LAST AND ALONE, after all four other nodes are quorate:**
- toc is the 5th cluster vote AND the GPU host for cortex (VM150) AND the management/Claude Code host. Rebooting toc removes a vote and kills the box you'd diagnose from. It must NEVER overlap any other node reboot (toc down + one other = 3/5 → one corosync flap loses quorum).
- Name and TEST a concrete alternate control point (a non-cortex host with keys + tooling) BEFORE this window. Stabilize the tailnet well before it — never in the same period as a headscale upgrade.
- cortex steps IN ORDER:
1. `qm snapshot 150 pre-nvidia` + vzdump (from toc, before anything changes)
2. `apt --only-upgrade nvidia-container-toolkit`; `nvidia-ctk runtime configure`; restart docker → verify `--version`=1.19 + `docker run --gpus all nvidia-smi`
3. `apt --only-upgrade nvidia-driver-580 nvidia-dkms-580`; DKMS rebuild → verify `dkms status` shows installed + `nvidia-smi`=580.167 BEFORE rebooting toc. Keep the old driver package installed as a reinstall fallback.
4. cortex Phase 1 security apt (if not already done in its own window)
- Then reboot toc → verify toc rejoins (`pvecm status` 5/5) → verify VM150 autostart + cortex GPU passthrough return → final AI-stack health check (driver 580.167; all 5 containers; ollama/tei/qdrant/open-webui; vault-tagger+TEI engines; TS online; Docker 29.6).
- **The Claude Code host goes down here** — run this step from the alternate control point.
**pi-nas — standalone window, TWO reboots:**
- Confirm physical/serial console access (no remote KVM on an RPi); export `config.xml` off-box; confirm pi-nas is not mid-sync as a Syncthing/backup target before each reboot. Back up `/boot` + `/boot/firmware`; keep the old kernel installed as fallback.
- Reboot 1: OMV 8.1→8.4 via OMV's own update path (NOT raw `apt full-upgrade` — use OMV UI Update Mgmt or `omv-upgrade`) → reboot → verify shares/SMB/NFS/omv-salt healthy.
- Reboot 2 (AFTER reboot 1 verified): kernel 6.12→6.18 — `apt install linux-image-arm64` (unhold); update bootloader; reboot ON APPROVAL → verify `uname -r`=6.18; LAN+tailnet up; shares mount; Docker up. Rollback: boot retained 6.12; restore `/boot`; on-site SD reflash.
- Independent of cluster windows.
**edge2 — CONDITIONAL (CT107 boot-survival blocker CLEARED 2026-06-20):**
- ~~Previously also blocked on the CT107 headscale boot-survival issue~~ — that blocker is resolved; headscale survives a CT107 or edge2 reboot and self-heals in ~45 seconds (see Critical Finding 1, fixed 2026-06-20).
- Reboot only if Phase 0 resolved the DirtyFrag/copy.fail CVE question as "yes, needs a newer kernel" (Open Decision #5 recorded "no action / no reboot" — edge2 host already fully patched). If a reboot is ever required: give it a dedicated window AFTER Authentik CT105 + headscale CT107 are upgraded and verified; announce SSO+tailnet downtime; drive from a non-tailnet path; confirm all CTs auto-start. edge2 is a single SPOF for ingress — no failover. Do NOT fold into cluster windows.
### Phase 4 — Cross-cutting (fold into earlier phases)
- **Tailscale 1.94→1.98** fleet-wide — rides along Phase 1.
- **Docker CE →29.6** wherever installed — rides along Phase 1/2 under the same snapshots.
### Tracking checklist (tick per target)
### Enumerated Target Checklist
| Phase | Target | Action | Snapshot | Update method | Verify | Rollback | Depends-on |
|---|---|---|---|---|---|---|---|
| 0 | Break-glass / OOB access matrix (ALL nodes) | Verify non-Tailscale, non-Caddy path to every node + each lockout-critical CT (headscale, Authentik CT105, Caddy, Forgejo CT103); confirm PVE/Proxmox web auth is PAM/local not Authentik | None (read-only) | Log into Contabo/PVE console; `pct enter` 105/106/107; confirm LAN `192.168.1.x` SSH; locate break-glass creds in `.ref/credentials`; TEST Authentik akadmin + each app local-admin fallback | Every node + CT shell reachable WITHOUT Authentik/headscale/Caddy; akadmin login screenshotted | N/A | FIRST Phase 0 step; gates everything mutating |
| 0 | Per-node free space + snapshot backend | `pvesm status` + `df -h` on data/utility/cloud/media/toc/edge2; record backend (LVM-thin/ZFS/dir-qcow2) + failure mode; size vs largest guests | None | Read-only; set numeric rule: never snapshot below ~20% free, require ≥2× write-delta | Free GB recorded per pool; cloud node measured; vzdump target located | N/A | Precedes all snapshot-bearing steps |
| 0 | `data` host — disk remediation (92% full, ~73 GB) | Free space sized to largest planned snapshot (VM1130) | None (cleanup; confirm deletions are cache/backup not live) | Prune stale vzdump/snapshots/ISO; `docker image prune` on recon-vm (keep `nominatim:4.5`) | Usage <~85% on snapshot-backing storage; test snapshot succeeds then removed | Restore from Forge/Syncthing/backup if a needed file removed | Free-space capture; BLOCKS data host step + canary use |
| 0 | Off-host restorable backups (stateful guests) | Verify last-good vzdump/PBS off the node for central, OTS, forgejo, matrix, nextcloud, edge2 livesync CouchDB | N/A | Read-only verification + on-demand `pg_dump`/app backup copied off-host | ≥1 backup off the changing node, restore-testable | N/A | Precedes all mutating stateful steps |
| 0 | Baseline HEALTH capture (all targets) | Record up/down, container states, endpoint 200s, `apt upgradable`/security counts; note oddities | None | `pct/qm status`, `docker ps`, curl probes, `apt list --upgradable` | Baseline recorded; oddities logged (jellyseerr dev tag, CT108 stopped, nominatim pin, CT118 rpcbind) | N/A | Precedes Phase 1 |
| 0 | Decision gates | Close daemon-restart tolerance sign-off (Open Dec #7); verify post-cutoff CVE claims vs primary advisories w/ owner+URL; resolve edge2 kernel question (uname -r vs repo availability); define+announce Boise maintenance windows with user-facing blip list | None | Read-only / sign-off | Each decision recorded before Phase 1/2/3 can start | N/A | Gates Phase 1/2/3 |
| 0 | ~~edge2 CT108 mailcow~~ | ✅ **DONE — destroyed 2026-06-20** (`pct destroy 108 --purge`); backup durable on pi-nas (sha256-verified); edge1 confirmed serving mail | N/A | N/A | CT108 purged; edge1 mail in+out confirmed | N/A | Complete |
| 0 | edge2 host kernel CVE question | Decide if DirtyFrag/copy.fail require a reboot | None (record `pveversion -v`) | Compare installed proxmox-kernel vs verified-advisory fixed versions; escalate if no-sub repo lacks fix (no repo changes w/o approval) | Decision (reboot yes/no) recorded | N/A | Gates edge2 Phase 3 row |
| 1 | Procedure canary — utility CT112 cobalt (idle) or CT102 searxng | Prove apt→snapshot→security upgrade→needrestart→verify on low-stakes guest | `pct snapshot` pre-phase1 | `apt-get update`; security pocket only; `needrestart -r l` then deliberate | Security-upgradable=0; service healthy; needrestart clear | `pct rollback` | Phase 0; runs BEFORE worst-first guests |
| 1 | utility CT119 mesh-territory (179/91, never patched) | Security-pocket apt | `pct snapshot 119` | `pct exec`; security pocket; `needrestart -r a` | sec count=0; meshwars up | rollback snapshot | After canary |
| 1 | utility CT108 meshai (109/79) | Security-pocket apt | `pct snapshot 108` | `pct exec`; security pocket; needrestart | sec=0; meshai up | rollback snapshot | After canary |
| 1 | utility CT104 central (PG16 16.13→16.14 + NATS/JetStream) | Security apt incl. PG minor; STAGED, controlled bounce | `pct snapshot 104` + `pg_dumpall` | `NEEDRESTART_MODE=l`; stage pkgs; quiesce JetStream publishers; bounce PG then NATS deliberately | sec=0; `select version()`=16.14; JetStream/MQTT consumers reconnected | rollback snapshot; restore dump | Phase 0 #8; stateful carve-out |
| 1 | utility CT109 opentakserver (OS layer, 31 sec) | Security apt; STAGED, do NOT touch RabbitMQ/MediaMTX/Mumble here | `pct snapshot 109` | `NEEDRESTART_MODE=l`; stage; deliberate bounce | sec=0; TAK clients reconnect | rollback snapshot | Precedes CT109 Phase 2 app work |
| 1 | media CT110 peertube (81/37, native PG16+redis) | Security apt; STAGED controlled bounce | `pct snapshot 110` (PG rides in CT snap) | `NEEDRESTART_MODE=l`; stop import/transcode jobs; snapshot redis; deliberate bounce | sec=0; PeerTube plays; PG16/redis/nginx up; nordvpn up | rollback snapshot | Phase 0 #8 |
| 1 | data VM1130 recon-vm (34/10, PG16+navi-backend) | Security apt; STAGED controlled bounce | `qm snapshot 1130 pre-phase1-apt` | `NEEDRESTART_MODE=l`; deliberate PG + navi-backend bounce | sec=0; psql overture/padus; navi-backend/recon.py/photon/kiwix up | `qm rollback` | data disk remediation done first |
| 1 | cloud CT120 immich (191/101) | Guest security apt; Tailscale 1.94→1.98 + Docker CE→29.6 ride along | `pct snapshot 120 pre-phase1-os` | `pct exec`; security pocket; docker bounce acceptable | sec=0; immich stack Up; web reachable; TS=1.98 | `pct rollback` | Before cloud host step + CT120 app upgrade |
| 1 | cloud CT121 nextcloud AIO (98/75) | Guest security apt; TS+Docker ride along | `pct snapshot 121 pre-phase1-os` | `pct exec`; security pocket | sec=0; 12 AIO containers Up; web reachable | `pct rollback` | Before cloud host step + AIO app upgrade |
| 1 | media VM105 arr (43/9) | Guest security apt; TS+Docker ride along | `qm snapshot 105 pre-os-sec` | `apt`; security pocket; needrestart | sec=0; 8 containers Up; Samba + all UIs load | `qm rollback` | Before VM105 Phase 2 apps |
| 1 | media CT111 mcc (51/13, local caddy+postfix) | Guest security apt | `pct snapshot 111 pre-os-sec` | `pct exec`; security pocket | sec=0; caddy+postfix active | `pct rollback` | Independent (no Phase 2 app) |
| 1 | utility CT100/101/102/103/106/107/112/118 (OS layer) | Security-pocket apt per CT; lockout-critical (Caddy CT101, headscale CT106) carved to LIST mode + LAST; CT118 leave rpcbind alone | `pct snapshot` each + headscale DB backup for CT106 | `pct exec`; security pocket; `needrestart -r l` for control-plane CTs | sec=0 each; ingress spot-check; `headscale nodes list` joined | `pct rollback` | Phase 0; control-plane CTs last as singletons |
| 1 | edge2 CT100-105 (OS layer, via `pct exec` from edge2 host) | Security-pocket apt per CT; DB dumps for CT101 WP MariaDB, CT104 CouchDB, CT105/CT106 PG | `pct snapshot` each pre-phase1 + DB dumps | From `root@184.174.35.153` `pct exec`; security pocket; defer reboot | sec=0; each app reachable via Caddy front door | `pct rollback`; restore DB dump | Phase 0; edge2 host fully patched (host = no-op) |
| 1 | edge2 CT106 matrix / CT107 headscale (NO Tailscale client) | Security apt driven via `pct exec` from edge2 HOST only; singletons | `pct snapshot` + Synapse PG dump (106) + headscale DB backup (107) | From edge2 host shell; `needrestart -r l`; do NOT restart headscaled/synapse unless required | sec=0; matrix federation+login; `headscale nodes list` all joined; Caddy ingress works | `pct rollback`; restore DB | OOB path confirmed first; NOT in edge2 batch |
| 1 | cortex VM150 (PROTECTED, 90 upgradable) | Security apt; manual in toc+cortex window; TS+Docker ride along | `qm snapshot 150 pre-sec-apt` (from toc) | `apt --only-upgrade` security pkgs; `needrestart -r a`; defer reboot | sec=0; 5 AI containers Up; `nvidia-smi` ok | `qm rollback` | Phase 0; protected — not bulk |
| 1 | pi-nas host OS (subset of 130) | Debian security pocket; HOLD kernel 6.12→6.18 + OMV pkgs | Off-box `config.xml` + `dpkg --get-selections` baseline | `apt-mark hold linux-image-* openmediavault*`; `-t trixie-security upgrade`; unhold; needrestart | sec=0; shares mount; OMV UI loads; no reboot now | Reinstall prior pkg from cache; restore config.xml | Phase 0; precedes pi-nas kernel reboot |
| 1 | Cross-cutting Tailscale 1.94→1.98 (all TS-client nodes) | Bump client; rides along Phase 1; EXCLUDE edge2 CT106/CT107 | Covered by per-target Phase 1 snapshot | `apt install --only-upgrade tailscale`; daemon restart | `tailscale version`=1.98; node still joined | Reinstall 1.94; snapshot rollback | Folds into Phase 1; LAN/console fallback confirmed |
| 1 | Cross-cutting Docker CE→29.6 (Docker hosts) | Bump engine; rides along Phase 1/2 | Covered by per-target snapshot | `apt install --only-upgrade docker-ce docker-ce-cli containerd.io`; daemon restart | `docker version`=29.6; containers return | Downgrade pkg; snapshot rollback | Folds into Phase 1; before compose pulls |
| 2 | edge2 CT105 Authentik 2025.12.4→2025.12.6 (hop 1) | SSO hop #1 (FIRST load-bearing app); pre-flight: dedup group names, `/media``/data/media` if local storage | `pct snapshot 105` + `pg_dump` (per-hop) | Edit tag; `compose pull && up -d`; run migrations; NO skipping | UI=12.6; migrations clean; dependent login works | `compose down`; restore tag+dump; or `pct rollback` | Phase 0 break-glass proven; window |
| 2 | edge2 CT105 Authentik 2025.12.6→2026.2.x (hop 2) | SSO hop #2 cross-major | `pct snapshot` + fresh `pg_dump` | tag→2026.2.x; pull/up; migrations | UI=2026.2.x; logins re-verified | restore prior tag+dump | after hop 1 |
| 2 | edge2 CT105 Authentik 2026.2.x→2026.5.3 (hop 3, final) | SSO hop #3; verify CVE/GHSA IDs; insert any mandatory intermediate | `pct snapshot` + fresh `pg_dump` | tag→2026.5.3; pull/up; final migrations | UI=2026.5.3; full sweep of dependent logins; close window | restore prior tag+dump | after hop 2; verify dependents BEFORE upgrading them |
| 2 | utility CT101 + media CT111 Caddy (ingress, EARLY) | Caddy bump (before home-app verification); map Caddy↔Authentik path; document direct Authentik admin URL | `pct snapshot` + Caddyfile/certs backup | `caddy validate`; restart in low-traffic window | home services reachable; certs valid | rollback snapshot/Caddyfile | After CT101 Phase 1; separate window from Authentik |
| 2 | edge2 CT103 Forgejo 14.0.5→15.0.3 (branch migration) | Major migration; pre: off-edge2 `git clone --mirror` echo6-docs + pause autocommit cron + `forgejo doctor check --all --fix` | `pct snapshot 103` + repos volume + `pg_dump` + app.ini | `doctor check`→repair→backup→bump tag 14→15; pull/up; migrations | reports 15.0.3; clone/push works; CI/webhooks ok | restore 14.0.5 image + volume + dump | After CT103 Phase 1; NOT same window as Authentik |
| 2 | utility CT109 OTS app | OTS updater (app+webUI+schema) — does NOT migrate RabbitMQ | `pct snapshot 109` + PG dump | Run OTS upgrade script | OTS web/API up; CoT works; clients reconnect | rollback snapshot + dump | After CT109 Phase 1 OS |
| 2 | utility CT109 RabbitMQ 3.12.1→4.2.8 (SEPARATE) | Decoupled EOL major: Erlang≥26→3.13.x→`enable_feature_flag all`→4.x | `pct snapshot` + `rabbitmqctl export_definitions` | staged hops; confirm flags enabled; no classic-mirroring config | 4.2.x; queues intact; TAK clients reconnect | restore snapshot + definitions | GATES MediaMTX/Mumble accept; after OTS app |
| 2 | utility CT106 + edge2 CT107 Headscale 0.28→0.29 | Lower-stakes instance FIRST (rehearsal), front-door LAST; disable key-expiry; OOB-driven | `pct snapshot` + headscale DB dump off-tailnet | stop→backup→replace binary→config migrate→start; headplane to match | version=0.29.1; `nodes list` all joined; tunnels intact | restore snapshot+DB via console | one-at-a-time; non-tailnet path; not same window as Caddy/Authentik/Forgejo |
| 2 | cloud CT120 Immich 2.5.6→2.7.5 (+Valkey 9.0.2→9.1.0, vectorchord PG) | Stack upgrade; prune images first; do NOT interrupt first-boot migration; disable bg regen | `pct snapshot 120 pre-immich-2.7.5` + `pg_dump` | `compose pull && up -d`; let migrations run; no skip across breaking migration | web+login; jobs process; valkey 9.1.0 ping; version 2.7.5 | restore tag (only with dump) or `pct rollback` | After CT120 Phase 1; free space confirmed |
| 2 | cloud CT121 Nextcloud AIO (master ~12.5→13.2.1, NC 32.0.4→32.0.11) | AIO orchestrated update; borg backup as rollback artifact | AIO borg backup + `pct snapshot 121` | AIO backup→stop→update mastercontainer→UI :8080 "Update containers" | 12 containers green; `occ status` 32.0.11; integrity clean | restore from AIO borg; else `pct rollback` | After CT121 Phase 1; AIO self-gates intermediate versions |
| 2 | media CT110 PeerTube 8.0.2→8.2.1 (native) | Native upgrade script; quiesce jobs | `pct snapshot 110` + `pg_dump peertube_prod` | back up DB → run PeerTube `upgrade.sh`/version steps → restart | reports 8.2.1; video plays; PG16/redis/nginx healthy | `pct rollback` + restore dump | After CT110 Phase 1 |
| 2 | media VM105 jellyfin 10.11.6→10.11.11 | Patch bump | `qm snapshot 105` (pre-app batch) | `compose pull jellyfin && up -d` (pin tag) | UI=10.11.11; plays/transcodes | redeploy prior digest; `qm rollback` | After VM105 Phase 1 |
| 2 | media VM105 SABnzbd 4.5.5→5.0.4 (MAJOR) | Major; pause queue; audit post-proc scripts | `qm snapshot 105 pre-sabnzbd` + `sabnzbd.ini` | pin 5.0.4; pull/up; migrate on start | UI=5.0.4; queue intact; test NZB; arr integrations ok | redeploy 4.5.5 + restore ini (queue repair); `qm rollback` | After VM105 Phase 1; before Lidarr |
| 2 | media VM105 Lidarr 2.x→3.x (MAJOR) | Major branch migration | `qm snapshot 105 pre-lidarr` + `lidarr.db` | pin v3; pull/up; DB migration on start | UI=v3; library+indexers+download client ok | restore db (migration one-way); `qm rollback` | After SABnzbd (verify download-client link) |
| 2 | media VM105 Sonarr/Radarr/Prowlarr (1-2 ver) | Routine bumps | pre-app VM snapshot + config DBs | `compose pull && up -d` | UIs load; prowlarr sync; test grab | redeploy prior tags + DBs | After SABnzbd/Lidarr |
| 2 | media VM105 Navidrome 0.60.3→0.62.0 | Routine bump | pre-app snapshot + DB | `compose pull navidrome && up -d` | UI=0.62.0; library scans; track streams | redeploy 0.60.3 + DB | Independent of arr chain |
| 2 | media VM105 jellyseerr preview-OIDC→stable 3.3.0 | DECISION GATE: schema downgrade + correct artifact (jellyseerr vs Seerr) + OIDC parity | `qm snapshot 105 pre-jellyseerr` + config dir | test stable vs DB COPY first; swap tag only if compatible | UI=3.3.0; requests intact; OIDC/login works | restore preview tag + config; `qm rollback` | After jellyfin/arr; STOP+report if OIDC parity unclear |
| 2 | edge2 CT104 livesync CouchDB 3.4→3.5.2 | Lower-urgency image bump | `pct snapshot 104` + CouchDB volume backup | bump tag; `compose pull && up -d` | reports 3.5.2; `_up` healthy; LiveSync replicates test edit | restore 3.4 + volume; `pct rollback` | After CT104 Phase 1; triple-sync verify |
| 2 | utility CT104 NATS 2.14.0→2.14.2 | Lower-urgency patch | `pct snapshot 104` (JetStream dir in snap) | replace binary; restart | 2.14.2; JetStream/MQTT intact | `pct rollback` | After CT104 Phase 1 (PG already done); stateful blip |
| 2 | utility CT100 meshmonitor / CT102 searxng+valkey-8 (8.1.5→8.1.8) | Lower-urgency image refresh | `pct snapshot` + record digests | `compose pull && up -d` | containers Up; function ok | redeploy prior digest; `pct rollback` | After respective Phase 1 |
| 2 | edge2 CT100 PDM / CT101 WP core / CT102 Vaultwarden / CT106 Synapse | Verify-current / no-op (already at/past fix); WP plugin/theme = separate WP-CLI follow-up | None (Phase 1 snaps) | optional digest re-pull only | versions confirmed; logged as current | redeploy prior digest | Closes as no-op |
| 2 | cortex Ollama 0.16.1→0.30.10 (PROTECTED) | Image bump; no downgrade after new pull | `qm snapshot 150 pre-ollama` + models volume | `compose pull ollama && up -d`; test inference | `/api/version`=0.30.10; models intact; GPU used; vault-tagger works | re-pin 0.16.1; restore snapshot if model format broke | After cortex Phase 1 + Docker + toolkit |
| 2 | cortex Open-WebUI 0.8.1→0.9.6 | Image bump (auth surface) | `qm snapshot 150 pre-openwebui` | bump tag; pull/up | UI loads; login works; reaches Ollama | re-pin 0.8.1; snapshot | After Ollama |
| 2 | cortex Qdrant 1.16.3→1.18.2 / TEI 1.7.4→1.9.3 | Lower-urgency image bumps | `qm snapshot 150` per app | bump tag; pull/up | health endpoints ok; collections/embeddings intact; docs engine works | re-pin prior; snapshot | After cortex Docker+toolkit; verify engines |
| 2 | cortex obsidian-remote v1.12.7 | Verify-current / no-op | None | none | container up; UI reachable | N/A | — |
| 2 | utility CT118 archivist rpcbind 0.0.0.0:111 | Remediate exposure (network change — APPROVAL) | `pct snapshot 118` | per approved option: disable rpcbind / bind localhost+TS / firewall | port 111 not on 0.0.0.0; app functions; external scan closed | re-enable / `pct rollback` | After CT118 Phase 1; approval-gated (Open Dec #6) |
| 2 | ~~edge2 CT108 mailcow~~ | **REMOVED — CT108 destroyed 2026-06-20 (not retained)** | N/A | N/A | N/A | N/A | N/A |
| 3 | Per-node corosync/HA pre-flight (data/utility/cloud/media/toc) | Before EACH reboot | None | `pvecm status` (Quorate:Yes, 5 votes); `ha-manager status` | quorate + expected votes=5; HA implications known | abort if not quorate | Gates each Phase 3 node reboot |
| 3 | Canary node (lowest blast + headroom; media, NOT data until disk freed) | Full 9.2/QEMU11/LXC7/kernel reboot to prove the path | vzdump all guests to EXTERNAL target + record `pveversion -v` | `apt update && apt dist-upgrade` → reboot; NO live-migration in mixed window | `pveversion`=9.2/QEMU11/LXC7/kernel 6.17.13; all guests `onboot` return; quorate | boot prior kernel (GRUB); restore guests from vzdump | Phase 1+2 done; canary before high-stakes nodes |
| 3 | media host PVE 9.1.1→9.2 (+QEMU11/LXC7/kernel) | Platform + reboot | vzdump VM105/CT110/CT111 + snapshots | `dist-upgrade` → reboot | guests return; arr/peertube/caddy healthy; quorate | GRUB prior kernel; vzdump restore | After media Phase 1+2; one node at a time |
| 3 | cloud host PVE 9.2 platform + reboot | Platform + reboot | vzdump CT120/CT121 + record versions | `dist-upgrade` → reboot | immich+nextcloud return healthy; quorate | GRUB prior kernel; vzdump restore | After cloud Phase 1+2; after canary proven |
| 3 | utility host PVE 9.2 platform + reboot | Platform + reboot (known full home-ingress + mesh + central outage) | vzdump all 12 guests + versions | `dist-upgrade` → reboot | all 12 return; Caddy ingress + headscale + central back; quorate | GRUB prior kernel; vzdump restore | LAST of early nodes; low-traffic window |
| 3 | data host PVE 9.2 platform + reboot | Platform + reboot | vzdump VM1130 + versions | `dist-upgrade` → reboot | VM1130 returns; NFS/Samba serve; quorate | GRUB prior kernel; vzdump restore | After disk remediation; one at a time |
| 3 | Hypervisor host OS-security (folded in) | Apply held host libc/openssl WITH the platform full-upgrade per node | covered by per-node vzdump | included in `dist-upgrade` (not applied back in Phase 1) | sec=0 host-side post-reboot; daemons up | per-node rollback | Fold-in decision (Open Dec #2) |
| 3 | Post-reboot app re-verification (per node) | Re-run Phase-2 health checks for majored guests on each rebooted node | None | health probes | CT120/CT121/CT110/VM105/CT109 healthy on new substrate | re-snapshot/redeploy degraded component | After each node reboot |
| 3 | cortex nvidia-container-toolkit 1.18→1.19 | Toolkit bump (pair w/ driver) | `qm snapshot 150 pre-nvtoolkit` | `apt --only-upgrade nvidia-container-toolkit`; `nvidia-ctk runtime configure`; restart docker | `--version`=1.19; `docker run --gpus all nvidia-smi` ok | downgrade 1.18; snapshot | toc+cortex window |
| 3 | cortex NVIDIA driver 580.159→580.167 + DKMS (reboot) | Driver + DKMS as discrete reversible step BEFORE toc reboot | `qm snapshot 150 pre-nvidia` + vzdump | `apt --only-upgrade nvidia-driver-580 nvidia-dkms-580`; DKMS rebuild; verify BEFORE toc reboot | `nvidia-smi`=580.167; `dkms status` installed; containers see GPU | restore snapshot; reinstall 580.159 + DKMS | Keep old driver pkg; after toolkit |
| 3 | toc host PVE 9.2 + QEMU11 + LXC7 + kernel (reboot — LAST, drops cortex) | Platform + reboot; only guest is VM150 | vzdump/snapshot VM150 + record `pveversion` | snapshot VM150 → `dist-upgrade` toc → reboot → VM150 returns → cortex driver reboot | `pveversion`=9.2; toc rejoins (5/5 votes); VM150 + GPU healthy | restore VM150 from vzdump; GRUB prior kernel | LAST + alone; alternate control point TESTED; all 4 other nodes quorate |
| 3 | cortex final stack health verification | Post-window full AI-stack + engine check | keep pre-window snaps until verified | verification only | driver 580.167; all 5 containers; ollama/tei/qdrant/open-webui; vault-tagger+TEI engines ok; TS online; Docker 29.6 | restore degraded component; worst case VM150 vzdump | LAST action of node group |
| 3 | ~~pi-nas OMV 8.1→8.4 (reboot #1, separate window)~~ | ✅ **DONE 2026-06-21**`apt full-upgrade` (kernel/firmware 6 pkgs held); OMV 8.4.0-3 installed; omv-salt deploy ran (14/0 OK); initramfs regenerated for 6.12.62 only; rebooted; uname -r still 6.12.62; all 5 NFS exports serving; immich-nfs-OK, nextcloud-nfs-OK; nfs-kernel-server + docker active; `degraded` = pre-existing quotaon "File exists" false alarm, unrelated to upgrade | N/A | N/A | N/A | Complete |
| 3 | pi-nas kernel 6.12→6.18 (reboot #2, separate window) | Kernel bump → reboot; activates deferred security | back up `/boot`+`/boot/firmware`; keep 6.12 installed; fresh config.xml | `apt install linux-image-arm64` (unhold); update bootloader; reboot ON APPROVAL | `uname -r`=6.18; LAN+tailnet up; shares mount; Docker up | boot retained 6.12; restore `/boot`; on-site SD reflash | After OMV reboot verified; physical access |
| 3 | edge2 host PVE 8.4 reboot — CONDITIONAL (CT107 boot-survival blocker CLEARED 2026-06-20) | ~~Previously blocked on CT107 headscale boot-survival~~ — cleared (self-heals in ~45s). Now gated only on Phase 0 kernel-CVE decision (Open Dec #5 = no reboot needed — host fully patched). If ever required: own no-failover window AFTER CT105/CT107 verified | vzdump all guests + `pveversion` | scoped kernel/security upgrade → reboot in approved window | `uname -r`=fixed; all CTs return; Caddy-fronted services reachable | GRUB prior kernel; vzdump restore | GATED on Phase 0; announce SSO+tailnet downtime; non-tailnet path |
| — | Nominatim 4.5→5.3 + Photon 1.1→1.2 (SEPARATE PROJECT) | Full re-import into separate DB/instance; keep 4.5 until validated; Photon AFTER nominatim validated | `qm snapshot 1130` + retain 4.5 image/data | import to new DB; size flat-nodes vs free space (NOT on data); cut over after verify | 5.x geocode correct; Photon 1.2 reads v5 | re-pin nominatim:4.5 + old data | Wholly separate from patch campaign; after disk remediation |
---
## Key Facts & Top Risks
### Cluster & control-plane facts
- **One corosync cluster.** data, utility, cloud, media, toc are `echo6-cluster` — quorum 3 of 5. Every Phase 3 reboot is gated on `pvecm status` (Quorate:Yes, expected votes=5); ONE node at a time; confirm 5/5 rejoin before next.
- **toc coupling.** toc is the 5th cluster vote + GPU host for cortex VM150 + the management/Claude Code host. toc+cortex are the FINAL Phase 3 action, alone — never overlap another node reboot.
- **Headscale topology (resolved).** Main fleet tailnet (34 nodes, ControlURL `vpn.echo6.co`) = edge2 CT107 (Headscale 0.28.0) — HIGH lockout risk; upgrade LAST, drive from out-of-band. Separate IdahoMesh sub-tailnet (`vpn.idahomesh.com`, 3 nodes) = utility CT106 — low risk; upgrade FIRST as rehearsal. No routing through old-Contabo.
### Top risks (critical first)
1. **Headscale edge2 CT107 self-lockout** (CRITICAL) — upgrading the main tailnet over the tailnet itself loses the control path to 34 nodes. Mitigation: out-of-band only (edge2 host console / direct SSH 184.174.35.153), key-expiry extended, DB dumped off-tailnet, boot chain verified.
2. **Corosync quorum loss** (CRITICAL) — two nodes down simultaneously = quorum lost, `/etc/pve` read-only fleet-wide. Mitigation: one node at a time, `pvecm status` gated before each, HA disabled during window.
3. **Authentik SSO migration chain** (CRITICAL) — 4-hop irreversible Django migration sequence; a failed hop with no per-hop pg_dump can strand the entire SSO surface. Mitigation: pg_dump per hop, verified-login gate, break-glass proven before starting.
4. **data node snapshot-space exhaustion** (CRITICAL) — 92% full; a snapshot failure mid-upgrade leaves the guest in a half-upgraded, unrollbackable state. Mitigation: disk remediated to <85% before any snapshot on that node; do not use as Phase 3 canary until resolved.
5. **RabbitMQ EOL/won't-boot** (HIGH) — 3.12→4.x direct jump refuses to start; 3.x has no security backports. Mitigation: staged hops (3.12→3.13.x→feature-flags→4.x), decoupled from OTS updater.
6. **Forward-only DB migrations / false-rollback assumption** (HIGH) — "redeploy previous image tag" is NOT a valid rollback for authentik, forgejo, nextcloud, peertube, immich after a migration runs. Rollback = quiesced logical dump + old binary together.
7. **toc reboot drops cortex + management host** (HIGH) — cortex is the Claude Code host; losing it during diagnosis is double-jeopardy. Mitigation: alternate control point tested before window; cortex NVIDIA/DKMS verified before toc reboot; toc+cortex done last and alone.
8. **pi-nas non-boot after combined kernel+OMV change** (HIGH) — RPi has no remote KVM; a bad combined update requires on-site SD reflash. Mitigation: two separate reboots (OMV first, kernel second); retain 6.12; `/boot`+`/boot/firmware` backed up; physical access confirmed before each.
---
## Full Inventory
Complete point-in-time state of every node, guest, and container service.
### data (PVE 9.1.1)
- Host: 126 upgradable / 51 security; no Docker installed; roles: NAS, NFS, Samba
- **VM1130 recon-vm** (Ubuntu 24.04) — 34 upgradable / 10 security
- PostgreSQL 16.14 (DBs: overture, padus), photon, kiwix, recon.py, 7x navi-backend, nginx, Apache, Samba
- Docker: valhalla:latest, nominatim:4.5 (stale 14 months), zimit:latest (not running)
### utility (PVE 9.1.1)
- Host: 148 upgradable / 38 security; PVE 9.2 platform update pending; 12 LXC guests
| CT | Name | Upgradable / Sec | Services |
|----|------|-----------------|----------|
| CT100 | meshmonitor | 62 / 15 | ghcr.io/yeraze/meshmonitor:latest |
| CT101 | caddy (home ingress) | 59 / 18 | — |
| CT102 | searxng | 54 / 15 | searxng/searxng:latest + valkey/valkey:8-alpine |
| CT103 | argus | 29 / 26 | RF capture / transcribe / viewer |
| CT104 | central | 50 / 38 | PostgreSQL 16 + NATS/MQTT/JetStream |
| CT106 | meshtastic-hs | 54 / 17 | headscale control plane |
| CT107 | mesh-bridge | 53 / 15 | dual tailscaled |
| CT108 | meshai | 109 / 79 | work-meshai local build |
| CT109 | opentakserver | 34 / 31 | nginx / PG16 / rabbitmq / mumble / mediamtx / CoT |
| CT112 | cobalt | 35 / 29 | build/CI (idle) |
| CT118 | archivist | 62 / 17 | archivist + rpcbind (FLAG: port 111 on 0.0.0.0) |
| CT119 | mesh-territory | 179 / 91 | meshwars:latest (never patched) |
### cloud (PVE 9.1.1)
- Host: 113 upgradable / 38 security; 2 LXC guests
| CT | Name | Upgradable / Sec | Services |
|----|------|-----------------|----------|
| CT120 | immich | 191 / 101 | immich_server, immich_machine_learning, valkey/valkey:9, immich postgres (14-vectorchord) — all drifted |
| CT121 | nextcloud AIO | 98 / 75 | 12 containers: mastercontainer + apache + nextcloud + postgresql + redis + collabora + clamav + imaginary + fulltextsearch + notify-push + whiteboard + docker-socket-proxy; NC 32.0.4; mastercontainer behind |
### media (PVE 9.1.1)
- Host: 107 upgradable / 36 security; 3 guests
| Guest | Name | Upgradable / Sec | Services |
|-------|------|-----------------|----------|
| VM105 | arr | 43 / 9 (Ubuntu 24.04) | jellyfin / sonarr / radarr / prowlarr / sabnzbd / lidarr / navidrome / jellyseerr (preview-OIDC) all :latest + Samba |
| CT110 | peertube | 81 / 37 | v8.0.2; nginx / PG16 / redis / peertube / pt-downloader + importer + monitor / nordvpn |
| CT111 | mcc | 51 / 13 | caddy + postfix |
### toc (PVE 9.1.1) — PROTECTED
- Host: 189 upgradable / 43 security; PVE 9.2 platform update pending; reboot required
- Hosts only VM150 cortex — coordinate any toc work with cortex maintenance window
### cortex (VM150, Ubuntu 24.04) — PROTECTED GPU / Claude Code host
- 90 apt upgradable
- NVIDIA driver 580.159 → 580.167 + DKMS (reboot required)
- nvidia-container-toolkit 1.18 → 1.19
- Docker: ollama / tei 1.7 / qdrant / open-webui / obsidian — all drifted
### pi-nas (Debian 13, arm64, RPi + OMV)
- 130 apt upgradable; kernel 6.12 → 6.18 (reboot required); OMV 8.1 → 8.4
- Docker engine installed; 0 containers running
### edge2 (PVE 8.4.19) — Host Fully Patched
- 8 LXC guests (CT108 mailcow decommissioned 2026-06-20)
| CT | Name | Services / Status |
|----|------|-------------------|
| CT100 | pdm | PDM 1.1.4, current (native) |
| CT101 | wordpress | Apache 2.4.67 / PHP 8.4 / MariaDB 11.8.6 / WP core 7.0 — plugin/theme status needs WP-CLI |
| CT102 | vaultwarden | vaultwarden/server:latest — drift indeterminate |
| CT103 | forgejo | 14.0.5 (forgejo:14 + postgres:16-alpine drifted) |
| CT104 | livesync | couchdb:3.4 drifted + local provisioner |
| CT105 | authentik | 2025.12.4 (server/worker/postgres) → upgrade to 2026.2.0 |
| CT106 | matrix | Synapse 1.155.0 / Element / MAS + mautrix-signal + postgres; OS apt security updates pending; no Tailscale client |
| CT107 | headscale | 0.28.0 → 0.29.0 + headplane; no Tailscale client |
| ~~CT108~~ | ~~mailcow~~ | **Decommissioned 2026-06-20** — destroyed (`pct destroy 108 --purge`); superseded by edge1, backup on pi-nas |
---
## Coverage Notes
- **Complete:** all five PVE-9 nodes (data, utility, cloud, media, toc), edge2, cortex, pi-nas, and all discovered guests/containers.
- **Excluded:** edge1 (the rebuilt ex-Contabo, mail-only) — mid-rebuild/maintenance at time of audit; re-audit once it settles.
- **Incomplete:** WordPress (CT101) plugin and theme status — requires WP-CLI; not assessed.
- **Method:** read-only throughout — SSH / `pct exec`, `apt list --upgradable`, `docker manifest inspect` for same-tag drift. Floating/pinned-tag caveats noted inline where drift could not be confirmed.