echo6-docs/vault/projects/fleet-patch-audit.md
echo6-autocommit 586e16de88 auto: docs sync 2026-07-17T18:00:20+00:00
Files changed: vault/.obsidian/workspace.json vault/docs/hardware/environment.md vault/docs/hardware/ip-allocation.md vault/docs/services/services.md vault/glossary.md vault/projects/fleet-patch-audit.md
2026-07-17 18:00:20 +00:00

88 KiB
Raw Blame History

title type tags aliases related updated status
Fleet Patch Audit — 2026-06-19 project
proxmox
fleet-platform-baseline
lxc-service-migration
caddy
services
ip-allocation
2026-07-17 complete

Fleet Patch Audit — 2026-06-19

Read-only audit snapshot as of 2026-06-19. Nothing has been applied — this is a planning document to build the patch plan from.

Topology note: the old Contabo VPS has been rebuilt as edge1 (mail-only); edge2 is now the front door for everything else. edge1 is excluded from this audit (mid-rebuild/maintenance). Headscale: edge2 CT107 is the main fleet tailnet (34 nodes, vpn.echo6.co, self-hosted Headscale 0.28.0); utility CT106 is a separate IdahoMesh sub-tailnet (vpn.idahomesh.com, 3 nodes, low-risk). No services route through old-Contabo. Mailcow CT108: destroyed 2026-06-20 (pct destroy 108 --purge); backup preserved durably on pi-nas (…/contabo-prewipe-2026-06/mailcow/, sha256-verified); live mail on edge1 (MX/A for mail.echo6.co → 5.189.158.149). CT101 update (2026-07-17): migrated from WordPress/MariaDB to Grav CMS 2.0.11 (flat-file, no database — MariaDB purged from the container); serves idahomesh.com (not intermountainmesh.com — that domain has always been parked at a third party and never reached this container); hostname wordpress unchanged. See environment / services for current state. All WordPress/MariaDB references below reflect the pre-migration state as audited on 2026-06-19 and are left as historical record.


Campaign complete (2026-06-22)

Phases 13 are fully done. The entire fleet is on the current platform.

  • Phase 1 (guest/VM security apt) — COMPLETE 2026-06-20. 26 guests patched, ~600+ security packages cleared, zero data loss.
  • Phase 2 (app/container updates) — COMPLETE 2026-06-21. All app upgrades done: authentik 2025.12.4→2026.5.3 (sequential), Forgejo 14→15, Headscale 0.28→0.29.1 (both instances), Immich 2.5.6→2.7.5, Nextcloud AIO→NC 33.0.5, media stack (Jellyfin/SABnzbd/arr), cortex AI stack (Ollama/TEI/Qdrant/Open-WebUI), and the low-urgency batch.
  • Phase 3 (platform/reboot windows) — COMPLETE 2026-06-22. All 5 PVE nodes on 9.2.3/kernel 7.0.12-1-pve (including toc+cortex); pi-nas on OMV 8.4/kernel 6.18; cortex NVIDIA driver 580.167.08 + DKMS + nvidia-container-toolkit 1.19.1; GPU passthrough (vfio) survived the 7.0 kernel; cluster 5/5 quorate.

Intentionally deferred / out of scope (not failures):

  • RabbitMQ — left on 3.12 by decision; OTS does not support 4.x and exposure is localhost-bound. Revisit only if/when OTS officially supports RabbitMQ 4.x.
  • Nominatim v5 — full re-import project, spun off to nominatim-v5-reimport.
  • Optional cosmetic cleanups: navidrome cert renewal (navidrome.echo6.co expired unmanaged cert); MediaMTX deprecated config param rename (protocolsrtspTransports, encryptionrtspEncryption); retained /root rollback artifacts (DB dumps, binary backups on their CTs); vestigial utility exit-node route (0.0.0.0/0, harmless); rpi-eeprom still held on pi-nas (SPI bootloader, intentionally untouched).

Prioritized Backlog

Tier 1 — Security-Urgent Guest OS

These containers have the highest raw security-update counts and have not been patched recently (or never). Address before any platform work.

Host Guest Upgradable / Security Notes
utility CT119 mesh-territory 179 / 91 sec Never patched
utility CT108 meshai 109 / 79
cloud CT120 immich guest-OS 191 / 101
cloud CT121 nextcloud guest-OS 98 / 75
media CT110 peertube 81 / 37
utility CT109 opentakserver 34 / 31
utility CT104 central 50 / 38 Includes PostgreSQL 16.13 → 16.14

Tier 2 — App / Container Updates

Updates where the application or its Docker images have drifted from current upstream, ordered roughly by operational risk.

Scope Guest Item Notes
edge2 CT105 authentik 2025.12.4 → 2026.5.3 #1 security item — 7 CVEs + 5 GHSAs in gap; sequential upgrade (min: 2025.12.6)
edge2 CT107 headscale 0.28.0 → 0.29.1 Also a 2nd headscale on utility CT106
edge2 CT103 forgejo 14.0.5 → 15.0.3 14.x EOL 2026-04-30 — migrate branch, not just patch
edge2 CT106 synapse 1.155.0 / Element / MAS Image drift + pending OS apt security updates
edge2 CT104 livesync couchdb:3.4 Docker image drift
edge2 CT108 mailcow (18 containers) Decommissioned 2026-06-20 — superseded by edge1; no longer an update target
cloud CT120 immich — server/ml/valkey:9/postgres(14-vectorchord) 4 images drifted
cloud CT121 nextcloud AIO — mastercontainer + NC app 32.0.4 Mastercontainer behind; 12-container stack
cortex VM150 ollama / tei(1.7) / qdrant / open-webui / obsidian 5 AI containers drifted
media VM105 arr stack — 8 containers jellyfin/sonarr/radarr/prowlarr/sabnzbd/lidarr/navidrome/jellyseerr all :latest

Tier 3 — Platform / Reboot Windows

Reboot-bearing updates. Coordinate maintenance windows carefully. cortex and toc are PROTECTED hosts.

Scope Item Detail
data, utility, cloud, media, toc (PVE 9 nodes) PVE 9.1.1 → 9.2 Reboot required
data, utility, cloud, media, toc QEMU 10 → 11 Reboot required
data, utility, cloud, media, toc LXC 6 → 7 Reboot required
data, utility, cloud, media, toc Kernel 6.17.2 → 6.17.13 Reboot required
edge2 PVE 8.4.19 Already fully patched — no action needed
cortex VM150 (PROTECTED) NVIDIA driver 580.159 → 580.167 + DKMS Reboot required
cortex VM150 (PROTECTED) nvidia-container-toolkit 1.18 → 1.19
pi-nas Kernel 6.12.62 → 6.18.34+rpt-rpi-2712 DONE 2026-06-22 — operator present; /boot backup at /root/boot-backup-pre6.18-20260622.tar.gz; booted clean; rpi-eeprom held
pi-nas OMV 8.1 → 8.4 DONE 2026-06-21 — 8.4.0-3

Cross-Cutting (All / Most Hosts)

Item Detail
Tailscale 1.94 → 1.98 nearly everywhere
Docker CE → 29.6 wherever Docker is installed

Non-Update Flags

Issues noted that are not package/image updates but warrant attention.

Host / Guest Flag Detail
data Disk 92% full ~73 GB / 938 GB free; address before patching
utility CT118 archivist rpcbind on 0.0.0.0:111 No Tailscale client or firewall on this CT; exposed port
media VM105 jellyseerr Non-stable image Running preview-OIDC tag, not a stable release
data VM1130 nominatim Stale image (14 months) nominatim:4.5, pinned; confirm intentional
edge2 CT106 matrix / CT107 headscale No Tailscale client Ingress via caddy; verify internal routing before patching

Application Currency (2026-06-19)

Running application version vs latest stable upstream, per app — the "is everything up to date?" view (distinct from OS-package/image-drift above). Read-only. Caveat: version deltas were measured directly from the running apps and are reliable; the specific CVE/advisory IDs below were surfaced by the audit agents and are post-knowledge-cutoff — verify against primary advisories before acting on them. GitHub's anonymous API rate-limited a few "latest" lookups (noted inline).

Security-relevant — prioritize

App Where Running → Latest Behind Security note (verify)
Authentik edge2 CT105 2025.12.4 → 2026.5.3 ~6 mo / 5 majors #1 — claimed 7 CVEs + 5 GHSAs across two May-2026 waves. Min-disruption: 2025.12.6 (same branch, backported fixes); full: 2026.5.3. SSO — upgrade sequentially.
Immich cloud CT120 2.5.6 → 2.7.5 9 rel shared-link ACL bypass (2.6.0) + stored XSS panorama viewer (2.7.0)
Valkey cloud CT120 9.0.2 → 9.1.0 3 patch 6 CVEs, "SECURITY" urgency (use-after-free, DoS, RESP injection)
RabbitMQ utility CT109 3.12.1 → 4.2.8 EOL major 3.x abandoned upstream; 4 advisories 2026-06-18, no 3.x backports
Nominatim data/recon-vm 4.5.0 → 5.3.2 ~19 mo, major v4→v5 data-model break; needs image swap + full re-import (pair w/ Photon 1.2.0)
Ollama cortex 0.16.1 → 0.30.10 14 minor history of SSRF / path-traversal CVEs
Nextcloud cloud CT121 32.0.4 → 32.0.11; AIO ~v12.5 → 13.2.1 7 patch / ~8 AIO resource-exhaustion fix (32.0.10); AIO image ~5 mo stale
Forgejo edge2 CT103 14.0.5 → 15.0.3 EOL branch 14.x EOL 2026-04-30; no further backports — migrate to 15.0.3 (LTS)
SABnzbd media VM105 4.5.5 → 5.0.4 major 5.x security hardening over 4.x
Jellyfin media VM105 10.11.6 → 10.11.11 5 patch 10.11.x patch cycle has carried CVE fixes
Jellyseerr media VM105 preview-OIDC dev → 3.3.0 unreleased running unpinned dev code — no stable security posture
PeerTube media CT110 8.0.2 → 8.2.1 2 patch 8.x patch releases include security fixes
Open-WebUI cortex 0.8.1 → 0.9.6 ~1.5 minor touches auth/session surface
MediaMTX utility CT109 1.13.0 → 1.19.1 6 minor RTSP/WebRTC streaming attack surface
Headscale ×2 utility CT106 + edge2 CT107 0.28.0 → 0.29.1 1 minor upgrade-guide-required; two separate instances; no security flag noted

Operational / host-level (not a simple app bump)

Item Where Finding
Mailcow edge2 CT108 DECOMMISSIONED 2026-06-20 — destroyed (pct destroy 108 --purge); backup on pi-nas, live mail on edge1.
Host kernel edge2 host Agent flagged DirtyFrag (CVE-2026-43284/-43500) + copy.fail (CVE-2026-31431, claimed CISA KEV) as host-kernel LPE. Tension: the host audit showed edge2 fully patched (0 upgradable) on its repo — verify whether these need a kernel newer than the no-subscription repo provides.

Current / already past the fix (no action)

Vaultwarden 1.36.0 (edge2 CT102 — has the SSO-takeover/org-access CVE fixes) · PDM 1.1.4 (edge2 CT100 — past the RCE PSA) · WordPress 7.0 core + all plugins/themes (edge2 CT101 — superseded 2026-07-17, see topology note above) · Synapse 1.155.0 / Element / MAS (edge2 CT106 — current, only minor :latest digest drift) · obsidian-remote v1.12.7 (cortex) · PostgreSQL 16.14 (recon-vm).

Lower urgency

Mumble 1.5.517→1.5.901 · Caddy 2.10.2/2.11.3→2.11.4 · Qdrant 1.16.3→1.18.2 · TEI 1.7.4→1.9.3 · Valhalla 3.6.3→3.7.0 · Photon 1.1.0→1.2.0 · kiwix 3.7.0→3.8.2 · CouchDB 3.4.3→3.5.2 (livesync) · Navidrome 0.60.3→0.62.0 · Sonarr/Radarr/Prowlarr/Lidarr 12 versions · NATS 2.14.0→2.14.2 · PostgreSQL 16.12/16.13→16.14 · meshmonitor (~1 mo, exact ver undeterminable) · searxng (rolling, ~4.5 mo) + valkey-8 sidecar 8.1.5→8.1.8 · mautrix_signal v0.2603.0.

Internal echo6 apps (no upstream to track)

central-, meshai, archivist, meshwars, recon / recon-watchdog, navi- — running; version = current git head.


Proposed Patch Approach (NOT executed — for planning)

Lowest-risk changes first; everything reboot-bearing deferred to scheduled windows.

Principles

  • Security-pocket apt only in the first pass (openssl, openssh, gnutls, krb5, libc6, samba, nghttp2, etc.). No PVE 9.2 / QEMU / LXC / kernel, no Docker image pulls, no OMV/NVIDIA, no reboots.
  • Protected hosts (cortex, toc) never go in a bulk pass — handled individually in their own window. Note: toc hosts cortex (VM150), so a toc reboot drops cortex — the two must be coordinated together.
  • edge1 (mail) excluded while it's mid-rebuild.
  • needrestart will bounce affected daemons after glibc/openssl upgrades — seconds of blip per guest, no data risk.

Phased plan

Phase Scope Reboot? Notes
1 — Guest/VM security apt utility CT100,101,102,103,104,106,107,108,109,112,118,119 · cloud CT120,121 · media VM105,CT110,CT111 · data VM1130 · edge2 CT100107 No Lowest blast radius. Worst-first: CT119, CT108, immich/nextcloud guest-OS, peertube.
2 — Hypervisor host OS security data, utility, cloud, media host OSes (not toc) No One node at a time; restarts hypervisor-side daemons (smbd etc.). edge2 host already patched.
3 — App / container updates (Tier 2) per-app, native updater each Per-app COMPLETE 2026-06-21 — all app upgrades done (security-critical, low-urgency batch, and cortex AI stack). See Phase 2 Execution Log.
4 — Reboot windows (Tier 3) PVE 9.2 + QEMU 11 + LXC 7 + kernel on the 5 PVE-9 nodes; pi-nas kernel + OMV; cortex NVIDIA/DKMS Yes COMPLETE (2026-06-22) — all 5 PVE nodes on 9.2.3/kernel 7.0.12; pi-nas on OMV 8.4/kernel 6.18; cortex NVIDIA 580.167/DKMS/toolkit 1.19.1; GPU passthrough verified. See Phase 3 Execution Log.
Cross-cutting Tailscale 1.94→1.98 fleet-wide No Can ride along Phase 1/2.

Special handling — do NOT bulk-patch these; use the native updater

  • mailcow (edge2 CT108) → decommissioned 2026-06-20, no action
  • nextcloud AIO (cloud CT121) → update mastercontainer, then in-UI update button (port 8080)
  • immich (cloud CT120) → docker compose pull && up -d from its compose dir
  • authentik (edge2 CT105) → sequential version upgrades with migrations; 2025.12.4 → 2026.2.0 cannot skip releases
  • pi-nas OMV → OMV's own update path, not raw apt
  • cortex / toc → manual, protected, own window

Open Decisions

Resolved (2026-06-19/20)

  1. RabbitMQ 3.12→4.xACCEPTED as a decoupled sub-task (Erlang ≥26 → 3.13.x → enable feature-flags → 4.x); the OpenTAKServer bump does NOT cover it. In scope for this campaign.
  2. Break-glassVALIDATED: all 7 LAN SSH paths work independent of the tailnet; every PVE node (incl. edge2) has local pam/pve realms; Authentik akadmin, Forgejo matt, Nextcloud admin local logins confirmed. ⚠️ One gap → PDM (edge2 CT100) has only openid:authentik in domains.cfg; confirm root@pam login at https://100.64.0.28:8443 (or add a pam: stanza) before the Authentik upgrade.
  3. Phase scopeALL PHASES; no-reboot work (Phase 12) first, platform/reboot (Phase 3) scheduled separately.
  4. Daemon-restart toleranceDOWNTIME ACCEPTED; needrestart blips OK on the stateful guests (still take logical DB dumps for safety, but no holding restarts for availability).
  5. edge2 host kernel (DirtyFrag/copy.fail)NO ACTION / NO REBOOT: running 6.8.12-30-pve, the newest its repos offer; no update available, no reboot-required flag. Not actionable without a repo/branch change. edge2 needs no Phase 3 reboot.
  6. CVE verificationNOT GATING: being behind on versions is sufficient justification; post-cutoff CVE IDs are not chased or relied on for ordering.
  7. Maintenance windowsNO CONSTRAINT (single user); disruptive steps may run anytime, no scheduling/announcement needed.
  8. Nominatim 4→5DEFERRED / OUT OF SCOPE: spun off to nominatim-v5-reimport as its own project.
  9. Headscale locationRESOLVED: fleet control plane = edge2 CT107 (vpn.echo6.co, 34 nodes — high-risk); IdahoMesh sub-tailnet = utility CT106 (vpn.idahomesh.com, 3 nodes — low-risk). No Contabo routing.
  10. mailcow CT108RESOLVED: superseded by edge1; safe to decommission (pending backup-off-CT verify below).
  11. Phase 2 host-OS sequencingRESOLVED: folded into Phase 3 reboot window.
  12. In-scope app campaign — Forgejo 14→15, Matrix/Synapse, Authentik chain, RabbitMQ, media stack, immich, nextcloud, headscale, cortex AI. (Nominatim excluded per #8.)

Pre-flight gates — status (2026-06-20)

  • PDM break-glass DONE: local admin@pam (Administrator) created on PDM CT100, login verified via API, Authentik realm untouched; cred in credentials, config backup at CT100 /root/access.bak-2026-06-20.
  • CT118 archivist rpcbind DONE: nftables rule restricts port 111 to source 192.168.1.240 (NFS server) only; NFS mount healthy, ruleset persisted. Now 100% LAN-internal.
  • mailcow CT108 DONE: backup durable on pi-nas (…/contabo-prewipe-2026-06/mailcow/, sha256-verified), live mail confirmed on edge1; CT108 destroyed (pct destroy 108 --purge, 2026-06-20) — config + disk image purged. edge2 now hosts CT100107.
  • data disk (92%)⏏️ DE-SCOPED: not a gate. OS package updates are roll-back-able (reinstall prior version); the migration-heavy apps that need real rollback (Authentik, Forgejo, Nextcloud, PeerTube) live on cloud/edge2, not data. Optional cleanup only (~6 GB obviously-safe: stale ISO, zimit temp) if ever wanted.

Rollback model (corrected): OS packages → reinstall the prior version (no VM snapshot needed). App DB-migration upgrades (Authentik/Forgejo/Nextcloud/PeerTube) → restore a quiesced DB dump (cheap), since reinstalling the old binary won't unwind a migrated schema.


Phase 1 — Execution Log (2026-06-20)

Status: COMPLETE. 26 guests patched (security-pocket only), every service validated back up, ~600+ security packages cleared, zero data loss. No host reboots (guests only). Method: snapshot-or-DB-dump → security-only apt (apt-get install --only-upgrade from the security pocket) → pct reboot → service-by-service validation.

Guests patched

  • utility: CT100, CT101, CT102, CT103, CT104, CT106, CT107, CT108, CT109, CT112, CT118, CT119
  • cloud: CT120, CT121
  • media: VM105, CT110, CT111
  • data: VM1130
  • edge2: CT100, CT101, CT102, CT103, CT104, CT105, CT106, CT107

(CT107 = headscale — see Critical Finding 1 below.)

DB dumps / rollback points captured (reusable in Phase 2)

Guest Location Notes
authentik CT105 /root/authentik-db-20260620.sql (126 MB)
matrix CT106 /root/matrix-db-20260620.sql (60 MB)
forgejo CT103 /root/forgejo-db-20260620.sql
wordpress CT101 dump taken on CT101
peertube CT110 DB dump
opentakserver CT109 PG dump + LVM snapshot
central CT104 PG dump + snapshot
immich CT120 /root/immich-predates-20260620.sql
headscale CT107 /root/headscale-db-20260620.sqlite

Note: many edge2 and cloud CTs' storage does NOT support pct snapshot — DB dump or package-reinstall is the rollback model for those.

Incidental fixes made during Phase 1

  • (a) CT110 peertube — immutable /etc/resolv.conf blocked reboot. The file had chattr +i set (intentional NordVPN dns protection). Cleared the immutable flag to allow the reboot, then verified the flag was restored and DNS remained healthy after boot.
  • (b) CT111 mcc — DNS hijacked to unreachable MagicDNS. Tailscale accept-dns was redirecting DNS to a MagicDNS address that was not reachable from this CT. Disabled tailscale accept-dns, set 1.1.1.1 / 8.8.8.8 persistently.
  • (c) PostgreSQL on central CT104 moved 16.13→16.14 as part of the security-pocket apt pass.
  • (d) Fleet-wide stale /etc/hosts fix — see Critical Finding 2.

Critical Finding 1 — CT107 headscale reboot-survival — FIXED 2026-06-20

On pct reboot 107, the fleet headscale coordinator crash-looped and required approximately 15 minutes of manual recovery. Two root causes:

  • (a) Docker compose bridge headscale_default came up linkdown after the unprivileged-LXC reboot.
  • (b) Bootstrap chicken-and-egg: Docker port-binds headscale to 100.64.0.38 (tailscale0), but tailscale0 needs headscale (the coordinator) to come up first.

Recovery required: compose downdocker restart → manually add 100.64.0.38/32 to tailscale0 → compose up → re-auth CT107's own tailscale node with a fresh preauthkey → temporarily DNAT vpn.echo6.co through CT107's internal IP to bootstrap, then revert Caddy config and clean up iptables.

Resolution (2026-06-20): Root cause was a triple boot deadlock: (a) headscale/headplane Docker ports were bound to the tailscale IP 100.64.0.38 (which only exists after tailscaled connects to headscale — circular); (b) the headscale-stack.service systemd unit waited on tailscale-online.target (also needs headscale up); (c) only_start_if_oidc_is_available: true blocked startup on Authentik reachability.

Fix applied (all originals backed up .bak-20260620):

  1. Rebound both ports 100.64.0.38:8084/310010.10.10.25:8084/3100 (CT's static internal IP, always up at boot) in /opt/headscale/docker-compose.yml.
  2. headscale-stack.serviceAfter=docker.service only (dropped tailscale-online.target dependency).
  3. only_start_if_oidc_is_available: false in /opt/headscale/config.yaml.
  4. Repointed edge2 /etc/caddy/Caddyfile vpn.echo6.co upstream (and /admin* headplane) to 10.10.10.25.

PROVEN by a controlled pct reboot 107: headscale self-healed in ~45 seconds with ZERO manual intervention/health 200, RestartCount 0, clean logs, 32 fleet nodes online.

Known minor leftover (cosmetic, non-blocking): CT107's OWN tailscale client node (100.64.0.38) stays in NoState because its tailscaled ControlURL points at its own IP (http://100.64.0.38:8084) — a self-bootstrap chicken-and-egg. The headscale SERVICE is fully healthy and the fleet is coordinated; CT107 is managed via the edge2 host OOB path, so its own tailnet membership is optional. Candidate follow-up: point CT107 tailscaled ControlURL at https://vpn.echo6.co (now always reachable via Caddy→10.10.10.25) so it self-registers cleanly.

🔴 Critical Finding 2 — fleet-wide stale /etc/hosts broke coordinator connectivity

7 fleet nodes — data, cloud, media, utility hosts, plus caddy CT101, cobalt CT112, and peertube CT110 — had a stale 5.189.158.149 vpn.echo6.co line in /etc/hosts left over from before the 2026-06-19 headscale migration to edge2. This pinned vpn.echo6.co to edge1 (now mail-only), so tailscaled hit Mailcow's TLS cert and could never reach the real coordinator — affected nodes showed OFFLINE in headscale while coasting on persistent WireGuard tunnels (still SSH-reachable, masking the problem).

Fixed 2026-06-20: removed the stale line and ran tailscale up on all 7; all confirmed ONLINE in the coordinator. /etc/hosts.bak-20260620 backups left on each host.

IMPACT on Phase 3: this class of bug means a reboot of an affected host while its control connection was stale would have failed to rejoin the tailnet — the primary lockout risk. Now that all nodes have valid control connections, reboots should re-register cleanly. However: verify each node is ONLINE in headscale before AND after any Phase-3 reboot.

Headscale stale-node cleanup: deleted dead nodes mailcow (destroyed CT108) and a stale peertube duplicate, both on 2026-06-20.


Phase 2 — Execution Log (2026-06-21)

Status: COMPLETE (2026-06-21). All app upgrades done — high-priority, security-critical, low-urgency batch, and cortex AI stack. Only Phase 3 (platform/reboots) and the deferred Nominatim project remain.

Completed upgrades

All targets below were validated healthy post-upgrade; rollback DB dumps retained on the respective CTs under /root.

Wave 1 — cloud apps

App Host From → To Notes
Immich cloud CT120 2.5.6 → 2.7.5 XSS + ACL-bypass fixes; stack includes Valkey 9 and vectorchord PG
PeerTube media CT110 8.0.2 → 8.2.1
Nextcloud AIO cloud CT121 NC 32.0.4 → 33.0.5 (major); mastercontainer updated Major NC version; AIO orchestrated update path

Wave 2 — media stack (media VM105)

App From → To Notes
Jellyfin 10.11.6 → 10.11.11
Sonarr → 4.0.17
Radarr → 6.2.1
Prowlarr → 2.4.0
Navidrome → 0.62.0
SABnzbd 4.5.5 → 5.0.4 (major) Major upgrade; clean migration
Lidarr DEFERRED Image maintainer hasn't shipped v3 — needs image swap to linuxserver to go v3; left on current
Jellyseerr INTENTIONALLY HELD on preview-OIDC tag Stable lacks OIDC support; switching would break SSO login

Authentik (edge2 CT105) 2025.12.4 → 2026.5.3

Done as 3 sequential hops: 12.4 → 12.6 → 2026.2.4 → 2026.5.3. pg_dump before each hop; migrations clean each time. SSO verified end-to-end by Matt logging into navi. Closes the May-2026 CVE waves. (ak app version 5.2.15.)

Forgejo (edge2 CT103) 14.0.5 → 15.0.3

EOL-branch migration; v15 schema migrations applied cleanly; web 200, API reports 15.0.3. Pre-15 DB dump at /root/forgejo-db-pre15-20260621.sql.

Headscale 0.28 → 0.29.1 (both instances)

  • CT106 (utility, IdahoMesh 3-node mesh) — upgraded first as rehearsal; native systemd binary.
  • CT107 (edge2, fleet coordinator, Docker) — 32/32 nodes reconnected post-upgrade; health 200; boot-survival fix preserved (ports remain on 10.10.10.25). 0.29 breaking changes documented in meshtastic-headscale-runbook (key changes: randomize_client_port removed — was a hard blocker; ephemeral key config nested; minimum Tailscale client 1.80.0; bare ACL * now tailnet-only).

OpenTAKServer (utility CT109) 1.7.10 → 1.7.12

Clean; all 9 TAK services healthy.

Key decision — RabbitMQ LEFT on 3.12.1 (EOL)

Empirically confirmed during the OTS update: updating OTS to 1.7.12 does not touch RabbitMQ — OTS 1.7.12 runs fine against RabbitMQ 3.12.1/Erlang OTP 25. Forcing RabbitMQ → 4.x is the wrong move: OTS is built and tested against 3.12, and 4.x has breaking changes (queue-mirroring removal, feature-flag requirements) that would likely break OTS, which does not appear to support 4.x. AMQP/MQTT ports are localhost-bound, so exposure is low. The EOL 3.12 is a low-priority latent risk to revisit only if/when OTS officially supports RabbitMQ 4.x. NOT a current action item.

Incidental fixes and side-work during Phase 2

  • CT107 boot-survival fix (applied earlier in the effort, during Phase 1 resolution) — rebound headscale/headplane ports to 10.10.10.25, dropped tailscale-online.target dependency, disabled only_start_if_oidc_is_available gate, repointed edge2 Caddy; proven by reboot self-heal in ~45 s. Also corrected CT107's own tailscale node ControlURL to vpn.echo6.co so it self-registers cleanly.
  • Utility node incident (resolved): a batch delete of 9 LVM-thin snapshots triggered an SSD TRIM/discard storm that spiked I/O and load transiently; compounded by CT103 argus running hot (transcription + docker-compose build churn). Matt migrated argus to the cloud node, resolving the issue; utility load returned to normal. LESSON: delete thin-pool snapshots one at a time — not in a batch — to avoid the discard storm.
  • Nextcloud: granted matt@echo6.co the NC admin role. (user_oidc has no group-claim sync, so this is durable across SSO logins.)
  • Radarr: set up a \\192.168.1.160\manual SMB drop folder on the same NFS export as the library (atomic-move imports) for manual movie filing.
  • Snapshot hygiene: all rollback snapshots cleaned up after validation — Phase 1 presec-*, Phase 2 prewave2-*, OTS pre-ots-* snapshots all removed.

Completed — low-urgency batch and cortex AI stack

Low-urgency batch — DONE:

  • Caddy (utility CT101): 2.10.2 → 2.11.4
  • CouchDB (edge2 CT104 livesync): 3.4 → 3.5.2 — all 5 DBs intact
  • valkey-8 sidecar (utility CT102 searxng): bumped to latest 8.x
  • NATS (utility CT104 central): 2.14.0 → 2.14.2 — all 12 JetStream streams intact
  • MediaMTX (CT109): 1.13.0 → 1.19.1
  • Mumble (CT109): already latest in Ubuntu repo (1.5.517 — no upstream action possible without going off-distro)

cortex AI stack — DONE (protected host; containers only; NO reboot, NO driver touch):

  • qdrant: 1.16.3 → 1.18.2
  • tei: 1.7 → 1.9 (bge-m3 on GPU)
  • ollama: 0.16.1 → 0.30.10 (vault-tagger 100% GPU)
  • open-webui: 0.8.1 → 0.9.6
  • Both .ref vault-engine deps (ollama vault-tagger + tei bge-m3) confirmed working on GPU.

Minor follow-ups noted (non-urgent)

  • Caddy expired cert: Caddy flagged an EXPIRED unmanaged cert for navidrome.echo6.co (pre-existing condition, not caused by the upgrade) — fix if that hostname matters.
  • MediaMTX deprecated config params: 1.19 uses deprecated param names (protocolsrtspTransports, encryptionrtspEncryption) — works now (warnings only); rename before a future MediaMTX release removes them.
  • Rollback artifact cleanup: retained rollback artifacts to clean once comfortable: /root DB dumps on their respective CTs (authentik hop1/2/3, forgejo, headscale107, immich, peertube, OTS), binary backups (caddy.bak, nats-server.bak, mediamtx.bak on their CTs), and openwebui DB backup on cortex (/home/zvx/openwebui-webui.db.bak-20260621).
  • Vestigial utility exit-node route: stale 0.0.0.0/0 exit-node route on utility — optional cleanup (harmless post-0.29 Headscale).

Separate deferred projects


Phase 3 — Execution Log (2026-06-21/22)

Status: COMPLETE (2026-06-22) — all five PVE nodes on 9.2.3/kernel 7.0.12, pi-nas on OMV 8.4/kernel 6.18, cortex NVIDIA driver + DKMS + container-toolkit updated, GPU passthrough verified. Phase 3 has no open items.

Pre-flight findings

  • Cluster echo6-cluster 5/5 quorate, quorum 3, NO HA configured (clean guest stop/start, no fencing).
  • BLOCKER found + fixed: data, cloud, media had NO Proxmox APT repo configured at all — added pve-no-subscription (trixie, matching utility/toc); all 5 then saw the 9.2 stack.
  • Note: PVE 9.2 ships kernel 7.0 as the new default (proxmox-default-kernel) — nodes boot 7.0.12-1-pve, not 6.17.13 as the audit predicted.

Cluster nodes — ALL upgraded to PVE 9.2.3 / kernel 7.0.12-1-pve (QEMU 11, LXC 7)

One at a time; cluster stayed 5/5 quorate throughout.

  • media — done first (canary). Incidental: CT110 peertube failed to auto-start (recurring /etc/resolv.conf immutable-flag vs LXC pre-start-hook conflict) → permanently fixed: the flag was an obsolete workaround (NordVPN set dns no longer overwrites resolv.conf), cleared it + enabled NordVPN auto-connect, so future reboots won't trip it.
  • data — recon-vm (VM1130) healthy; virtiofsd-{kiwix,library,nav} auto-recovered this time; nginx needed the one expected restart (pre-existing mesh-DNS startup race).
  • utility — all 11 CTs; central JetStream intact (12 streams); caddy proxying, mesh headscale (CT106, 3 nodes), mesh-bridge (CT107) both tailnets, OTS stack all healthy. CT102 searxng needed a manual pct start (transient auto-start miss, no persistent fault).
  • cloud — done last (per operator). immich (4 containers + API 200), nextcloud (12 AIO containers, occ healthy, v33.0.5), argus running. (Note: argus CT103 on cloud has argus-capture missing / argus-transcribe masked — operator's in-progress argus→cloud migration, not from the upgrade.)

pi-nas FULLY DONE (OMV 8.4 + kernel 6.18)

OMV (2026-06-21): OMV 8.1.0-2 → 8.4.0-3 + Debian userspace (Docker, OpenSSL, salt, tailscale). All 5 NFS exports (arr/immich/nextcloud/peertube/data) serving; immich+nextcloud mounts confirmed OK.

Kernel jump (2026-06-22): 6.12.62 → 6.18.34+rpt-rpi-2712 — operator physically present. Pre-jump: full /boot backup at /root/boot-backup-pre6.18-20260622.tar.gz (154 MB) + old kernel left in place as fallback. Booted clean in ~15 s; all 5 NFS exports serving; immich+nextcloud mounts recovered; SD card healthy. Only package remaining HELD: rpi-eeprom (SPI bootloader — intentionally untouched; optional to update later).

Important architecture note recorded: pi-nas (2.8 TB RAID1) is the NFS storage backend for immich's photo library (644 GB), nextcloud files, peertube, and the *arr library — so a pi-nas reboot stalls those services' storage.

toc + cortex DONE (2026-06-22, run from matt-desktop)

  • cortex: NVIDIA driver 580.159.03 → 580.167.08 + DKMS (built for running kernel) + nvidia-container-toolkit 1.19.1; all AI containers healthy on GPU — vault-tagger generates (Ollama), TEI bge-m3 embeds (200 OK from localhost:8090), Qdrant healthy.
  • toc: PVE 9.1 → 9.2.3 / kernel 7.0.12-1-pve. GPU passthrough (vfio) survived the 7.0 kernel — the key risk, verified.
  • Cluster 5/5 quorate with toc rejoined; cortex VM150 auto-started; no leftover snapshot.

Phase 3 has no open items.

Wrap-up (2026-06-22)

Rollback-artifact sweep done. Approximately 10.4 GB of campaign DB dumps and binary/config backups removed fleet-wide. Largest single item: a 7.79 GB central Postgres dump on utility CT104. Retained intentionally: /root/boot-backup-pre6.18-20260622.tar.gz on pi-nas (kernel rollback artifact; 154 MB) and the durable mailcow backup on pi-nas (…/contabo-prewipe-2026-06/mailcow/).

Accepted final state — operator-acknowledged decisions (2026-06-22). These are closed decisions, not TODOs.

  • Guest OS = security-only. LXC containers and VMs received security-pocket apt updates only (the intentional Phase 1 scope). Non-security package drift (e.g. Docker CE versions, miscellaneous libs) was deliberately not swept with a full apt full-upgrade. The PVE hosts, pi-nas, and cortex did receive full upgrades. Operator accepted this state. A full guest apt full-upgrade ("Phase 1.5") remains an option if ever wanted.
  • Apps capped by external factors (decisions, not failures): RabbitMQ 3.12 left by decision — OTS depends on it and 4.x would break it; Lidarr v2 — the lidarr-on-steroids image maintainer has not shipped v3, would require an image swap; Jellyseerr on preview-OIDC — kept because stable 3.3.0 lacks OIDC/SSO support; Mumble 1.5.517 — newest version in the Ubuntu 24.04 repo, upstream 1.5.901 is not available without going off-distro.
  • Deferred project: Nominatim v5 re-import — spun off to nominatim-v5-reimport.
  • Optional cosmetic follow-ups (non-urgent, non-blocking): navidrome.echo6.co expired unmanaged cert; MediaMTX deprecated config param names (protocols/encryption); rpi-eeprom held on pi-nas; vestigial utility exit-node route (0.0.0.0/0).
  • Kernel summary: all 5 PVE hosts on kernel 7.0.12-1-pve; LXC containers share the host kernel (7.0); VM guests (recon-vm, arr VM105) have their own Ubuntu kernels (security-patched during Phase 1, not necessarily absolute-latest upstream); pi-nas on 6.18.34+rpt-rpi-2712; cortex kernel updated as part of the toc+cortex window.

Execution Runbook (Meticulous)

Step-by-step plan to bring every application and package current. Per-app target versions live in Application Currency above; this section is the how/when/order. Execution model: scoped Sonnet task per target → Opus verifies output → report to Matt; approval gate before every load-bearing or reboot step.

Standing guardrails (apply to EVERY step)

  • Protected hosts — cortex & toc — never in a bulk pass. They get a dedicated, manual window. Note the coupling: toc reboot drops cortex (VM150), which is the Claude Code host — expect to lose the control session during that window; drive it from elsewhere or accept the outage.
  • One target at a time. Verify health before moving to the next.
  • Snapshot/backup before each mutating step: pct snapshot / qm snapshot (or vzdump) the guest first; DB dump before any DB-bearing upgrade; run the app's own backup where it has one (Nextcloud AIO, mailcow).
  • No host reboots outside the explicit Phase 3 windows. Security upgrades that set reboot-required are applied but activation deferred to Phase 3. Host OS-security apt for the four cluster nodes is NOT applied in Phase 1 — it rides the Phase 3 full-upgrade window.
  • Rollback for DB-bearing apps = quiesced logical dump + old binary together. For authentik, forgejo, nextcloud, peertube, and immich the rollback unit is a pg_dump / AIO borg / peertube dump taken with the app stopped and the old binary still in place. "Redeploy previous image tag" is NOT valid after forward-only migrations — Django/TypeORM migrations do not reverse, and a pct snapshot of a hot separate-volume DB may restore torn. Record before/after version in the tracking table.
  • Every result must survive a reboot (standing infra policy).

Load-bearing dependencies (drive the ordering)

  • Authentik (edge2 CT105) = SSO. Upgrading it briefly breaks login to everything behind it. Confirm break-glass admin access first; do in a low-traffic window; verify dependent-service logins after.
  • Caddy (utility CT101) = home-services ingress. A restart blips all home services — fold its update into a deliberate moment, not mid-day.
  • Headscale ×2: edge2 CT107 = main fleet tailnet (34 nodes, vpn.echo6.co, ControlURL https://vpn.echo6.co → 184.174.35.153) — HIGH lockout risk. utility CT106 = IdahoMesh sub-tailnet (vpn.idahomesh.com, 3 nodes) — low risk. Upgrade CT106 first as rehearsal. Edge2 CT107 upgrade MUST be driven from an out-of-band path (edge2 host console / direct SSH to 184.174.35.153, NOT over vpn.echo6.co); extend node key-expiry first; dump DB off-tailnet before touching it.
  • Nominatim 4→5 is NOT a routine update — it's a re-import project (see Phase 2C).

Phase 0 — Pre-flight (no app changes yet)

STEP 1 (FIRST — gates everything mutating): Verify break-glass and out-of-band access. Confirm a non-Tailscale, non-Caddy path to every node: Proxmox/Contabo console for edge2, LAN 192.168.1.x SSH for home nodes, PVE noVNC/pct enter for each CT. Confirm PVE/Proxmox web auth is PAM/local, NOT behind Authentik. Prove it: log into Contabo/PVE console; pct enter 105/106/107; confirm LAN SSH; locate break-glass creds in .ref/credentials. Verify Authentik akadmin (private browser session) + each protected app has a local-admin fallback (Forgejo admin user, Nextcloud occ, PDM/Vaultwarden local). Every node + CT shell must be reachable WITHOUT Authentik/headscale/Caddy. Losing the management path mid-run is the single highest-consequence failure mode.

STEP 2: Measure snapshot headroom (per node). Run pvesm status + df -h on data/utility/cloud/media/toc/edge2. Record the storage backend per node (LVM-thin/ZFS/dir-qcow2) and its over-provision failure mode. Hard numeric rule: never start a snapshot below ~20% free; require ≥2× the largest guest's expected write-delta. Cloud node CT120/CT121 free space must be measured now — not assumed. Identify where vzdump lands; if it's data, freeing data is a fleet-wide rollback prerequisite.

STEP 3: Free space on data (92% full, ~73 GB free) — sized to the largest planned snapshot (VM1130 recon-vm / nominatim+overture+padus). Prune stale vzdump/snapshots/ISOs; docker image prune on recon-vm (do NOT remove the pinned nominatim:4.5 image). Confirm nothing deleted is the only copy of a backup or live NAS data. Verify: usage <~85% on snapshot-backing storage; test snapshot succeeds then remove. Blocks data host Phase 3 and canary use.

STEP 4: Verify off-host restorable backups (distinct from per-step snapshots) for every stateful guest: central PG/NATS (CT104), opentakserver PG + RabbitMQ definitions (CT109), forgejo PG + repos (CT103), matrix Synapse PG (CT106), nextcloud AIO borg (CT121), edge2 livesync CouchDB (CT104). At least one copy must live off the node being changed.

STEP 5: Capture pre-change baseline HEALTH snapshot. Record per-target up/down, container states (docker ps/pct/qm status), key endpoint 200 checks, and apt list --upgradable/security counts. Note known oddities: jellyseerr on preview-OIDC dev tag, mailcow CT108 already STOPPED, pinned-stale nominatim, CT118 archivist rpcbind on 0.0.0.0:111.

STEP 6: Decision gates. Close genuinely-open decisions before Phase 1/2/3 can start: daemon-restart tolerance sign-off (Open Dec #7); verify post-cutoff CVE claims against primary advisories with owner+URL per claim; resolve edge2 host-kernel CVE question fully (pin uname -r, check repo kernel availability, decide reboot yes/no — not left conditional); define and announce maintenance windows in America/Boise naming user-facing blips.

STEP 7: edge2 CT108 mailcow — DONE. CT108 destroyed 2026-06-20 (pct destroy 108 --purge); config + disk image purged. Backup verified durable on pi-nas (…/contabo-prewipe-2026-06/mailcow/, sha256-verified); edge1 confirmed serving mail (MX/A → 5.189.158.149). edge2 now hosts CT100107.

Phase 1 — Guest OS security packages (no reboot; hosts folded into Phase 3)

Guests only in Phase 1 (low blast radius). The four cluster hosts' OS-security apt rides their Phase 3 full-upgrade window — do NOT patch data/utility/cloud/media host OS in Phase 1 (avoids a mixed-state libc/openssl across the entire Phase 2 app campaign; both phases get cleaned up in one reboot anyway). edge2 host is already fully patched — no action. toc excluded (Phase 3, with cortex).

Per guest target: apt-get update → snapshot → apply security-pocket upgrades → needrestart policy (see below) → verify service health and security-upgradable count hits 0. Kernel/libc security that flags reboot-required → applied, reboot activation deferred to Phase 3.

needrestart policy:

  • Stateful guests — NEEDRESTART_MODE=l (list-only), then deliberate bounce: CT104 central (PG16/NATS/JetStream), CT109 opentakserver, CT110 peertube, VM1130 recon-vm. Stage the packages; take a logical DB dump with the app quiesced; bounce each service deliberately in controlled order. A JetStream restart mid-write loses messages on at-most-once pipelines; a RabbitMQ/redis bounce mid-job is not a clean "blip."
  • Lockout-critical daemons — NEEDRESTART_MODE=l + deliberate controlled bounce: Authentik CT105, headscale instance(s), Caddy CT101/CT111, sshd on any host you're connected through. Run needrestart -r l to see what would restart; bounce deliberately so a libc/openssl bump can't auto-bounce auth/tailnet/SSH daemons out from under the operator.
  • All other guests: needrestart -r a (automatic) is acceptable — seconds of blip, no data risk.

Ordering:

  1. Canary first (utility CT112 cobalt idle, or CT102 searxng) — prove the full procedure on a low-stakes guest before touching never-patched targets.
  2. Worst-first (after canary proven): utility CT119 mesh-territory (91 sec, never patched), CT108 meshai (79 sec), cloud CT120/CT121 guest-OS (101/75 sec), media CT110 peertube (37 sec), utility CT109 opentakserver (31 sec), CT104 central (38 sec, incl. PG 16.13→16.14 + NATS).
  3. Remaining guests: all other utility/cloud/media/edge2 CTs + data VM1130 (recon-vm). Gate data VM1130 step on data disk remediation verified DONE.
  4. edge2 CT106 matrix / CT107 headscale (NO Tailscale client): drive via pct exec from the edge2 HOST (root@184.174.35.153) ONLY — never over service ingress. Treat as own singleton steps, not part of edge2 batch. needrestart -r l; do NOT restart headscaled/synapse unless required.
  5. Control-plane singletons LAST: home-ingress Caddy CT101 and headscale instance(s) — each with DB backup + node-stays-joined verification, never inside a batch.
  6. cortex VM150 (PROTECTED): security apt manually in the toc+cortex window (Phase 3 coordination); not in bulk.
  7. Cross-cutting: Tailscale 1.94→1.98 and Docker CE→29.6 ride along Phase 1 under the same snapshots; EXCLUDE edge2 CT106/CT107 from Tailscale bump (no TS client).
  8. pi-nas: Debian security pocket with linux-image-* and openmediavault* packages held; kernel upgrade rides Phase 3.

Phase 2 — Application updates

Hard isolation rule: {Headscale, Authentik, Caddy, Forgejo} are NEVER in the same maintenance window — at least one of {tailnet, SSO, ingress, git} must always be a known-good recovery path.

2A — Load-bearing SSO and control-plane apps (in this order):

  • Authentik (edge2 CT105) — FIRST, sequential hops, pg_dump per hop.

    • Pre-flight before any hop: (1) Prove break-glass: log in as akadmin in a private browser session; confirm every protected app has a working local-admin fallback (Forgejo admin user, Nextcloud occ, PDM/Vaultwarden local, PVE PAM). (2) Audit/rename duplicate group names — the Django migration fails loudly on dupes. (3) If local /media storage is used, stop the service, mv ./media ./data/media, rewrite compose volume, before starting any new version.
    • Hop sequence (no skipping): 2025.12.4 → 2025.12.6 (backported security) → 2026.2.x2026.5.3. One hop at a time: pct snapshot + pg_dump (labelled per hop) → pull new tag → compose up -d → migrations run → verified-login gate (all dependent SSO logins tested) → abort+restore on FIRST migration error. Verify all dependent SSO logins BEFORE upgrading any of the dependents.
    • Rollback: compose down → restore prior tag + dump; or pct rollback for the full CT.
  • Caddy (utility CT101 + media CT111) — early in Phase 2, before home-service app verification. Pair CT101/CT111 in one deliberate low-traffic window (separate from Authentik, Forgejo, and headscale windows). Validate caddy validate; spot-check Caddy↔Authentik path; document a direct Caddy-bypassing Authentik admin URL as break-glass. Verify all home services reachable through the new ingress before proceeding with other app upgrades.

2B — Self-managed updaters (snapshot → run native updater → verify):

  • OpenTAKServer (utility CT109) — app upgrade first, RabbitMQ SEPARATE:

    • Run OTS upgrade script (upgrades OTS + webUI + schema only). Snapshot first. Verify OTS web/API up; CoT works; TAK clients reconnect.
    • RabbitMQ 3.12.1→4.2.8 is a SEPARATE sub-task (decoupled, after OTS app step). The OTS updater does NOT migrate RabbitMQ, Erlang, or feature flags — the EOL RabbitMQ 3.12 would silently remain. Decoupled upgrade: pct snapshot + rabbitmqctl export_definitions → upgrade Erlang to ≥26 → 3.12→3.13.x → rabbitmqctl enable_feature_flag all (confirm all enabled, no classic-queue-mirroring config remains) → only then 4.x. A naive 3.12→4.x jump refuses to boot. Gate "OTS done" on RabbitMQ 4.x up + TAK clients reconnecting before accepting the bundled MediaMTX/Mumble bumps.
    • Rollback: restore snapshot + definitions export.
  • Immich (cloud CT120) — prune old images + docker image prune before snapshot; disable Immich background regeneration until free space is confirmed; docker compose pull && up -d. Do NOT interrupt first-boot 2.5→2.7 DB migration. Also clears Valkey 9.0.2→9.1.0. Rollback = restore pct snapshot + pg_dump taken with the app stopped (not a hot-DB snapshot alone). Verify web, mobile sync, ML.

  • Nextcloud AIO (cloud CT121) — AIO borg backup → stop → update mastercontainer → trigger child update via AIO UI (:8080). A stale AIO may need two cycles. Rollback unit = AIO borg backup (mastercontainer downgrade is not supported — borg is the only path back). Verify occ status 32.0.11; 12 containers green; integrity clean.

2C — Versioned apps (snapshot → bump/upgrade → verify):

  • Forgejo (edge2 CT103) — branch migration 14.x→15.0.3 (separate window from Authentik).

    • Pre-steps: git clone --mirror echo6-docs (and other critical repos) to an OFF-edge2 location; pause the echo6-docs-autocommit cron so recovery docs survive an edge2/Forgejo failure; run forgejo doctor check --all [--fix] to repair stopwatch/tracked_time inconsistencies BEFORE backup+bump. Read 15.0 breaking changes + forward-only FK migration notes.
    • pct snapshot 103 + repos volume + pg_dump + app.ini → bump tag 14→15 → pull/up → migrations → verify.
    • Rollback: restore 14.0.5 image + volume + dump.
  • Headscale 0.28→0.29.1 — utility CT106 FIRST (rehearsal), edge2 CT107 LAST.

    • CT106 (IdahoMesh, 3 nodes, low risk): standard 0.28→0.29 upgrade guide; pct snapshot + DB backup; stop → upgrade binary → config migrate → start. Verify nodes list all joined; verify headplane compatibility.
    • CT107 (main fleet, 34 nodes, HIGH lockout risk — edge2 front door): extend/disable node key-expiry on all 34 nodes BEFORE starting; pct snapshot 107 + dump headscale DB off-tailnet (NOT via vpn.echo6.co); keep a second authenticated SSH session open to 184.174.35.153. Drive entirely from edge2 host console / direct SSH to root@184.174.35.153 — NOT over the tailnet being upgraded. Verify boot chain: edge2 host → CT107 autostart → headscale up → vpn.echo6.co resolves → Caddy proxies. Rollback: restore snapshot via console (not via tailnet).
    • NOT in the same window as Caddy, Authentik, or Forgejo.
  • Media stack (media VM105):

    • qm snapshot 105 before batch. SABnzbd 4.5→5.0: pause/empty queue; audit custom post-proc scripts (empty_postproc removed in 5.x, scripts now run on failed jobs); pin 5.0.4 tag; pull/up; verify queue intact. Lidarr 2.x→3.x: pin v3; pull/up; DB migration on start; verify library+indexers+download client. Sonarr/Radarr/Prowlarr (1-2 ver): routine bumps after SABnzbd/Lidarr. Jellyfin 10.11.6→10.11.11: pin tag; pull/up. Navidrome: routine bump.
    • Jellyseerr preview-OIDC→stable 3.3.0: DECISION GATE before retag. preview-OIDC → stable 3.3.0 is a schema DOWNGRADE; TypeORM stable may refuse to start against a dev-migrated DB. Confirm the correct destination artifact (jellyseerr stable vs "Seerr" successor) and, if OIDC is in active use, confirm 3.3.0 parity. Test stable image against a COPY of the DB first. If parity unclear → STOP and report.
    • Verify each UI; snapshot before, per-app snapshots for majors.
  • PeerTube (media CT110) — quiesce import/transcode jobs; pct snapshot 110 + pg_dump peertube_prod; run PeerTube upgrade script 8.0.2→8.2.1; restart; verify. Rollback = pct rollback + restore dump.

  • cortex AI stack (PROTECTED — manual, own window): ollama 0.16.1→0.30.10 (snapshot models volume — the only rollback after pulling new models; upgrade; verify existing ollama run without re-pull; verify vault-tagger at localhost:11434 and TEI at localhost:8090 still work); open-webui 0.8.1→0.9.6; qdrant 1.16→1.18; tei 1.7→1.9. Pull images, verify.

2D — Lower-urgency batch: caddy (minor), couchdb livesync (3.4→3.5.2), valhalla, photon, kiwix, NATS 2.14.0→2.14.2, valkey-8 sidecar (8.1.5→8.1.8) — snapshot + update as convenient.

Vaultwarden, PDM, WordPress, Synapse/Element/MAS, obsidian, recon-vm Postgres → already current, skip.

2E — Separate project (not in this campaign):

  • Nominatim 4.5→5.3 + Photon 1.1→1.2 (recon-vm) — requires full OSM re-import into a SEPARATE DB/instance; keep nominatim:4.5 image + data until 5.x validated; transient import disk (flat-nodes tens of GB) cannot live on the 92%-full data pool; Photon 1.2 AFTER nominatim 5.x validated, never concurrently. Schedule as its own effort.

Phase 3 — Platform & reboot windows (coordinated, approval-gated)

🔴 BLOCKED — edge2 reboot was NOT permitted until the CT107 headscale boot-survival fix was in place. CLEARED 2026-06-20 — the CT107 headscale boot-survival fix is applied and proven (see Critical Finding 1 in the Phase 1 Execution Log). The triple boot deadlock (tailscale-IP port bind + tailscale-online.target wait + OIDC gate) has been resolved; headscale now self-heals in ~45 seconds after a CT107 or edge2 reboot with zero manual intervention. The Phase-3 edge2 reboot blocker is lifted. The remaining gates below (corosync quorum + Finding-2 headscale-ONLINE check per node) still apply.

🔴 GATE for every Phase-3 host reboot (Critical Finding 2): Before AND after rebooting any node, confirm that node is ONLINE in the headscale coordinator (headscale nodes list). Stale /etc/hosts entries previously masked coordinator disconnects behind coasting WireGuard tunnels — the stale entries have been removed fleet-wide, but verify ONLINE status at each reboot step to catch any regression. This gate is IN ADDITION TO the corosync-quorum gate below.

Corosync cluster rules (apply to EVERY node reboot in this section):

  • The five PVE nodes (data, utility, cloud, media, toc) are ONE echo6-cluster, quorum = 3 of 5. Losing quorum makes /etc/pve read-only fleet-wide and freezes all VM/CT start/stop/config on every surviving node.
  • Before EACH node reboot: pvecm status must show Quorate: Yes and Expected votes: 5. Also run ha-manager status — if any guest is HA-managed, the reboot triggers fencing/auto-migration rather than a clean local stop/start.
  • Reboot ONE node at a time. Wait for full 5/5 rejoin before the next node.
  • Disable HA migration and NO live migration during the mixed-version window (QEMU 10/11 boundary — running VMs keep QEMU 10 machine type until cold-started; cross-boundary migration fails).

Per-node procedure: apt update && apt full-upgrade (so proxmox-ve/qemu/lxc metapackages pull) → this also clears the held host OS-security packages → reboot → verify pveversion + all guests return with onboot=1. Confirm pvecm status shows 5/5 before proceeding.

Pre-window for each node: vzdump all guests to an EXTERNAL target (not the local pool, never data) so a guest that won't start under LXC 7 can be restored elsewhere. Read PVE 9.2 + LXC 7 release notes for unprivileged-CT/cgroupv2/apparmor breaking changes before the first node.

Canary node order (lowest blast-radius with headroom first): media (single arr VM + peertube, non-auth/non-DB-critical) → cloud (immich/nextcloud, after media proven) → data (after disk remediation verified; its near-full pool can fail the mandatory pre-reboot snapshot — do NOT use as canary until remediated) → utility LAST of the four (carries Caddy ingress + mesh/central stack — deliberately-scheduled full home-ingress + mesh-coordination outage, announce in advance) → toc+cortex LAST AND ALONE (see below).

Post-reboot re-verification gate (per node): after each node boots into 9.2/LXC7, re-run the Phase-2 health checks for every guest that received a major app upgrade on that node (CT120 immich, CT121 nextcloud, CT110 peertube, VM105 arr, CT109 OTS). The substrate changed under them.

toc + cortex — DONE LAST AND ALONE, after all four other nodes are quorate:

  • toc is the 5th cluster vote AND the GPU host for cortex (VM150) AND the management/Claude Code host. Rebooting toc removes a vote and kills the box you'd diagnose from. It must NEVER overlap any other node reboot (toc down + one other = 3/5 → one corosync flap loses quorum).
  • Name and TEST a concrete alternate control point (a non-cortex host with keys + tooling) BEFORE this window. Stabilize the tailnet well before it — never in the same period as a headscale upgrade.
  • cortex steps IN ORDER:
    1. qm snapshot 150 pre-nvidia + vzdump (from toc, before anything changes)
    2. apt --only-upgrade nvidia-container-toolkit; nvidia-ctk runtime configure; restart docker → verify --version=1.19 + docker run --gpus all nvidia-smi
    3. apt --only-upgrade nvidia-driver-580 nvidia-dkms-580; DKMS rebuild → verify dkms status shows installed + nvidia-smi=580.167 BEFORE rebooting toc. Keep the old driver package installed as a reinstall fallback.
    4. cortex Phase 1 security apt (if not already done in its own window)
  • Then reboot toc → verify toc rejoins (pvecm status 5/5) → verify VM150 autostart + cortex GPU passthrough return → final AI-stack health check (driver 580.167; all 5 containers; ollama/tei/qdrant/open-webui; vault-tagger+TEI engines; TS online; Docker 29.6).
  • The Claude Code host goes down here — run this step from the alternate control point.

pi-nas — standalone window, TWO reboots:

  • Confirm physical/serial console access (no remote KVM on an RPi); export config.xml off-box; confirm pi-nas is not mid-sync as a Syncthing/backup target before each reboot. Back up /boot + /boot/firmware; keep the old kernel installed as fallback.
  • Reboot 1: OMV 8.1→8.4 via OMV's own update path (NOT raw apt full-upgrade — use OMV UI Update Mgmt or omv-upgrade) → reboot → verify shares/SMB/NFS/omv-salt healthy.
  • Reboot 2 (AFTER reboot 1 verified): kernel 6.12→6.18 — apt install linux-image-arm64 (unhold); update bootloader; reboot ON APPROVAL → verify uname -r=6.18; LAN+tailnet up; shares mount; Docker up. Rollback: boot retained 6.12; restore /boot; on-site SD reflash.
  • Independent of cluster windows.

edge2 — CONDITIONAL (CT107 boot-survival blocker CLEARED 2026-06-20):

  • Previously also blocked on the CT107 headscale boot-survival issue — that blocker is resolved; headscale survives a CT107 or edge2 reboot and self-heals in ~45 seconds (see Critical Finding 1, fixed 2026-06-20).
  • Reboot only if Phase 0 resolved the DirtyFrag/copy.fail CVE question as "yes, needs a newer kernel" (Open Decision #5 recorded "no action / no reboot" — edge2 host already fully patched). If a reboot is ever required: give it a dedicated window AFTER Authentik CT105 + headscale CT107 are upgraded and verified; announce SSO+tailnet downtime; drive from a non-tailnet path; confirm all CTs auto-start. edge2 is a single SPOF for ingress — no failover. Do NOT fold into cluster windows.

Phase 4 — Cross-cutting (fold into earlier phases)

  • Tailscale 1.94→1.98 fleet-wide — rides along Phase 1.
  • Docker CE →29.6 wherever installed — rides along Phase 1/2 under the same snapshots.

Tracking checklist (tick per target)

Enumerated Target Checklist

Phase Target Action Snapshot Update method Verify Rollback Depends-on
0 Break-glass / OOB access matrix (ALL nodes) Verify non-Tailscale, non-Caddy path to every node + each lockout-critical CT (headscale, Authentik CT105, Caddy, Forgejo CT103); confirm PVE/Proxmox web auth is PAM/local not Authentik None (read-only) Log into Contabo/PVE console; pct enter 105/106/107; confirm LAN 192.168.1.x SSH; locate break-glass creds in .ref/credentials; TEST Authentik akadmin + each app local-admin fallback Every node + CT shell reachable WITHOUT Authentik/headscale/Caddy; akadmin login screenshotted N/A FIRST Phase 0 step; gates everything mutating
0 Per-node free space + snapshot backend pvesm status + df -h on data/utility/cloud/media/toc/edge2; record backend (LVM-thin/ZFS/dir-qcow2) + failure mode; size vs largest guests None Read-only; set numeric rule: never snapshot below ~20% free, require ≥2× write-delta Free GB recorded per pool; cloud node measured; vzdump target located N/A Precedes all snapshot-bearing steps
0 data host — disk remediation (92% full, ~73 GB) Free space sized to largest planned snapshot (VM1130) None (cleanup; confirm deletions are cache/backup not live) Prune stale vzdump/snapshots/ISO; docker image prune on recon-vm (keep nominatim:4.5) Usage <~85% on snapshot-backing storage; test snapshot succeeds then removed Restore from Forge/Syncthing/backup if a needed file removed Free-space capture; BLOCKS data host step + canary use
0 Off-host restorable backups (stateful guests) Verify last-good vzdump/PBS off the node for central, OTS, forgejo, matrix, nextcloud, edge2 livesync CouchDB N/A Read-only verification + on-demand pg_dump/app backup copied off-host ≥1 backup off the changing node, restore-testable N/A Precedes all mutating stateful steps
0 Baseline HEALTH capture (all targets) Record up/down, container states, endpoint 200s, apt upgradable/security counts; note oddities None pct/qm status, docker ps, curl probes, apt list --upgradable Baseline recorded; oddities logged (jellyseerr dev tag, CT108 stopped, nominatim pin, CT118 rpcbind) N/A Precedes Phase 1
0 Decision gates Close daemon-restart tolerance sign-off (Open Dec #7); verify post-cutoff CVE claims vs primary advisories w/ owner+URL; resolve edge2 kernel question (uname -r vs repo availability); define+announce Boise maintenance windows with user-facing blip list None Read-only / sign-off Each decision recorded before Phase 1/2/3 can start N/A Gates Phase 1/2/3
0 edge2 CT108 mailcow DONE — destroyed 2026-06-20 (pct destroy 108 --purge); backup durable on pi-nas (sha256-verified); edge1 confirmed serving mail N/A N/A CT108 purged; edge1 mail in+out confirmed N/A Complete
0 edge2 host kernel CVE question Decide if DirtyFrag/copy.fail require a reboot None (record pveversion -v) Compare installed proxmox-kernel vs verified-advisory fixed versions; escalate if no-sub repo lacks fix (no repo changes w/o approval) Decision (reboot yes/no) recorded N/A Gates edge2 Phase 3 row
1 Procedure canary — utility CT112 cobalt (idle) or CT102 searxng Prove apt→snapshot→security upgrade→needrestart→verify on low-stakes guest pct snapshot pre-phase1 apt-get update; security pocket only; needrestart -r l then deliberate Security-upgradable=0; service healthy; needrestart clear pct rollback Phase 0; runs BEFORE worst-first guests
1 utility CT119 mesh-territory (179/91, never patched) Security-pocket apt pct snapshot 119 pct exec; security pocket; needrestart -r a sec count=0; meshwars up rollback snapshot After canary
1 utility CT108 meshai (109/79) Security-pocket apt pct snapshot 108 pct exec; security pocket; needrestart sec=0; meshai up rollback snapshot After canary
1 utility CT104 central (PG16 16.13→16.14 + NATS/JetStream) Security apt incl. PG minor; STAGED, controlled bounce pct snapshot 104 + pg_dumpall NEEDRESTART_MODE=l; stage pkgs; quiesce JetStream publishers; bounce PG then NATS deliberately sec=0; select version()=16.14; JetStream/MQTT consumers reconnected rollback snapshot; restore dump Phase 0 #8; stateful carve-out
1 utility CT109 opentakserver (OS layer, 31 sec) Security apt; STAGED, do NOT touch RabbitMQ/MediaMTX/Mumble here pct snapshot 109 NEEDRESTART_MODE=l; stage; deliberate bounce sec=0; TAK clients reconnect rollback snapshot Precedes CT109 Phase 2 app work
1 media CT110 peertube (81/37, native PG16+redis) Security apt; STAGED controlled bounce pct snapshot 110 (PG rides in CT snap) NEEDRESTART_MODE=l; stop import/transcode jobs; snapshot redis; deliberate bounce sec=0; PeerTube plays; PG16/redis/nginx up; nordvpn up rollback snapshot Phase 0 #8
1 data VM1130 recon-vm (34/10, PG16+navi-backend) Security apt; STAGED controlled bounce qm snapshot 1130 pre-phase1-apt NEEDRESTART_MODE=l; deliberate PG + navi-backend bounce sec=0; psql overture/padus; navi-backend/recon.py/photon/kiwix up qm rollback data disk remediation done first
1 cloud CT120 immich (191/101) Guest security apt; Tailscale 1.94→1.98 + Docker CE→29.6 ride along pct snapshot 120 pre-phase1-os pct exec; security pocket; docker bounce acceptable sec=0; immich stack Up; web reachable; TS=1.98 pct rollback Before cloud host step + CT120 app upgrade
1 cloud CT121 nextcloud AIO (98/75) Guest security apt; TS+Docker ride along pct snapshot 121 pre-phase1-os pct exec; security pocket sec=0; 12 AIO containers Up; web reachable pct rollback Before cloud host step + AIO app upgrade
1 media VM105 arr (43/9) Guest security apt; TS+Docker ride along qm snapshot 105 pre-os-sec apt; security pocket; needrestart sec=0; 8 containers Up; Samba + all UIs load qm rollback Before VM105 Phase 2 apps
1 media CT111 mcc (51/13, local caddy+postfix) Guest security apt pct snapshot 111 pre-os-sec pct exec; security pocket sec=0; caddy+postfix active pct rollback Independent (no Phase 2 app)
1 utility CT100/101/102/103/106/107/112/118 (OS layer) Security-pocket apt per CT; lockout-critical (Caddy CT101, headscale CT106) carved to LIST mode + LAST; CT118 leave rpcbind alone pct snapshot each + headscale DB backup for CT106 pct exec; security pocket; needrestart -r l for control-plane CTs sec=0 each; ingress spot-check; headscale nodes list joined pct rollback Phase 0; control-plane CTs last as singletons
1 edge2 CT100-105 (OS layer, via pct exec from edge2 host) Security-pocket apt per CT; DB dumps for CT101 WP MariaDB, CT104 CouchDB, CT105/CT106 PG pct snapshot each pre-phase1 + DB dumps From root@184.174.35.153 pct exec; security pocket; defer reboot sec=0; each app reachable via Caddy front door pct rollback; restore DB dump Phase 0; edge2 host fully patched (host = no-op)
1 edge2 CT106 matrix / CT107 headscale (NO Tailscale client) Security apt driven via pct exec from edge2 HOST only; singletons pct snapshot + Synapse PG dump (106) + headscale DB backup (107) From edge2 host shell; needrestart -r l; do NOT restart headscaled/synapse unless required sec=0; matrix federation+login; headscale nodes list all joined; Caddy ingress works pct rollback; restore DB OOB path confirmed first; NOT in edge2 batch
1 cortex VM150 (PROTECTED, 90 upgradable) Security apt; manual in toc+cortex window; TS+Docker ride along qm snapshot 150 pre-sec-apt (from toc) apt --only-upgrade security pkgs; needrestart -r a; defer reboot sec=0; 5 AI containers Up; nvidia-smi ok qm rollback Phase 0; protected — not bulk
1 pi-nas host OS (subset of 130) Debian security pocket; HOLD kernel 6.12→6.18 + OMV pkgs Off-box config.xml + dpkg --get-selections baseline apt-mark hold linux-image-* openmediavault*; -t trixie-security upgrade; unhold; needrestart sec=0; shares mount; OMV UI loads; no reboot now Reinstall prior pkg from cache; restore config.xml Phase 0; precedes pi-nas kernel reboot
1 Cross-cutting Tailscale 1.94→1.98 (all TS-client nodes) Bump client; rides along Phase 1; EXCLUDE edge2 CT106/CT107 Covered by per-target Phase 1 snapshot apt install --only-upgrade tailscale; daemon restart tailscale version=1.98; node still joined Reinstall 1.94; snapshot rollback Folds into Phase 1; LAN/console fallback confirmed
1 Cross-cutting Docker CE→29.6 (Docker hosts) Bump engine; rides along Phase 1/2 Covered by per-target snapshot apt install --only-upgrade docker-ce docker-ce-cli containerd.io; daemon restart docker version=29.6; containers return Downgrade pkg; snapshot rollback Folds into Phase 1; before compose pulls
2 edge2 CT105 Authentik 2025.12.4→2025.12.6 (hop 1) SSO hop #1 (FIRST load-bearing app); pre-flight: dedup group names, /media/data/media if local storage pct snapshot 105 + pg_dump (per-hop) Edit tag; compose pull && up -d; run migrations; NO skipping UI=12.6; migrations clean; dependent login works compose down; restore tag+dump; or pct rollback Phase 0 break-glass proven; window
2 edge2 CT105 Authentik 2025.12.6→2026.2.x (hop 2) SSO hop #2 cross-major pct snapshot + fresh pg_dump tag→2026.2.x; pull/up; migrations UI=2026.2.x; logins re-verified restore prior tag+dump after hop 1
2 edge2 CT105 Authentik 2026.2.x→2026.5.3 (hop 3, final) SSO hop #3; verify CVE/GHSA IDs; insert any mandatory intermediate pct snapshot + fresh pg_dump tag→2026.5.3; pull/up; final migrations UI=2026.5.3; full sweep of dependent logins; close window restore prior tag+dump after hop 2; verify dependents BEFORE upgrading them
2 utility CT101 + media CT111 Caddy (ingress, EARLY) Caddy bump (before home-app verification); map Caddy↔Authentik path; document direct Authentik admin URL pct snapshot + Caddyfile/certs backup caddy validate; restart in low-traffic window home services reachable; certs valid rollback snapshot/Caddyfile After CT101 Phase 1; separate window from Authentik
2 edge2 CT103 Forgejo 14.0.5→15.0.3 (branch migration) Major migration; pre: off-edge2 git clone --mirror echo6-docs + pause autocommit cron + forgejo doctor check --all --fix pct snapshot 103 + repos volume + pg_dump + app.ini doctor check→repair→backup→bump tag 14→15; pull/up; migrations reports 15.0.3; clone/push works; CI/webhooks ok restore 14.0.5 image + volume + dump After CT103 Phase 1; NOT same window as Authentik
2 utility CT109 OTS app OTS updater (app+webUI+schema) — does NOT migrate RabbitMQ pct snapshot 109 + PG dump Run OTS upgrade script OTS web/API up; CoT works; clients reconnect rollback snapshot + dump After CT109 Phase 1 OS
2 utility CT109 RabbitMQ 3.12.1→4.2.8 (SEPARATE) Decoupled EOL major: Erlang≥26→3.13.x→enable_feature_flag all→4.x pct snapshot + rabbitmqctl export_definitions staged hops; confirm flags enabled; no classic-mirroring config 4.2.x; queues intact; TAK clients reconnect restore snapshot + definitions GATES MediaMTX/Mumble accept; after OTS app
2 utility CT106 + edge2 CT107 Headscale 0.28→0.29 Lower-stakes instance FIRST (rehearsal), front-door LAST; disable key-expiry; OOB-driven pct snapshot + headscale DB dump off-tailnet stop→backup→replace binary→config migrate→start; headplane to match version=0.29.1; nodes list all joined; tunnels intact restore snapshot+DB via console one-at-a-time; non-tailnet path; not same window as Caddy/Authentik/Forgejo
2 cloud CT120 Immich 2.5.6→2.7.5 (+Valkey 9.0.2→9.1.0, vectorchord PG) Stack upgrade; prune images first; do NOT interrupt first-boot migration; disable bg regen pct snapshot 120 pre-immich-2.7.5 + pg_dump compose pull && up -d; let migrations run; no skip across breaking migration web+login; jobs process; valkey 9.1.0 ping; version 2.7.5 restore tag (only with dump) or pct rollback After CT120 Phase 1; free space confirmed
2 cloud CT121 Nextcloud AIO (master ~12.5→13.2.1, NC 32.0.4→32.0.11) AIO orchestrated update; borg backup as rollback artifact AIO borg backup + pct snapshot 121 AIO backup→stop→update mastercontainer→UI :8080 "Update containers" 12 containers green; occ status 32.0.11; integrity clean restore from AIO borg; else pct rollback After CT121 Phase 1; AIO self-gates intermediate versions
2 media CT110 PeerTube 8.0.2→8.2.1 (native) Native upgrade script; quiesce jobs pct snapshot 110 + pg_dump peertube_prod back up DB → run PeerTube upgrade.sh/version steps → restart reports 8.2.1; video plays; PG16/redis/nginx healthy pct rollback + restore dump After CT110 Phase 1
2 media VM105 jellyfin 10.11.6→10.11.11 Patch bump qm snapshot 105 (pre-app batch) compose pull jellyfin && up -d (pin tag) UI=10.11.11; plays/transcodes redeploy prior digest; qm rollback After VM105 Phase 1
2 media VM105 SABnzbd 4.5.5→5.0.4 (MAJOR) Major; pause queue; audit post-proc scripts qm snapshot 105 pre-sabnzbd + sabnzbd.ini pin 5.0.4; pull/up; migrate on start UI=5.0.4; queue intact; test NZB; arr integrations ok redeploy 4.5.5 + restore ini (queue repair); qm rollback After VM105 Phase 1; before Lidarr
2 media VM105 Lidarr 2.x→3.x (MAJOR) Major branch migration qm snapshot 105 pre-lidarr + lidarr.db pin v3; pull/up; DB migration on start UI=v3; library+indexers+download client ok restore db (migration one-way); qm rollback After SABnzbd (verify download-client link)
2 media VM105 Sonarr/Radarr/Prowlarr (1-2 ver) Routine bumps pre-app VM snapshot + config DBs compose pull && up -d UIs load; prowlarr sync; test grab redeploy prior tags + DBs After SABnzbd/Lidarr
2 media VM105 Navidrome 0.60.3→0.62.0 Routine bump pre-app snapshot + DB compose pull navidrome && up -d UI=0.62.0; library scans; track streams redeploy 0.60.3 + DB Independent of arr chain
2 media VM105 jellyseerr preview-OIDC→stable 3.3.0 DECISION GATE: schema downgrade + correct artifact (jellyseerr vs Seerr) + OIDC parity qm snapshot 105 pre-jellyseerr + config dir test stable vs DB COPY first; swap tag only if compatible UI=3.3.0; requests intact; OIDC/login works restore preview tag + config; qm rollback After jellyfin/arr; STOP+report if OIDC parity unclear
2 edge2 CT104 livesync CouchDB 3.4→3.5.2 Lower-urgency image bump pct snapshot 104 + CouchDB volume backup bump tag; compose pull && up -d reports 3.5.2; _up healthy; LiveSync replicates test edit restore 3.4 + volume; pct rollback After CT104 Phase 1; triple-sync verify
2 utility CT104 NATS 2.14.0→2.14.2 Lower-urgency patch pct snapshot 104 (JetStream dir in snap) replace binary; restart 2.14.2; JetStream/MQTT intact pct rollback After CT104 Phase 1 (PG already done); stateful blip
2 utility CT100 meshmonitor / CT102 searxng+valkey-8 (8.1.5→8.1.8) Lower-urgency image refresh pct snapshot + record digests compose pull && up -d containers Up; function ok redeploy prior digest; pct rollback After respective Phase 1
2 edge2 CT100 PDM / CT101 WP core / CT102 Vaultwarden / CT106 Synapse Verify-current / no-op (already at/past fix); WP plugin/theme = separate WP-CLI follow-up None (Phase 1 snaps) optional digest re-pull only versions confirmed; logged as current redeploy prior digest Closes as no-op
2 cortex Ollama 0.16.1→0.30.10 (PROTECTED) Image bump; no downgrade after new pull qm snapshot 150 pre-ollama + models volume compose pull ollama && up -d; test inference /api/version=0.30.10; models intact; GPU used; vault-tagger works re-pin 0.16.1; restore snapshot if model format broke After cortex Phase 1 + Docker + toolkit
2 cortex Open-WebUI 0.8.1→0.9.6 Image bump (auth surface) qm snapshot 150 pre-openwebui bump tag; pull/up UI loads; login works; reaches Ollama re-pin 0.8.1; snapshot After Ollama
2 cortex Qdrant 1.16.3→1.18.2 / TEI 1.7.4→1.9.3 Lower-urgency image bumps qm snapshot 150 per app bump tag; pull/up health endpoints ok; collections/embeddings intact; docs engine works re-pin prior; snapshot After cortex Docker+toolkit; verify engines
2 cortex obsidian-remote v1.12.7 Verify-current / no-op None none container up; UI reachable N/A
2 utility CT118 archivist rpcbind 0.0.0.0:111 Remediate exposure (network change — APPROVAL) pct snapshot 118 per approved option: disable rpcbind / bind localhost+TS / firewall port 111 not on 0.0.0.0; app functions; external scan closed re-enable / pct rollback After CT118 Phase 1; approval-gated (Open Dec #6)
2 edge2 CT108 mailcow REMOVED — CT108 destroyed 2026-06-20 (not retained) N/A N/A N/A N/A N/A
3 Per-node corosync/HA pre-flight (data/utility/cloud/media/toc) Before EACH reboot None pvecm status (Quorate:Yes, 5 votes); ha-manager status quorate + expected votes=5; HA implications known abort if not quorate Gates each Phase 3 node reboot
3 Canary node (lowest blast + headroom; media, NOT data until disk freed) Full 9.2/QEMU11/LXC7/kernel reboot to prove the path vzdump all guests to EXTERNAL target + record pveversion -v apt update && apt dist-upgrade → reboot; NO live-migration in mixed window pveversion=9.2/QEMU11/LXC7/kernel 6.17.13; all guests onboot return; quorate boot prior kernel (GRUB); restore guests from vzdump Phase 1+2 done; canary before high-stakes nodes
3 media host PVE 9.1.1→9.2 (+QEMU11/LXC7/kernel) Platform + reboot vzdump VM105/CT110/CT111 + snapshots dist-upgrade → reboot guests return; arr/peertube/caddy healthy; quorate GRUB prior kernel; vzdump restore After media Phase 1+2; one node at a time
3 cloud host PVE 9.2 platform + reboot Platform + reboot vzdump CT120/CT121 + record versions dist-upgrade → reboot immich+nextcloud return healthy; quorate GRUB prior kernel; vzdump restore After cloud Phase 1+2; after canary proven
3 utility host PVE 9.2 platform + reboot Platform + reboot (known full home-ingress + mesh + central outage) vzdump all 12 guests + versions dist-upgrade → reboot all 12 return; Caddy ingress + headscale + central back; quorate GRUB prior kernel; vzdump restore LAST of early nodes; low-traffic window
3 data host PVE 9.2 platform + reboot Platform + reboot vzdump VM1130 + versions dist-upgrade → reboot VM1130 returns; NFS/Samba serve; quorate GRUB prior kernel; vzdump restore After disk remediation; one at a time
3 Hypervisor host OS-security (folded in) Apply held host libc/openssl WITH the platform full-upgrade per node covered by per-node vzdump included in dist-upgrade (not applied back in Phase 1) sec=0 host-side post-reboot; daemons up per-node rollback Fold-in decision (Open Dec #2)
3 Post-reboot app re-verification (per node) Re-run Phase-2 health checks for majored guests on each rebooted node None health probes CT120/CT121/CT110/VM105/CT109 healthy on new substrate re-snapshot/redeploy degraded component After each node reboot
3 cortex nvidia-container-toolkit 1.18→1.19 Toolkit bump (pair w/ driver) qm snapshot 150 pre-nvtoolkit apt --only-upgrade nvidia-container-toolkit; nvidia-ctk runtime configure; restart docker --version=1.19; docker run --gpus all nvidia-smi ok downgrade 1.18; snapshot toc+cortex window
3 cortex NVIDIA driver 580.159→580.167 + DKMS (reboot) Driver + DKMS as discrete reversible step BEFORE toc reboot qm snapshot 150 pre-nvidia + vzdump apt --only-upgrade nvidia-driver-580 nvidia-dkms-580; DKMS rebuild; verify BEFORE toc reboot nvidia-smi=580.167; dkms status installed; containers see GPU restore snapshot; reinstall 580.159 + DKMS Keep old driver pkg; after toolkit
3 toc host PVE 9.2 + QEMU11 + LXC7 + kernel (reboot — LAST, drops cortex) Platform + reboot; only guest is VM150 vzdump/snapshot VM150 + record pveversion snapshot VM150 → dist-upgrade toc → reboot → VM150 returns → cortex driver reboot pveversion=9.2; toc rejoins (5/5 votes); VM150 + GPU healthy restore VM150 from vzdump; GRUB prior kernel LAST + alone; alternate control point TESTED; all 4 other nodes quorate
3 cortex final stack health verification Post-window full AI-stack + engine check keep pre-window snaps until verified verification only driver 580.167; all 5 containers; ollama/tei/qdrant/open-webui; vault-tagger+TEI engines ok; TS online; Docker 29.6 restore degraded component; worst case VM150 vzdump LAST action of node group
3 pi-nas OMV 8.1→8.4 (reboot #1, separate window) DONE 2026-06-21apt full-upgrade (kernel/firmware 6 pkgs held); OMV 8.4.0-3 installed; omv-salt deploy ran (14/0 OK); initramfs regenerated for 6.12.62 only; rebooted; uname -r still 6.12.62; all 5 NFS exports serving; immich-nfs-OK, nextcloud-nfs-OK; nfs-kernel-server + docker active; degraded = pre-existing quotaon "File exists" false alarm, unrelated to upgrade N/A N/A N/A Complete
3 pi-nas kernel 6.12→6.18 (reboot #2, separate window) Kernel bump → reboot; activates deferred security back up /boot+/boot/firmware; keep 6.12 installed; fresh config.xml apt install linux-image-arm64 (unhold); update bootloader; reboot ON APPROVAL uname -r=6.18; LAN+tailnet up; shares mount; Docker up boot retained 6.12; restore /boot; on-site SD reflash After OMV reboot verified; physical access
3 edge2 host PVE 8.4 reboot — CONDITIONAL (CT107 boot-survival blocker CLEARED 2026-06-20) Previously blocked on CT107 headscale boot-survival — cleared (self-heals in ~45s). Now gated only on Phase 0 kernel-CVE decision (Open Dec #5 = no reboot needed — host fully patched). If ever required: own no-failover window AFTER CT105/CT107 verified vzdump all guests + pveversion scoped kernel/security upgrade → reboot in approved window uname -r=fixed; all CTs return; Caddy-fronted services reachable GRUB prior kernel; vzdump restore GATED on Phase 0; announce SSO+tailnet downtime; non-tailnet path
Nominatim 4.5→5.3 + Photon 1.1→1.2 (SEPARATE PROJECT) Full re-import into separate DB/instance; keep 4.5 until validated; Photon AFTER nominatim validated qm snapshot 1130 + retain 4.5 image/data import to new DB; size flat-nodes vs free space (NOT on data); cut over after verify 5.x geocode correct; Photon 1.2 reads v5 re-pin nominatim:4.5 + old data Wholly separate from patch campaign; after disk remediation

Key Facts & Top Risks

Cluster & control-plane facts

  • One corosync cluster. data, utility, cloud, media, toc are echo6-cluster — quorum 3 of 5. Every Phase 3 reboot is gated on pvecm status (Quorate:Yes, expected votes=5); ONE node at a time; confirm 5/5 rejoin before next.
  • toc coupling. toc is the 5th cluster vote + GPU host for cortex VM150 + the management/Claude Code host. toc+cortex are the FINAL Phase 3 action, alone — never overlap another node reboot.
  • Headscale topology (resolved). Main fleet tailnet (34 nodes, ControlURL vpn.echo6.co) = edge2 CT107 (Headscale 0.28.0) — HIGH lockout risk; upgrade LAST, drive from out-of-band. Separate IdahoMesh sub-tailnet (vpn.idahomesh.com, 3 nodes) = utility CT106 — low risk; upgrade FIRST as rehearsal. No routing through old-Contabo.

Top risks (critical first)

  1. Headscale edge2 CT107 self-lockout (CRITICAL) — upgrading the main tailnet over the tailnet itself loses the control path to 34 nodes. Mitigation: out-of-band only (edge2 host console / direct SSH 184.174.35.153), key-expiry extended, DB dumped off-tailnet, boot chain verified.
  2. Corosync quorum loss (CRITICAL) — two nodes down simultaneously = quorum lost, /etc/pve read-only fleet-wide. Mitigation: one node at a time, pvecm status gated before each, HA disabled during window.
  3. Authentik SSO migration chain (CRITICAL) — 4-hop irreversible Django migration sequence; a failed hop with no per-hop pg_dump can strand the entire SSO surface. Mitigation: pg_dump per hop, verified-login gate, break-glass proven before starting.
  4. data node snapshot-space exhaustion (CRITICAL) — 92% full; a snapshot failure mid-upgrade leaves the guest in a half-upgraded, unrollbackable state. Mitigation: disk remediated to <85% before any snapshot on that node; do not use as Phase 3 canary until resolved.
  5. RabbitMQ EOL/won't-boot (HIGH) — 3.12→4.x direct jump refuses to start; 3.x has no security backports. Mitigation: staged hops (3.12→3.13.x→feature-flags→4.x), decoupled from OTS updater.
  6. Forward-only DB migrations / false-rollback assumption (HIGH) — "redeploy previous image tag" is NOT a valid rollback for authentik, forgejo, nextcloud, peertube, immich after a migration runs. Rollback = quiesced logical dump + old binary together.
  7. toc reboot drops cortex + management host (HIGH) — cortex is the Claude Code host; losing it during diagnosis is double-jeopardy. Mitigation: alternate control point tested before window; cortex NVIDIA/DKMS verified before toc reboot; toc+cortex done last and alone.
  8. pi-nas non-boot after combined kernel+OMV change (HIGH) — RPi has no remote KVM; a bad combined update requires on-site SD reflash. Mitigation: two separate reboots (OMV first, kernel second); retain 6.12; /boot+/boot/firmware backed up; physical access confirmed before each.

Full Inventory

Complete point-in-time state of every node, guest, and container service.

data (PVE 9.1.1)

  • Host: 126 upgradable / 51 security; no Docker installed; roles: NAS, NFS, Samba
  • VM1130 recon-vm (Ubuntu 24.04) — 34 upgradable / 10 security
    • PostgreSQL 16.14 (DBs: overture, padus), photon, kiwix, recon.py, 7x navi-backend, nginx, Apache, Samba
    • Docker: valhalla:latest, nominatim:4.5 (stale 14 months), zimit:latest (not running)

utility (PVE 9.1.1)

  • Host: 148 upgradable / 38 security; PVE 9.2 platform update pending; 12 LXC guests
CT Name Upgradable / Sec Services
CT100 meshmonitor 62 / 15 ghcr.io/yeraze/meshmonitor:latest
CT101 caddy (home ingress) 59 / 18
CT102 searxng 54 / 15 searxng/searxng:latest + valkey/valkey:8-alpine
CT103 argus 29 / 26 RF capture / transcribe / viewer
CT104 central 50 / 38 PostgreSQL 16 + NATS/MQTT/JetStream
CT106 meshtastic-hs 54 / 17 headscale control plane
CT107 mesh-bridge 53 / 15 dual tailscaled
CT108 meshai 109 / 79 work-meshai local build
CT109 opentakserver 34 / 31 nginx / PG16 / rabbitmq / mumble / mediamtx / CoT
CT112 cobalt 35 / 29 build/CI (idle)
CT118 archivist 62 / 17 archivist + rpcbind (FLAG: port 111 on 0.0.0.0)
CT119 mesh-territory 179 / 91 meshwars:latest (never patched)

cloud (PVE 9.1.1)

  • Host: 113 upgradable / 38 security; 2 LXC guests
CT Name Upgradable / Sec Services
CT120 immich 191 / 101 immich_server, immich_machine_learning, valkey/valkey:9, immich postgres (14-vectorchord) — all drifted
CT121 nextcloud AIO 98 / 75 12 containers: mastercontainer + apache + nextcloud + postgresql + redis + collabora + clamav + imaginary + fulltextsearch + notify-push + whiteboard + docker-socket-proxy; NC 32.0.4; mastercontainer behind

media (PVE 9.1.1)

  • Host: 107 upgradable / 36 security; 3 guests
Guest Name Upgradable / Sec Services
VM105 arr 43 / 9 (Ubuntu 24.04) jellyfin / sonarr / radarr / prowlarr / sabnzbd / lidarr / navidrome / jellyseerr (preview-OIDC) all :latest + Samba
CT110 peertube 81 / 37 v8.0.2; nginx / PG16 / redis / peertube / pt-downloader + importer + monitor / nordvpn
CT111 mcc 51 / 13 caddy + postfix

toc (PVE 9.1.1) — PROTECTED

  • Host: 189 upgradable / 43 security; PVE 9.2 platform update pending; reboot required
  • Hosts only VM150 cortex — coordinate any toc work with cortex maintenance window

cortex (VM150, Ubuntu 24.04) — PROTECTED GPU / Claude Code host

  • 90 apt upgradable
  • NVIDIA driver 580.159 → 580.167 + DKMS (reboot required)
  • nvidia-container-toolkit 1.18 → 1.19
  • Docker: ollama / tei 1.7 / qdrant / open-webui / obsidian — all drifted

pi-nas (Debian 13, arm64, RPi + OMV)

  • 130 apt upgradable; kernel 6.12 → 6.18 (reboot required); OMV 8.1 → 8.4
  • Docker engine installed; 0 containers running

edge2 (PVE 8.4.19) — Host Fully Patched

  • 8 LXC guests (CT108 mailcow decommissioned 2026-06-20)
CT Name Services / Status
CT100 pdm PDM 1.1.4, current (native)
CT101 wordpress Apache 2.4.67 / PHP 8.4 / MariaDB 11.8.6 / WP core 7.0 — plugin/theme status needs WP-CLI (superseded 2026-07-17 — migrated to Grav CMS, MariaDB removed; see topology note above)
CT102 vaultwarden vaultwarden/server:latest — drift indeterminate
CT103 forgejo 14.0.5 (forgejo:14 + postgres:16-alpine drifted)
CT104 livesync couchdb:3.4 drifted + local provisioner
CT105 authentik 2025.12.4 (server/worker/postgres) → upgrade to 2026.2.0
CT106 matrix Synapse 1.155.0 / Element / MAS + mautrix-signal + postgres; OS apt security updates pending; no Tailscale client
CT107 headscale 0.28.0 → 0.29.0 + headplane; no Tailscale client
CT108 mailcow Decommissioned 2026-06-20 — destroyed (pct destroy 108 --purge); superseded by edge1, backup on pi-nas

Coverage Notes

  • Complete: all five PVE-9 nodes (data, utility, cloud, media, toc), edge2, cortex, pi-nas, and all discovered guests/containers.
  • Excluded: edge1 (the rebuilt ex-Contabo, mail-only) — mid-rebuild/maintenance at time of audit; re-audit once it settles.
  • Incomplete: WordPress (CT101) plugin and theme status — requires WP-CLI; not assessed. (Moot as of 2026-07-17 — CT101 migrated to Grav CMS, no plugins/DB; see topology note above.)
  • Method: read-only throughout — SSH / pct exec, apt list --upgradable, docker manifest inspect for same-tag drift. Floating/pinned-tag caveats noted inline where drift could not be confirmed.