echo6-docs/vault/projects/fleet-patch-audit.md
echo6-autocommit 66a214ed34 auto: docs sync 2026-06-19T18:00:06+00:00
Files changed: .claude/settings.local.json engine/lint-report.md vault/.obsidian/workspace.json vault/projects/fleet-patch-audit.md
2026-06-19 18:00:06 +00:00

63 KiB
Raw Blame History

title type tags related updated status
Fleet Patch Audit — 2026-06-19 project
proxmox
ai
2026-06-19 active

Fleet Patch Audit — 2026-06-19

Read-only audit snapshot as of 2026-06-19. Nothing has been applied — this is a planning document to build the patch plan from.

Topology note: the old Contabo VPS has been rebuilt as edge1 (mail-only); edge2 is now the front door for everything else. edge1 is excluded from this audit (mid-rebuild/maintenance). The CLAUDE.md cheat-sheet still lists the old Contabo layout and is stale — refreshing it is a follow-up task (see Open Decisions).


Prioritized Backlog

Tier 1 — Security-Urgent Guest OS

These containers have the highest raw security-update counts and have not been patched recently (or never). Address before any platform work.

Host Guest Upgradable / Security Notes
utility CT119 mesh-territory 179 / 91 sec Never patched
utility CT108 meshai 109 / 79
cloud CT120 immich guest-OS 191 / 101
cloud CT121 nextcloud guest-OS 98 / 75
media CT110 peertube 81 / 37
utility CT109 opentakserver 34 / 31
utility CT104 central 50 / 38 Includes PostgreSQL 16.13 → 16.14

Tier 2 — App / Container Updates

Updates where the application or its Docker images have drifted from current upstream, ordered roughly by operational risk.

Scope Guest Item Notes
edge2 CT105 authentik 2025.12.4 → 2026.5.3 #1 security item — 7 CVEs + 5 GHSAs in gap; sequential upgrade (min: 2025.12.6)
edge2 CT107 headscale 0.28.0 → 0.29.1 Also a 2nd headscale on utility CT106
edge2 CT103 forgejo 14.0.5 → 15.0.3 14.x EOL 2026-04-30 — migrate branch, not just patch
edge2 CT106 Synapse 1.155.0 / Element / MAS Image drift + pending OS apt security updates
edge2 CT104 livesync couchdb:3.4 Docker image drift
edge2 CT108 mailcow (18 containers) Upgrade via update.sh only
cloud CT120 immich — server/ml/valkey:9/postgres(14-vectorchord) 4 images drifted
cloud CT121 nextcloud AIO — mastercontainer + NC app 32.0.4 Mastercontainer behind; 12-container stack
cortex VM150 ollama / tei(1.7) / qdrant / open-webui / obsidian 5 AI containers drifted
media VM105 arr stack — 8 containers jellyfin/sonarr/radarr/prowlarr/sabnzbd/lidarr/navidrome/jellyseerr all :latest

Tier 3 — Platform / Reboot Windows

Reboot-bearing updates. Coordinate maintenance windows carefully. cortex and toc are PROTECTED hosts.

Scope Item Detail
data, utility, cloud, media, toc (PVE 9 nodes) PVE 9.1.1 → 9.2 Reboot required
data, utility, cloud, media, toc QEMU 10 → 11 Reboot required
data, utility, cloud, media, toc LXC 6 → 7 Reboot required
data, utility, cloud, media, toc Kernel 6.17.2 → 6.17.13 Reboot required
edge2 PVE 8.4.19 Already fully patched — no action needed
cortex VM150 (PROTECTED) NVIDIA driver 580.159 → 580.167 + DKMS Reboot required
cortex VM150 (PROTECTED) nvidia-container-toolkit 1.18 → 1.19
pi-nas Kernel 6.12 → 6.18 Reboot required
pi-nas OMV 8.1 → 8.4

Cross-Cutting (All / Most Hosts)

Item Detail
Tailscale 1.94 → 1.98 nearly everywhere
Docker CE → 29.6 wherever Docker is installed

Non-Update Flags

Issues noted that are not package/image updates but warrant attention.

Host / Guest Flag Detail
data Disk 92% full ~73 GB / 938 GB free; address before patching
utility CT118 archivist rpcbind on 0.0.0.0:111 No Tailscale client or firewall on this CT; exposed port
media VM105 jellyseerr Non-stable image Running preview-OIDC tag, not a stable release
data VM1130 nominatim Stale image (14 months) nominatim:4.5, pinned; confirm intentional
edge2 CT106 matrix / CT107 headscale No Tailscale client Ingress via Caddy; verify internal routing before patching

Application Currency (2026-06-19)

Running application version vs latest stable upstream, per app — the "is everything up to date?" view (distinct from OS-package/image-drift above). Read-only. Caveat: version deltas were measured directly from the running apps and are reliable; the specific CVE/advisory IDs below were surfaced by the audit agents and are post-knowledge-cutoff — verify against primary advisories before acting on them. GitHub's anonymous API rate-limited a few "latest" lookups (noted inline).

Security-relevant — prioritize

App Where Running → Latest Behind Security note (verify)
Authentik edge2 CT105 2025.12.4 → 2026.5.3 ~6 mo / 5 majors #1 — claimed 7 CVEs + 5 GHSAs across two May-2026 waves. Min-disruption: 2025.12.6 (same branch, backported fixes); full: 2026.5.3. SSO — upgrade sequentially.
Immich cloud CT120 2.5.6 → 2.7.5 9 rel shared-link ACL bypass (2.6.0) + stored XSS panorama viewer (2.7.0)
Valkey cloud CT120 9.0.2 → 9.1.0 3 patch 6 CVEs, "SECURITY" urgency (use-after-free, DoS, RESP injection)
RabbitMQ utility CT109 3.12.1 → 4.2.8 EOL major 3.x abandoned upstream; 4 advisories 2026-06-18, no 3.x backports
Nominatim data/recon-vm 4.5.0 → 5.3.2 ~19 mo, major v4→v5 data-model break; needs image swap + full re-import (pair w/ Photon 1.2.0)
Ollama cortex 0.16.1 → 0.30.10 14 minor history of SSRF / path-traversal CVEs
Nextcloud cloud CT121 32.0.4 → 32.0.11; AIO ~v12.5 → 13.2.1 7 patch / ~8 AIO resource-exhaustion fix (32.0.10); AIO image ~5 mo stale
Forgejo edge2 CT103 14.0.5 → 15.0.3 EOL branch 14.x EOL 2026-04-30; no further backports — migrate to 15.0.3 (LTS)
SABnzbd media VM105 4.5.5 → 5.0.4 major 5.x security hardening over 4.x
Jellyfin media VM105 10.11.6 → 10.11.11 5 patch 10.11.x patch cycle has carried CVE fixes
Jellyseerr media VM105 preview-OIDC dev → 3.3.0 unreleased running unpinned dev code — no stable security posture
PeerTube media CT110 8.0.2 → 8.2.1 2 patch 8.x patch releases include security fixes
Open-WebUI cortex 0.8.1 → 0.9.6 ~1.5 minor touches auth/session surface
MediaMTX utility CT109 1.13.0 → 1.19.1 6 minor RTSP/WebRTC streaming attack surface
Headscale ×2 utility CT106 + edge2 CT107 0.28.0 → 0.29.1 1 minor upgrade-guide-required; two separate instances; no security flag noted

Operational / host-level (not a simple app bump)

Item Where Finding
Mailcow edge2 CT108 STOPPED — very likely superseded by the new edge1 mail-only node. Confirm and decommission rather than update.
Host kernel edge2 host Agent flagged DirtyFrag (CVE-2026-43284/-43500) + copy.fail (CVE-2026-31431, claimed CISA KEV) as host-kernel LPE. Tension: the host audit showed edge2 fully patched (0 upgradable) on its repo — verify whether these need a kernel newer than the no-subscription repo provides.

Current / already past the fix (no action)

Vaultwarden 1.36.0 (edge2 CT102 — has the SSO-takeover/org-access CVE fixes) · PDM 1.1.4 (edge2 CT100 — past the RCE PSA) · WordPress 7.0 core + all plugins/themes (edge2 CT101) · Synapse 1.155.0 / Element / MAS (edge2 CT106 — current, only minor :latest digest drift) · obsidian-remote v1.12.7 (cortex) · PostgreSQL 16.14 (recon-vm).

Lower urgency

Mumble 1.5.517→1.5.901 · Caddy 2.10.2/2.11.3→2.11.4 · Qdrant 1.16.3→1.18.2 · TEI 1.7.4→1.9.3 · Valhalla 3.6.3→3.7.0 · Photon 1.1.0→1.2.0 · kiwix 3.7.0→3.8.2 · CouchDB 3.4.3→3.5.2 (livesync) · Navidrome 0.60.3→0.62.0 · Sonarr/Radarr/Prowlarr/Lidarr 12 versions · NATS 2.14.0→2.14.2 · PostgreSQL 16.12/16.13→16.14 · meshmonitor (~1 mo, exact ver undeterminable) · searxng (rolling, ~4.5 mo) + valkey-8 sidecar 8.1.5→8.1.8 · mautrix-signal v0.2603.0.

Internal echo6 apps (no upstream to track)

central-, meshai, archivist, meshwars, recon / recon-watchdog, navi- — running; version = current git head.


Proposed Patch Approach (NOT executed — for planning)

Lowest-risk changes first; everything reboot-bearing deferred to scheduled windows.

Principles

  • Security-pocket apt only in the first pass (openssl, openssh, gnutls, krb5, libc6, samba, nghttp2, etc.). No PVE 9.2 / QEMU / LXC / kernel, no Docker image pulls, no OMV/NVIDIA, no reboots.
  • Protected hosts (cortex, toc) never go in a bulk pass — handled individually in their own window. Note: toc hosts cortex (VM150), so a toc reboot drops cortex — the two must be coordinated together.
  • edge1 (mail) excluded while it's mid-rebuild.
  • needrestart will bounce affected daemons after glibc/openssl upgrades — seconds of blip per guest, no data risk.

Phased plan

Phase Scope Reboot? Notes
1 — Guest/VM security apt utility CT100,101,102,103,104,106,107,108,109,112,118,119 · cloud CT120,121 · media VM105,CT110,CT111 · data VM1130 · edge2 CT100108 No Lowest blast radius. Worst-first: CT119, CT108, immich/nextcloud guest-OS, peertube.
2 — Hypervisor host OS security data, utility, cloud, media host OSes (not toc) No One node at a time; restarts hypervisor-side daemons (smbd etc.). edge2 host already patched.
3 — App / container updates (Tier 2) per-app, native updater each Per-app See "special handling" below — not a generic docker pull.
4 — Reboot windows (Tier 3) PVE 9.2 + QEMU 11 + LXC 7 + kernel on the 5 PVE-9 nodes; pi-nas kernel + OMV; cortex NVIDIA/DKMS Yes Schedule deliberately; toc+cortex coordinated.
Cross-cutting Tailscale 1.94→1.98 fleet-wide No Can ride along Phase 1/2.

Special handling — do NOT bulk-patch these; use the native updater

  • mailcow (edge2 CT108) → ./update.sh (upgrades all 18 components in lockstep)
  • nextcloud AIO (cloud CT121) → update mastercontainer, then in-UI update button (port 8080)
  • immich (cloud CT120) → docker compose pull && up -d from its compose dir
  • authentik (edge2 CT105) → sequential version upgrades with migrations; 2025.12.4 → 2026.2.0 cannot skip releases
  • pi-nas OMV → OMV's own update path, not raw apt
  • cortex / toc → manual, protected, own window

Open Decisions (for tomorrow's plan)

  1. Phase 1 scope — all non-protected guests at once, or staged worst-first?
  2. Phase 2 — patch the four hypervisor host OSes now (no reboot), or fold into the Phase 4 reboot window?
  3. Reboot-window scheduling — order of the 5 PVE-9 nodes; toc + cortex must be done together (toc reboot drops cortex). edge2 needs none.
  4. Tier-2 app upgrades — which to take on: authentik major (2025→2026, migration-heavy), Matrix/Synapse, Forgejo, headscale 0.29, mailcow. Each is its own task.
  5. data disk at 92% — remediate before/independently of patching (operational risk regardless).
  6. CT118 archivist rpcbind on 0.0.0.0:111 with no Tailscale/firewall — treat as a separate exposure fix.
  7. Refresh the stale CLAUDE.md cheat-sheet to the edge1/edge2 topology — separate doc task.
  8. Daemon-restart tolerance — confirm brief blips are acceptable for the stateful services (central PG16/NATS, opentakserver, peertube) during Phase 1/2.
  9. edge2 CT108 mailcow is stopped — confirm it's superseded by the edge1 mail node and decommission it, vs. it being an unintended outage.
  10. edge2 host-kernel CVEs (DirtyFrag / copy.fail) — verify whether the flagged in-the-wild LPEs actually apply given the host shows fully patched; if real, this is a host-kernel + reboot action on edge2 (which otherwise needs none).
  11. App-currency CVE IDs are post-cutoff — verify the specific advisories (Authentik waves, Valkey, Immich, RabbitMQ) against primary sources before using them to justify urgency.

Execution Runbook (Meticulous)

Step-by-step plan to bring every application and package current. Per-app target versions live in Application Currency above; this section is the how/when/order. Execution model: scoped Sonnet task per target → Opus verifies output → report to Matt; approval gate before every load-bearing or reboot step.

Standing guardrails (apply to EVERY step)

  • Protected hosts — cortex & toc — never in a bulk pass. They get a dedicated, manual window. Note the coupling: toc reboot drops cortex (VM150), which is the Claude Code host — expect to lose the control session during that window; drive it from elsewhere or accept the outage.
  • One target at a time. Verify health before moving to the next.
  • Snapshot/backup before each mutating step: pct snapshot / qm snapshot (or vzdump) the guest first; DB dump before any DB-bearing upgrade; run the app's own backup where it has one (Nextcloud AIO, mailcow).
  • No host reboots outside the explicit Phase 3 windows. Security upgrades that set reboot-required are applied but activation deferred to Phase 3.
  • Rollback = restore the snapshot / redeploy the previous image tag. Record before/after version in the tracking table.
  • Every result must survive a reboot (standing infra policy).

Load-bearing dependencies (drive the ordering)

  • Authentik (edge2 CT105) = SSO. Upgrading it briefly breaks login to everything behind it. Confirm break-glass admin access first; do in a low-traffic window; verify dependent-service logins after.
  • Caddy (utility CT101) = home-services ingress. A restart blips all home services — fold its update into a deliberate moment, not mid-day.
  • Headscale ×2 (edge2 CT107 main + utility CT106 mesh) = tailnet coordination. Existing tunnels usually persist a restart, but edge2 is the front door and you may be managing over Tailscale — back up the DB and verify nodes stay joined.
  • Nominatim 4→5 is NOT a routine update — it's a re-import project (see Phase 2C).

Phase 0 — Pre-flight (no app changes yet)

  1. Free space on data (92% full) so snapshots/image pulls have room.
  2. Resolve edge2 CT108 mailcow — confirm it's superseded by edge1; if so, back up then stop/destroy → removes it from scope. If it's an unintended outage, that's a separate incident.
  3. Verify break-glass: local admin creds for Authentik-protected apps; PVE console access independent of the services being changed.
  4. Confirm snapshot headroom on each PVE node's storage.
  5. Resolve the edge2 host-kernel CVE question (DirtyFrag/copy.fail) → decides whether edge2 needs a Phase 3 reboot after all.
  6. Freeze the baseline — this doc's versions are the pre-change record.

Phase 1 — OS security packages (no reboot)

Guests first (low blast radius), then hypervisor hosts. Per target: apt-get update → snapshot → apply security-pocket upgrades → needrestart to bounce affected daemons → verify service health and that the security-upgradable count hits 0.

  • Worst-first guests: utility CT119 (91 sec), CT108 (79), cloud CT120/CT121 guest-OS (101/75), media CT110 (37), utility CT109 (31), CT104 (38, incl. PG 16.13→16.14).
  • Remaining guests: all other utility/cloud/media/edge2 CTs + data VM1130 (recon-vm).
  • Hosts (one at a time): data, utility, cloud, media. edge2 host already patched. toc excluded (Phase 3, with cortex).
  • Kernel/libc security that flags reboot-required → applied, reboot deferred to Phase 3.

Phase 2 — Application updates

2A — Self-managed updaters (snapshot → run native updater → verify):

  • OpenTAKServer (utility CT109) — update per OTS upgrade procedure; this is the agreed fix for the EOL RabbitMQ plus MediaMTX/Mumble currency. Snapshot first; verify TAK clients reconnect.
  • Immich (cloud CT120)docker compose pull && up -d in its dir (also clears Valkey 9.0.2→9.1.0). Verify web, mobile sync, ML.
  • Nextcloud AIO (cloud CT121) — run AIO backup → update mastercontainer → trigger child update via AIO UI (:8080). Verify occ status, apps.

2B — Versioned apps (snapshot → bump/upgrade → verify):

  • Authentik (edge2 CT105)sequential, no skipping: 2025.12.4 → 2025.12.6 (backported security, lowest risk) → 2026.2.x → 2026.5.3, running DB migrations at each hop. SSO window + break-glass. Verify dependent logins after each hop.
  • Forgejo (edge2 CT103)branch migration 14.x → 15.0.3, not a patch: read 15.0 breaking changes, back up repos + DB, bump tag, verify.
  • Headscale ×2 (edge2 CT107, utility CT106) — follow the 0.28→0.29 upgrade guide; back up DB; one at a time; verify nodes stay connected.
  • Media stack (media VM105)docker compose pull && up -d for jellyfin / sonarr / radarr / prowlarr / navidrome; read release notes for the majors — SABnzbd 4→5 and Lidarr 2→3; move jellyseerr off the preview-OIDC dev tag to stable 3.3.0. Snapshot VM first; verify each UI.
  • PeerTube (media CT110) — native install: back up DB → run PeerTube's upgrade script 8.0.2→8.2.1 → verify.
  • cortex AI stack (PROTECTED — manual, own window): ollama 0.16→0.30 (14 versions — check model compat), open-webui 0.8→0.9, qdrant, tei. Pull images, verify.
  • Lower-urgency batch: caddy, couchdb (livesync), valhalla, photon, kiwix, navidrome, NATS, valkey-8 sidecar — snapshot + update as convenient.
  • Vaultwarden, PDM, WordPress, Synapse/Element/MAS, obsidian, recon-vm Postgres → already current, skip.

2C — Separate project (not routine):

  • Nominatim 4.5 → 5.3 (recon-vm) — requires the mediagis/nominatim:5.x image + full OSM re-import, coordinated with Photon 1.1→1.2 (1.2 reads the v5 data format). Schedule as its own effort with disk/time budget.

Phase 3 — Platform & reboot windows (coordinated, approval-gated)

  • PVE-9 nodes (data, utility, cloud, media, toc): PVE 9.1.1→9.2, QEMU 10→11, LXC 6→7, kernel 6.17.2→6.17.13. Per node: snapshot/backup guests → apt full-upgrade the proxmox stack → reboot → verify all guests return. One node at a time. Validate on a low-stakes node (data or media) before utility/cloud.
  • toc + cortex = the special window, done LAST and together: toc gets PVE 9.2 + reboot; cortex gets NVIDIA 580.159→580.167 + DKMS + container-toolkit 1.18→1.19 + apt + reboot. The Claude Code host goes down here — run this step from a different control point.
  • pi-nas (standalone window): kernel 6.12→6.18 + OMV 8.1→8.4 (via OMV's update path) → reboot.
  • edge2: reboot only if Phase 0 step 5 confirmed it needs a newer kernel; otherwise none.

Phase 4 — Cross-cutting (fold into earlier phases)

  • Tailscale 1.94→1.98 fleet-wide — rides along Phase 1.
  • Docker CE →29.6 wherever installed — rides along Phase 1/2 under the same snapshots.

Tracking checklist (tick per target)

Enumerated Target Checklist

Phase Target Action Snapshot Update method Verify Rollback Depends-on
0 Break-glass / OOB access matrix (ALL nodes) Verify non-Tailscale, non-Caddy path to every node + each lockout-critical CT (headscale, Authentik CT105, Caddy, Forgejo CT103); confirm PVE/Proxmox web auth is PAM/local not Authentik None (read-only) Log into Contabo/PVE console; pct enter 105/106/107; confirm LAN 192.168.1.x SSH; locate break-glass creds in .ref/credentials; TEST Authentik akadmin + each app local-admin fallback Every node + CT shell reachable WITHOUT Authentik/headscale/Caddy; akadmin login screenshotted N/A FIRST Phase 0 step; gates everything mutating
0 Per-node free space + snapshot backend pvesm status + df -h on data/utility/cloud/media/toc/edge2; record backend (LVM-thin/ZFS/dir-qcow2) + failure mode; size vs largest guests None Read-only; set numeric rule: never snapshot below ~20% free, require ≥2× write-delta Free GB recorded per pool; cloud node measured; vzdump target located N/A Precedes all snapshot-bearing steps
0 data host — disk remediation (92% full, ~73 GB) Free space sized to largest planned snapshot (VM1130) None (cleanup; confirm deletions are cache/backup not live) Prune stale vzdump/snapshots/ISO; docker image prune on recon-vm (keep nominatim:4.5) Usage <~85% on snapshot-backing storage; test snapshot succeeds then removed Restore from Forge/Syncthing/backup if a needed file removed Free-space capture; BLOCKS data host step + canary use
0 Off-host restorable backups (stateful guests) Verify last-good vzdump/PBS off the node for central, OTS, forgejo, matrix, nextcloud, edge2 livesync CouchDB N/A Read-only verification + on-demand pg_dump/app backup copied off-host ≥1 backup off the changing node, restore-testable N/A Precedes all mutating stateful steps
0 Baseline HEALTH capture (all targets) Record up/down, container states, endpoint 200s, apt upgradable/security counts; note oddities None pct/qm status, docker ps, curl probes, apt list --upgradable Baseline recorded; oddities logged (jellyseerr dev tag, CT108 stopped, nominatim pin, CT118 rpcbind) N/A Precedes Phase 1
0 Decision gates Close Open Dec #8 (needrestart tolerance); verify post-cutoff CVE claims vs primary advisories w/ owner+URL; resolve edge2 kernel question; reconcile CT106-vs-CT107 headscale location; define+announce Boise windows None Read-only / sign-off Each decision recorded before Phase 1/2/3 can start N/A Gates Phase 1/2/3
0 edge2 CT108 mailcow (STOPPED) Confirm superseded by edge1, back up, decommission pct snapshot 108 predecommission + vzdump + mailcow native backup Confirm edge1 mail live + no MX at edge2; pct stop/pct destroy 108 ON APPROVAL CT108 gone/archived; edge1 mail in+out works; vzdump restorable pct restore 108 + start; re-point MX edge1 confirmed; approval (Open Dec #9)
0 edge2 host kernel CVE question Decide if DirtyFrag/copy.fail require a reboot None (record pveversion -v) Compare installed proxmox-kernel vs verified-advisory fixed versions; escalate if no-sub repo lacks fix (no repo changes w/o approval) Decision (reboot yes/no) recorded N/A Gates edge2 Phase 3 row
1 Procedure canary — utility CT112 cobalt (idle) or CT102 searxng Prove apt→snapshot→security upgrade→needrestart→verify on low-stakes guest pct snapshot pre-phase1 apt-get update; security pocket only; needrestart -r l then deliberate Security-upgradable=0; service healthy; needrestart clear pct rollback Phase 0; runs BEFORE worst-first guests
1 utility CT119 mesh-territory (179/91, never patched) Security-pocket apt pct snapshot 119 pct exec; security pocket; needrestart -r a sec count=0; meshwars up rollback snapshot After canary
1 utility CT108 meshai (109/79) Security-pocket apt pct snapshot 108 pct exec; security pocket; needrestart sec=0; meshai up rollback snapshot After canary
1 utility CT104 central (PG16 16.13→16.14 + NATS/JetStream) Security apt incl. PG minor; STAGED, controlled bounce pct snapshot 104 + pg_dumpall NEEDRESTART_MODE=l; stage pkgs; quiesce JetStream publishers; bounce PG then NATS deliberately sec=0; select version()=16.14; JetStream/MQTT consumers reconnected rollback snapshot; restore dump Phase 0 #8; stateful carve-out
1 utility CT109 opentakserver (OS layer, 31 sec) Security apt; STAGED, do NOT touch RabbitMQ/MediaMTX/Mumble here pct snapshot 109 NEEDRESTART_MODE=l; stage; deliberate bounce sec=0; TAK clients reconnect rollback snapshot Precedes CT109 Phase 2 app work
1 media CT110 peertube (81/37, native PG16+redis) Security apt; STAGED controlled bounce pct snapshot 110 (PG rides in CT snap) NEEDRESTART_MODE=l; stop import/transcode jobs; snapshot redis; deliberate bounce sec=0; PeerTube plays; PG16/redis/nginx up; nordvpn up rollback snapshot Phase 0 #8
1 data VM1130 recon-vm (34/10, PG16+navi-backend) Security apt; STAGED controlled bounce qm snapshot 1130 pre-phase1-apt NEEDRESTART_MODE=l; deliberate PG + navi-backend bounce sec=0; psql overture/padus; navi-backend/recon.py/photon/kiwix up qm rollback data disk remediation done first
1 cloud CT120 immich (191/101) Guest security apt; Tailscale 1.94→1.98 + Docker CE→29.6 ride along pct snapshot 120 pre-phase1-os pct exec; security pocket; docker bounce acceptable sec=0; immich stack Up; web reachable; TS=1.98 pct rollback Before cloud host step + CT120 app upgrade
1 cloud CT121 nextcloud AIO (98/75) Guest security apt; TS+Docker ride along pct snapshot 121 pre-phase1-os pct exec; security pocket sec=0; 12 AIO containers Up; web reachable pct rollback Before cloud host step + AIO app upgrade
1 media VM105 arr (43/9) Guest security apt; TS+Docker ride along qm snapshot 105 pre-os-sec apt; security pocket; needrestart sec=0; 8 containers Up; Samba + all UIs load qm rollback Before VM105 Phase 2 apps
1 media CT111 mcc (51/13, local caddy+postfix) Guest security apt pct snapshot 111 pre-os-sec pct exec; security pocket sec=0; caddy+postfix active pct rollback Independent (no Phase 2 app)
1 utility CT100/101/102/103/106/107/112/118 (OS layer) Security-pocket apt per CT; lockout-critical (Caddy CT101, headscale CT106) carved to LIST mode + LAST; CT118 leave rpcbind alone pct snapshot each + headscale DB backup for CT106 pct exec; security pocket; needrestart -r l for control-plane CTs sec=0 each; ingress spot-check; headscale nodes list joined pct rollback Phase 0; control-plane CTs last as singletons
1 edge2 CT100-105 (OS layer, via pct exec from edge2 host) Security-pocket apt per CT; DB dumps for CT101 WP MariaDB, CT104 CouchDB, CT105/CT106 PG pct snapshot each pre-phase1 + DB dumps From root@184.174.35.153 pct exec; security pocket; defer reboot sec=0; each app reachable via Caddy front door pct rollback; restore DB dump Phase 0; edge2 host fully patched (host = no-op)
1 edge2 CT106 matrix / CT107 headscale (NO Tailscale client) Security apt driven via pct exec from edge2 HOST only; singletons pct snapshot + Synapse PG dump (106) + headscale DB backup (107) From edge2 host shell; needrestart -r l; do NOT restart headscaled/synapse unless required sec=0; matrix federation+login; headscale nodes list all joined; Caddy ingress works pct rollback; restore DB OOB path confirmed first; NOT in edge2 batch
1 cortex VM150 (PROTECTED, 90 upgradable) Security apt; manual in toc+cortex window; TS+Docker ride along qm snapshot 150 pre-sec-apt (from toc) apt --only-upgrade security pkgs; needrestart -r a; defer reboot sec=0; 5 AI containers Up; nvidia-smi ok qm rollback Phase 0; protected — not bulk
1 pi-nas host OS (subset of 130) Debian security pocket; HOLD kernel 6.12→6.18 + OMV pkgs Off-box config.xml + dpkg --get-selections baseline apt-mark hold linux-image-* openmediavault*; -t trixie-security upgrade; unhold; needrestart sec=0; shares mount; OMV UI loads; no reboot now Reinstall prior pkg from cache; restore config.xml Phase 0; precedes pi-nas kernel reboot
1 Cross-cutting Tailscale 1.94→1.98 (all TS-client nodes) Bump client; rides along Phase 1; EXCLUDE edge2 CT106/CT107 Covered by per-target Phase 1 snapshot apt install --only-upgrade tailscale; daemon restart tailscale version=1.98; node still joined Reinstall 1.94; snapshot rollback Folds into Phase 1; LAN/console fallback confirmed
1 Cross-cutting Docker CE→29.6 (Docker hosts) Bump engine; rides along Phase 1/2 Covered by per-target snapshot apt install --only-upgrade docker-ce docker-ce-cli containerd.io; daemon restart docker version=29.6; containers return Downgrade pkg; snapshot rollback Folds into Phase 1; before compose pulls
2 edge2 CT105 Authentik 2025.12.4→2025.12.6 (hop 1) SSO hop #1 (FIRST load-bearing app); pre-flight: dedup group names, /media/data/media if local storage pct snapshot 105 + pg_dump (per-hop) Edit tag; compose pull && up -d; run migrations; NO skipping UI=12.6; migrations clean; dependent login works compose down; restore tag+dump; or pct rollback Phase 0 break-glass proven; window
2 edge2 CT105 Authentik 2025.12.6→2026.2.x (hop 2) SSO hop #2 cross-major pct snapshot + fresh pg_dump tag→2026.2.x; pull/up; migrations UI=2026.2.x; logins re-verified restore prior tag+dump after hop 1
2 edge2 CT105 Authentik 2026.2.x→2026.5.3 (hop 3, final) SSO hop #3; verify CVE/GHSA IDs; insert any mandatory intermediate pct snapshot + fresh pg_dump tag→2026.5.3; pull/up; final migrations UI=2026.5.3; full sweep of dependent logins; close window restore prior tag+dump after hop 2; verify dependents BEFORE upgrading them
2 utility CT101 + media CT111 Caddy (ingress, EARLY) Caddy bump (before home-app verification); map Caddy↔Authentik path; document direct Authentik admin URL pct snapshot + Caddyfile/certs backup caddy validate; restart in low-traffic window home services reachable; certs valid rollback snapshot/Caddyfile After CT101 Phase 1; separate window from Authentik
2 edge2 CT103 Forgejo 14.0.5→15.0.3 (branch migration) Major migration; pre: off-edge2 git clone --mirror echo6-docs + pause autocommit cron + forgejo doctor check --all --fix pct snapshot 103 + repos volume + pg_dump + app.ini doctor check→repair→backup→bump tag 14→15; pull/up; migrations reports 15.0.3; clone/push works; CI/webhooks ok restore 14.0.5 image + volume + dump After CT103 Phase 1; NOT same window as Authentik
2 utility CT109 OTS app OTS updater (app+webUI+schema) — does NOT migrate RabbitMQ pct snapshot 109 + PG dump Run OTS upgrade script OTS web/API up; CoT works; clients reconnect rollback snapshot + dump After CT109 Phase 1 OS
2 utility CT109 RabbitMQ 3.12.1→4.2.8 (SEPARATE) Decoupled EOL major: Erlang≥26→3.13.x→enable_feature_flag all→4.x pct snapshot + rabbitmqctl export_definitions staged hops; confirm flags enabled; no classic-mirroring config 4.2.x; queues intact; TAK clients reconnect restore snapshot + definitions GATES MediaMTX/Mumble accept; after OTS app
2 utility CT106 + edge2 CT107 Headscale 0.28→0.29 Lower-stakes instance FIRST (rehearsal), front-door LAST; disable key-expiry; OOB-driven pct snapshot + headscale DB dump off-tailnet stop→backup→replace binary→config migrate→start; headplane to match version=0.29.1; nodes list all joined; tunnels intact restore snapshot+DB via console one-at-a-time; non-tailnet path; not same window as Caddy/Authentik/Forgejo
2 cloud CT120 Immich 2.5.6→2.7.5 (+Valkey 9.0.2→9.1.0, vectorchord PG) Stack upgrade; prune images first; do NOT interrupt first-boot migration; disable bg regen pct snapshot 120 pre-immich-2.7.5 + pg_dump compose pull && up -d; let migrations run; no skip across breaking migration web+login; jobs process; valkey 9.1.0 ping; version 2.7.5 restore tag (only with dump) or pct rollback After CT120 Phase 1; free space confirmed
2 cloud CT121 Nextcloud AIO (master ~12.5→13.2.1, NC 32.0.4→32.0.11) AIO orchestrated update; borg backup as rollback artifact AIO borg backup + pct snapshot 121 AIO backup→stop→update mastercontainer→UI :8080 "Update containers" 12 containers green; occ status 32.0.11; integrity clean restore from AIO borg; else pct rollback After CT121 Phase 1; AIO self-gates intermediate versions
2 media CT110 PeerTube 8.0.2→8.2.1 (native) Native upgrade script; quiesce jobs pct snapshot 110 + pg_dump peertube_prod back up DB → run PeerTube upgrade.sh/version steps → restart reports 8.2.1; video plays; PG16/redis/nginx healthy pct rollback + restore dump After CT110 Phase 1
2 media VM105 jellyfin 10.11.6→10.11.11 Patch bump qm snapshot 105 (pre-app batch) compose pull jellyfin && up -d (pin tag) UI=10.11.11; plays/transcodes redeploy prior digest; qm rollback After VM105 Phase 1
2 media VM105 SABnzbd 4.5.5→5.0.4 (MAJOR) Major; pause queue; audit post-proc scripts qm snapshot 105 pre-sabnzbd + sabnzbd.ini pin 5.0.4; pull/up; migrate on start UI=5.0.4; queue intact; test NZB; arr integrations ok redeploy 4.5.5 + restore ini (queue repair); qm rollback After VM105 Phase 1; before Lidarr
2 media VM105 Lidarr 2.x→3.x (MAJOR) Major branch migration qm snapshot 105 pre-lidarr + lidarr.db pin v3; pull/up; DB migration on start UI=v3; library+indexers+download client ok restore db (migration one-way); qm rollback After SABnzbd (verify download-client link)
2 media VM105 Sonarr/Radarr/Prowlarr (1-2 ver) Routine bumps pre-app VM snapshot + config DBs compose pull && up -d UIs load; prowlarr sync; test grab redeploy prior tags + DBs After SABnzbd/Lidarr
2 media VM105 Navidrome 0.60.3→0.62.0 Routine bump pre-app snapshot + DB compose pull navidrome && up -d UI=0.62.0; library scans; track streams redeploy 0.60.3 + DB Independent of arr chain
2 media VM105 jellyseerr preview-OIDC→stable 3.3.0 DECISION GATE: schema downgrade + correct artifact (jellyseerr vs Seerr) + OIDC parity qm snapshot 105 pre-jellyseerr + config dir test stable vs DB COPY first; swap tag only if compatible UI=3.3.0; requests intact; OIDC/login works restore preview tag + config; qm rollback After jellyfin/arr; STOP+report if OIDC parity unclear
2 edge2 CT104 livesync CouchDB 3.4→3.5.2 Lower-urgency image bump pct snapshot 104 + CouchDB volume backup bump tag; compose pull && up -d reports 3.5.2; _up healthy; LiveSync replicates test edit restore 3.4 + volume; pct rollback After CT104 Phase 1; triple-sync verify
2 utility CT104 NATS 2.14.0→2.14.2 Lower-urgency patch pct snapshot 104 (JetStream dir in snap) replace binary; restart 2.14.2; JetStream/MQTT intact pct rollback After CT104 Phase 1 (PG already done); stateful blip
2 utility CT100 meshmonitor / CT102 searxng+valkey-8 (8.1.5→8.1.8) Lower-urgency image refresh pct snapshot + record digests compose pull && up -d containers Up; function ok redeploy prior digest; pct rollback After respective Phase 1
2 edge2 CT100 PDM / CT101 WP core / CT102 Vaultwarden / CT106 Synapse Verify-current / no-op (already at/past fix); WP plugin/theme = separate WP-CLI follow-up None (Phase 1 snaps) optional digest re-pull only versions confirmed; logged as current redeploy prior digest Closes as no-op
2 cortex Ollama 0.16.1→0.30.10 (PROTECTED) Image bump; no downgrade after new pull qm snapshot 150 pre-ollama + models volume compose pull ollama && up -d; test inference /api/version=0.30.10; models intact; GPU used; vault-tagger works re-pin 0.16.1; restore snapshot if model format broke After cortex Phase 1 + Docker + toolkit
2 cortex Open-WebUI 0.8.1→0.9.6 Image bump (auth surface) qm snapshot 150 pre-openwebui bump tag; pull/up UI loads; login works; reaches Ollama re-pin 0.8.1; snapshot After Ollama
2 cortex Qdrant 1.16.3→1.18.2 / TEI 1.7.4→1.9.3 Lower-urgency image bumps qm snapshot 150 per app bump tag; pull/up health endpoints ok; collections/embeddings intact; docs engine works re-pin prior; snapshot After cortex Docker+toolkit; verify engines
2 cortex obsidian-remote v1.12.7 Verify-current / no-op None none container up; UI reachable N/A
2 utility CT118 archivist rpcbind 0.0.0.0:111 Remediate exposure (network change — APPROVAL) pct snapshot 118 per approved option: disable rpcbind / bind localhost+TS / firewall port 111 not on 0.0.0.0; app functions; external scan closed re-enable / pct rollback After CT118 Phase 1; approval-gated (Open Dec #6)
2 edge2 CT108 mailcow ./update.sh — CONDITIONAL Only if RETAINED; else SKIP pct snapshot 108 + mailcow backup ./update.sh (native lockstep) 18 containers up; mail in/out ./update.sh restore; pct rollback GATED on Phase 0 decision (likely skipped)
3 Per-node corosync/HA pre-flight (data/utility/cloud/media/toc) Before EACH reboot None pvecm status (Quorate:Yes, 5 votes); ha-manager status quorate + expected votes=5; HA implications known abort if not quorate Gates each Phase 3 node reboot
3 Canary node (lowest blast + headroom; media, NOT data until disk freed) Full 9.2/QEMU11/LXC7/kernel reboot to prove the path vzdump all guests to EXTERNAL target + record pveversion -v apt update && apt dist-upgrade → reboot; NO live-migration in mixed window pveversion=9.2/QEMU11/LXC7/kernel 6.17.13; all guests onboot return; quorate boot prior kernel (GRUB); restore guests from vzdump Phase 1+2 done; canary before high-stakes nodes
3 media host PVE 9.1.1→9.2 (+QEMU11/LXC7/kernel) Platform + reboot vzdump VM105/CT110/CT111 + snapshots dist-upgrade → reboot guests return; arr/peertube/caddy healthy; quorate GRUB prior kernel; vzdump restore After media Phase 1+2; one node at a time
3 cloud host PVE 9.2 platform + reboot Platform + reboot vzdump CT120/CT121 + record versions dist-upgrade → reboot immich+nextcloud return healthy; quorate GRUB prior kernel; vzdump restore After cloud Phase 1+2; after canary proven
3 utility host PVE 9.2 platform + reboot Platform + reboot (known full home-ingress + mesh + central outage) vzdump all 12 guests + versions dist-upgrade → reboot all 12 return; Caddy ingress + headscale + central back; quorate GRUB prior kernel; vzdump restore LAST of early nodes; low-traffic window
3 data host PVE 9.2 platform + reboot Platform + reboot vzdump VM1130 + versions dist-upgrade → reboot VM1130 returns; NFS/Samba serve; quorate GRUB prior kernel; vzdump restore After disk remediation; one at a time
3 Hypervisor host OS-security (folded in) Apply held host libc/openssl WITH the platform full-upgrade per node covered by per-node vzdump included in dist-upgrade (not applied back in Phase 1) sec=0 host-side post-reboot; daemons up per-node rollback Fold-in decision (Open Dec #2)
3 Post-reboot app re-verification (per node) Re-run Phase-2 health checks for majored guests on each rebooted node None health probes CT120/CT121/CT110/VM105/CT109 healthy on new substrate re-snapshot/redeploy degraded component After each node reboot
3 cortex nvidia-container-toolkit 1.18→1.19 Toolkit bump (pair w/ driver) qm snapshot 150 pre-nvtoolkit apt --only-upgrade nvidia-container-toolkit; nvidia-ctk runtime configure; restart docker --version=1.19; docker run --gpus all nvidia-smi ok downgrade 1.18; snapshot toc+cortex window
3 cortex NVIDIA driver 580.159→580.167 + DKMS (reboot) Driver + DKMS as discrete reversible step BEFORE toc reboot qm snapshot 150 pre-nvidia + vzdump apt --only-upgrade nvidia-driver-580 nvidia-dkms-580; DKMS rebuild; verify BEFORE toc reboot nvidia-smi=580.167; dkms status installed; containers see GPU restore snapshot; reinstall 580.159 + DKMS Keep old driver pkg; after toolkit
3 toc host PVE 9.2 + QEMU11 + LXC7 + kernel (reboot — LAST, drops cortex) Platform + reboot; only guest is VM150 vzdump/snapshot VM150 + record pveversion snapshot VM150 → dist-upgrade toc → reboot → VM150 returns → cortex driver reboot pveversion=9.2; toc rejoins (5/5 votes); VM150 + GPU healthy restore VM150 from vzdump; GRUB prior kernel LAST + alone; alternate control point TESTED; all 4 other nodes quorate
3 cortex final stack health verification Post-window full AI-stack + engine check keep pre-window snaps until verified verification only driver 580.167; all 5 containers; ollama/tei/qdrant/open-webui; vault-tagger+TEI engines ok; TS online; Docker 29.6 restore degraded component; worst case VM150 vzdump LAST action of node group
3 pi-nas OMV 8.1→8.4 (reboot #1, separate window) OMV native update path (NOT raw apt) → reboot off-box config.xml export + plugin/layout record OMV UI Update Mgmt or omv-upgrade; reboot OMV=8.4; shares/SMB/NFS/RAID/mergerfs healthy; client mounts restore config.xml + dpkg --set-selections 8.1 + omv-salt deploy Physical access; not mid-sync; after Phase 1
3 pi-nas kernel 6.12→6.18 (reboot #2, separate window) Kernel bump → reboot; activates deferred security back up /boot+/boot/firmware; keep 6.12 installed; fresh config.xml apt install linux-image-arm64 (unhold); update bootloader; reboot ON APPROVAL uname -r=6.18; LAN+tailnet up; shares mount; Docker up boot retained 6.12; restore /boot; on-site SD reflash After OMV reboot verified; physical access
3 edge2 host PVE 8.4 reboot — CONDITIONAL Only if Phase 0 confirmed kernel fix needed; own no-failover window AFTER CT105/CT107 verified vzdump all guests + pveversion scoped kernel/security upgrade → reboot in approved window uname -r=fixed; all CTs return; Caddy-fronted services reachable GRUB prior kernel; vzdump restore GATED on Phase 0; announce SSO+tailnet downtime; non-tailnet path
Nominatim 4.5→5.3 + Photon 1.1→1.2 (SEPARATE PROJECT) Full re-import into separate DB/instance; keep 4.5 until validated; Photon AFTER nominatim validated qm snapshot 1130 + retain 4.5 image/data import to new DB; size flat-nodes vs free space (NOT on data); cut over after verify 5.x geocode correct; Photon 1.2 reads v5 re-pin nominatim:4.5 + old data Wholly separate from patch campaign; after disk remediation

Plan Hardening (from review)

The review surfaced corrections across six dimensions. They are merged and de-duplicated below, critical/high first, grouped by phase, with hosts/CTs named. Several issues recur because the same architectural fact (one corosync cluster; control plane runs over the tailnet; data is 92% full; SSO + ingress + git + tailnet are mutually entangled) drives multiple failure modes — the consolidated corrections address the root cause once.

Cluster & control-plane facts that the whole plan must respect

  • The five PVE nodes are ONE corosync cluster (echo6-cluster), not standalone hypervisors. Confirmed in vault/docs/hardware/environment.md (line 17) and vault/runbooks/proxmox-onboard-node.md. Quorum = 3 of 5. The plan never mentions corosync/quorum. Every Phase 3 node reboot must be gated on pvecm statusQuorate: Yes, expected votes = 5; reboot ONE node at a time; wait for full rejoin (5/5) before the next. Losing quorum makes /etc/pve read-only fleet-wide and freezes all VM/CT start/stop/config on every surviving node. Also run ha-manager status first — if any guest is HA-managed, a reboot triggers fencing/auto-migration, not a clean local stop/start.
  • toc is the 5th cluster vote AND the GPU host for cortex (VM150) AND cortex is the documented cluster "Management host" + the Claude Code control host. Rebooting toc removes a vote and kills the box you'd diagnose from. toc+cortex must be the final cluster action and must never overlap any other node reboot (toc down + a second node down = bare 3/5; one corosync flap loses quorum).
  • Reconcile the headscale location contradiction BEFORE touching any of it. The audit places headscale at utility CT106 and edge2 CT107 (and calls CT106 both "meshtastic-hs" and "headscale control plane"); the vault docs (headscale-onboard-node.md) describe headscale as a Docker container on Contabo (100.64.0.1) reached over the VPN. You cannot protect a control plane you have mislocated. Pin the real instance(s) and their reachability first; this gates every headscale step below.

Phase 0 — preflight (promote these to hard gates; some must move to first)

  • (MOVE TO FIRST) Verify a non-Tailscale, non-Caddy break-glass path to every node — Contabo/Proxmox console for edge2, LAN 192.168.1.x for the home nodes, PVE noVNC into each CT — and confirm PVE/Proxmox web auth is PAM/local, NOT behind Authentik. Losing the management path mid-run is the single highest-consequence failure on this fleet. This must pass before anything mutating.
  • Measure, don't hand-wave, snapshot headroom. Run pvesm status + df -h on data/utility/cloud/media/toc/edge2. Record the storage backend per node (LVM-thin/ZFS/dir-qcow2) and its over-provision failure mode. Set a hard numeric rule: never start a snapshot below ~20% free; require ≥2× the largest guest's expected write-delta. data at 92% (~73 GB free) cannot safely snapshot VM1130 recon-vm (nominatim/overture/padus), and cloud node free space (CT120 immich, CT121 nextcloud AIO) is not measured at all — add it. Identify where vzdump lands; if it's data, freeing data is a fleet-wide rollback prerequisite, not a data-node nicety.
  • data disk remediation is a blocking gate, sized from measurement. Identify consumers (old vzdump, snapshots, ISO/template store, recon-vm dangling Docker layers — do NOT remove the pinned nominatim:4.5 image), confirm nothing deleted is the only copy of a backup or live NAS data, and size cleanup to the largest planned snapshot on the node. Must complete before data's own host step and before it is used as a canary.
  • Verify off-host restorable backups exist (distinct from per-step snapshots) for every stateful guest: central PG/NATS, opentakserver PG + RabbitMQ definitions, forgejo PG + repos, matrix Synapse PG, nextcloud AIO (borg), edge2 livesync CouchDB. At least one copy must live off the node being changed.
  • Capture a pre-change baseline HEALTH snapshot (per-target up/down, docker ps/pct/qm status, key endpoint 200 checks, and apt list --upgradable/security counts) so "verify after" has a comparison and known oddities are recorded: jellyseerr on the preview-OIDC dev tag, mailcow CT108 already STOPPED, pinned-stale nominatim, CT118 archivist rpcbind on 0.0.0.0:111.
  • Close Open Decision #8 as a gate: explicit sign-off on needrestart-driven daemon bounces for central (CT104 PG16/NATS/JetStream), opentakserver (CT109), peertube (CT110), matrix Synapse, recon-vm VM1130 PG16 + navi-backend, and the auth/tailnet daemons — or set those guests NEEDRESTART_MODE=l (list-only) so the operator controls the bounce.
  • Verify the load-bearing post-cutoff CVE claims against primary advisories before they drive ordering (Authentik May-2026 waves justifying the mandatory 2025.12.6 hop; Valkey 9.1.0; Immich 2.6/2.7; RabbitMQ 4.x; edge2 DirtyFrag CVE-2026-43284/-43500 and copy.fail CVE-2026-31431). Assign an owner + primary-source URL per claim. Resolve the edge2 kernel question fully here (pin uname -r/proxmox-boot-tool kernel list, check repo kernel availability) so edge2 is either firmly scheduled for its own no-failover reboot window or firmly excluded — not left conditional inside Phase 3.
  • Decommission decision for mailcow CT108 (STOPPED, superseded by edge1): vzdump + mailcow native backup, confirm edge1 is handling mail and no MX points at edge2, then pct destroy only on approval. Removes CT108 from all later scope.
  • Define and announce maintenance windows in America/Boise, naming user-facing blips: Authentik SSO (the upgrade "briefly breaks login to everything"), all home services behind Caddy, matrix CT106, media VM105, immich/nextcloud/peertube.

Phase 1 — guest OS security sweep (ordering + needrestart policy)

  • Prove the procedure on a low-stakes guest first (utility CT112 cobalt idle, or CT102 searxng) before the never-patched worst-first targets (CT119 "never patched", 179 upgradable / 91 sec; CT108 meshai). Keep worst-first for security urgency only AFTER the procedure is proven; snapshot each never-patched guest immediately before its first-ever security upgrade.
  • Carve the stateful guests out of the bulk needrestart pass (CT104 central, CT109 opentakserver, CT110 peertube, VM1130 recon-vm): set NEEDRESTART_MODE=l, stage the packages, take a logical DB dump (pg_dump/JetStream snapshot) with the app quiesced, then bounce each service deliberately in a controlled order. A JetStream restart mid-write loses messages on an at-most-once pipeline; a RabbitMQ/redis bounce mid-job is not a clean "blip."
  • Carve the lockout-critical daemons out too (Authentik CT105, the headscale instance(s), Caddy, and sshd on any host you're connected over). Run needrestart -r l to see what would restart and bounce deliberately, so a libc/openssl bump can't auto-bounce the auth/tailnet/SSH daemons out from under the operator mid-pass.
  • Order control-plane guests LAST as singletons — the home-ingress Caddy and the headscale instance(s) — each with DB backup + node-stays-joined verification, never inside a batch. Before touching headscale, confirm the out-of-band path (LAN SSH / PVE console via pct exec) to every still-queued target.
  • edge2 CT106 matrix / CT107 headscale have NO Tailscale client (Caddy-only ingress). Drive their security-apt step via pct exec from the edge2 HOST (root@184.174.35.153), never over service ingress, so a needrestart bounce of a network/resolver daemon can't strand the only path in. Treat them as their own singleton steps, not part of the edge2 batch.
  • Gate data's host step on the disk remediation being verified DONE (its snapshot guardrail needs the space).

Phase 1 hosts vs Phase 3 (sequencing)

  • Fold the four cluster hosts' OS-security apt into their own Phase 3 reboot window rather than applying host libc/openssl in Phase 1 and living mixed-state (new libc / old kernel, bounced smbd) across the entire Phase 2 app campaign. Guests still get Phase-1 security immediately; each host transitions in one clean window. (Open Decision #2 — this is the sequencing-correct answer.)

Phase 2 — app upgrades (sequence so a failure is isolatable, and never co-schedule recovery paths)

  • Hard rule: {Headscale, Authentik, Caddy, Forgejo} are NEVER in the same maintenance window so at least one of {tailnet, SSO, ingress, git} is always a known-good recovery path.
  • Authentik first among load-bearing apps (it fronts Forgejo, Vaultwarden, WordPress, Nextcloud, Matrix). Prove break-glass with evidence first: log in with the local akadmin in a private session; confirm each protected app has a working local-admin fallback (Forgejo admin user, Nextcloud occ, PDM/Vaultwarden local). The chain 2025.12.4 → 2025.12.6 → 2026.2.x → 2026.5.3 runs irreversible Django migrations at each hop — pg_dump per hop (labelled), one hop at a time, verified-login gate between each, abort+restore on first migration error. Add the 2025.12 hard pre-flight: audit/rename duplicate group names (migration "fails loudly" on dupes), and if local /media storage is used, stop and mv ./media ./data/media + rewrite the compose volume before starting the new version. Verify all dependent SSO logins BEFORE upgrading the dependents.
  • Caddy early (utility ingress + media CT111), before home-service app verification, so backend sign-offs run through the new ingress; pair CT101/CT111 in one deliberate window. Map the Caddy↔Authentik path first and document a direct (Caddy-bypassing) Authentik admin URL as break-glass.
  • OpenTAKServer CT109 — decouple RabbitMQ from the OTS bump. The OTS updater only upgrades OTS + webUI + schema; it does NOT migrate RabbitMQ, Erlang, or feature flags (verified against docs.opentakserver.io/installation/upgrading.html), so the most security-urgent EOL item (RabbitMQ 3.12.1, no 3.x backports) would silently go unfixed. Write RabbitMQ as its own sub-task: snapshot + export_definitions → Erlang ≥26 → 3.12 → 3.13.x → rabbitmqctl enable_feature_flag all (confirm all enabled) → only then 4.x; confirm no classic-queue-mirroring config remains. A naive 3.12→4.x jump refuses to boot. Gate the OTS step on RabbitMQ 4.x up + TAK clients reconnecting before accepting the bundled MediaMTX/Mumble bumps.
  • Forgejo CT103 14→15 is a branch migration on an EOL branch (14.x EOL 2026-04-30) with forward-only FK migrations. Run forgejo doctor check --all [--fix] to repair stopwatch/tracked_time inconsistencies BEFORE the backup+bump. Pre-step: git clone --mirror echo6-docs (and other critical repos) to an OFF-edge2 location and pause the echo6-docs-autocommit cron so the recovery docs survive an edge2/Forgejo failure and the cron doesn't loop-error into a half-migrated instance.
  • Headscale: upgrade the lower-stakes instance FIRST as a rehearsal, the front-door LAST, each driven from a path that does NOT depend on the instance being upgraded (PVE console / direct edge2-host → CT internal IP). Before the front-door hop: disable/extend node key-expiry on all nodes so persistent tunnels don't drop on re-handshake mid-migration; pct snapshot + dump the headscale DB off-tailnet; keep a second already-authenticated SSH session open. Rollback restores the snapshot via console, not via tailnet. Verify headplane compatibility with 0.29.
  • Immich CT120 / Nextcloud AIO CT121 — method is correct, but size disk and protect the DB. Prune old images + docker image prune before the snapshot; do NOT interrupt Immich's first-boot 2.5→2.7 DB migration; AIO self-gates intermediate mastercontainer versions (a stale AIO may need two cycles) — use AIO's borg backup as the rollback artifact (mastercontainer downgrades unsupported). The rollback unit for every DB-bearing app is a quiesced logical dump + old binary together — "redeploy previous image tag" is NOT valid after forward-only migrations, and a pct snapshot of a hot, separate-volume DB can restore torn. Disable Immich background regeneration until free space is confirmed.
  • jellyseerr is a data-compatibility decision, not a retag. preview-OIDC (dev) → stable 3.3.0 is a schema DOWNGRADE; TypeORM stable may refuse to start against a dev-migrated DB. Confirm the correct destination artifact (jellyseerr stable vs the unified "Seerr" successor) and, if OIDC is in active use, confirm 3.3.0 parity — else STOP and report. Test the stable image against a COPY of the DB first.
  • cortex ollama 0.16→0.30: reframe — old local blobs read fine; the real trap is no clean downgrade after pulling any new model. Snapshot the models volume (the only rollback), upgrade, verify existing models ollama run without re-pull. Verify the vault-tagger engine (localhost:11434) and TEI related: engine (localhost:8090) still work after.
  • SABnzbd 4.5→5.0: pause/empty the queue, audit custom post-proc scripts (scripts now run on failed jobs), note empty_postproc removed and that downgrade needs a queue repair.
  • Nominatim 4→5 + Photon 1.1→1.2 is a separate project, not part of the patch campaign. Re-import into a SEPARATE DB/instance; keep the nominatim:4.5 image + data until 5.x is validated; size transient import disk (flat-nodes tens of GB) against real free space on a non-starved node — this cannot live on the 92%-full data pool; sequence Photon 1.2 AFTER nominatim 5.x is validated, never concurrently.

Phase 3 — platform & reboot windows

  • Per-node order: apt update && apt dist-upgrade (full-upgrade so proxmox-ve/qemu/lxc metapackages pull) → reboot → verify pveversion + all guests return. Hard rule for the mixed-version window: NO live migration and disable HA migration across the QEMU 10/11 (PVE 9.1/9.2) boundary — running VMs keep QEMU 10 machine type until cold-started, and migration across the boundary fails. Optionally bump VM machine types + cold-restart after all 5 are on 9.2 (noted follow-up).
  • Canary by blast radius + snapshot headroom, not label. data (92% full, hosts recon-vm overture/padus) must NOT be canary until its disk is remediated, and its near-full pool can fail the mandatory pre-reboot snapshot. Lowest-stakes is media (single arr VM + peertube, non-auth/non-DB-critical) — but note media carries freshly-majored apps; the reconciled order is: canary the genuinely lowest-blast node with headroom → then utility LAST among the early nodes (it carries Caddy ingress + the mesh/central stack — a deliberately-scheduled full home-ingress + mesh-coordination outage) → cloud (immich/nextcloud) → data → toc+cortex last and alone, all four other nodes confirmed quorate.
  • Pre-window: confirm every guest has onboot=1 so they actually return; take vzdump of guests to an EXTERNAL target (not the local pool, never data) so a guest that won't start under LXC 7 can be restored elsewhere. Read PVE 9.2 + LXC 7 release notes for unprivileged-CT/cgroupv2/apparmor breaking changes before the first node.
  • Add a post-Phase-3 re-verification gate: after each node boots into 9.2/LXC7, re-run the Phase-2 health check for every guest that received a major app upgrade on that node (CT120 immich, CT121 nextcloud, CT110 peertube, VM105 arr, CT109 OTS) — the substrate changed under them.
  • toc+cortex window: name and TEST a concrete alternate control point (a non-cortex host with keys + tooling to the fleet) BEFORE the window; stabilize the tailnet well before it (never in the same period as a headscale upgrade). Split cortex: snapshot VM150 from toc → NVIDIA 580.159→580.167 + DKMS + container-toolkit 1.19 as a discrete reversible step, confirm dkms status built and nvidia-smi BEFORE rebooting toc → reboot toc → confirm toc rejoins (5/5 votes) and cortex + GPU passthrough return. Keep the old driver package for reinstall.
  • edge2 is PVE 8.4, architecturally separate, single SPOF for ingress (no failover). Do NOT fold it into the cluster windows. If the Phase 0 kernel verification says it needs a reboot, give it a dedicated window AFTER Authentik CT105 + headscale CT107 are upgraded and verified, announce SSO+tailnet downtime, drive from a non-tailnet path, and confirm all CTs auto-start.
  • pi-nas: split into two reboots, not one. OMV 8.1→8.4 via OMV's own update path (NOT raw apt full-upgrade) → reboot → verify shares/SMB/NFS/omv-salt; THEN kernel 6.12→6.18 → reboot → verify boot + disk remount. Confirm physical/serial console access (no remote KVM on an RPi), back up /boot+/boot/firmware and keep the old kernel installed as fallback, export config.xml off-box, and confirm it's not mid-sync as a Syncthing/backup target before taking it down. Independent of the cluster windows.

Full Inventory

Complete point-in-time state of every node, guest, and container service.

data (PVE 9.1.1)

  • Host: 126 upgradable / 51 security; no Docker installed; roles: NAS, NFS, Samba
  • VM1130 recon-vm (Ubuntu 24.04) — 34 upgradable / 10 security
    • PostgreSQL 16.14 (DBs: overture, padus), photon, kiwix, recon.py, 7x navi-backend, nginx, Apache, Samba
    • Docker: valhalla:latest, nominatim:4.5 (stale 14 months), zimit:latest (not running)

utility (PVE 9.1.1)

  • Host: 148 upgradable / 38 security; PVE 9.2 platform update pending; 12 LXC guests
CT Name Upgradable / Sec Services
CT100 meshmonitor 62 / 15 ghcr.io/yeraze/meshmonitor:latest
CT101 caddy (home ingress) 59 / 18
CT102 searxng 54 / 15 searxng/searxng:latest + valkey/valkey:8-alpine
CT103 argus 29 / 26 RF capture / transcribe / viewer
CT104 central 50 / 38 PostgreSQL 16 + NATS/MQTT/JetStream
CT106 meshtastic-hs 54 / 17 headscale control plane
CT107 mesh-bridge 53 / 15 dual tailscaled
CT108 meshai 109 / 79 work-meshai local build
CT109 opentakserver 34 / 31 nginx / PG16 / rabbitmq / mumble / mediamtx / CoT
CT112 cobalt 35 / 29 build/CI (idle)
CT118 archivist 62 / 17 archivist + rpcbind (FLAG: port 111 on 0.0.0.0)
CT119 mesh-territory 179 / 91 meshwars:latest (never patched)

cloud (PVE 9.1.1)

  • Host: 113 upgradable / 38 security; 2 LXC guests
CT Name Upgradable / Sec Services
CT120 immich 191 / 101 immich_server, immich_machine_learning, valkey/valkey:9, immich postgres (14-vectorchord) — all drifted
CT121 nextcloud AIO 98 / 75 12 containers: mastercontainer + apache + nextcloud + postgresql + redis + collabora + clamav + imaginary + fulltextsearch + notify-push + whiteboard + docker-socket-proxy; NC 32.0.4; mastercontainer behind

media (PVE 9.1.1)

  • Host: 107 upgradable / 36 security; 3 guests
Guest Name Upgradable / Sec Services
VM105 arr 43 / 9 (Ubuntu 24.04) jellyfin / sonarr / radarr / prowlarr / sabnzbd / lidarr / navidrome / jellyseerr (preview-OIDC) all :latest + Samba
CT110 peertube 81 / 37 v8.0.2; nginx / PG16 / redis / peertube / pt-downloader + importer + monitor / nordvpn
CT111 mcc 51 / 13 caddy + postfix

toc (PVE 9.1.1) — PROTECTED

  • Host: 189 upgradable / 43 security; PVE 9.2 platform update pending; reboot required
  • Hosts only VM150 cortex — coordinate any toc work with cortex maintenance window

cortex (VM150, Ubuntu 24.04) — PROTECTED GPU / Claude Code host

  • 90 apt upgradable
  • NVIDIA driver 580.159 → 580.167 + DKMS (reboot required)
  • nvidia-container-toolkit 1.18 → 1.19
  • Docker: ollama / tei 1.7 / qdrant / open-webui / obsidian — all drifted

pi-nas (Debian 13, arm64, RPi + OMV)

  • 130 apt upgradable; kernel 6.12 → 6.18 (reboot required); OMV 8.1 → 8.4
  • Docker engine installed; 0 containers running

edge2 (PVE 8.4.19) — Host Fully Patched

  • 9 LXC guests
CT Name Services / Status
CT100 pdm PDM 1.1.4, current (native)
CT101 wordpress Apache 2.4.67 / PHP 8.4 / MariaDB 11.8.6 / WP core 7.0 — plugin/theme status needs WP-CLI
CT102 vaultwarden vaultwarden/server:latest — drift indeterminate
CT103 forgejo 14.0.5 (forgejo:14 + postgres:16-alpine drifted)
CT104 livesync couchdb:3.4 drifted + local provisioner
CT105 authentik 2025.12.4 (server/worker/postgres) → upgrade to 2026.2.0
CT106 matrix Synapse 1.155.0 / Element / MAS + mautrix-signal + postgres; OS apt security updates pending; no Tailscale client
CT107 headscale 0.28.0 → 0.29.0 + headplane; no Tailscale client
CT108 mailcow 18 containers (postfix/nginx/dovecot/rspamd/clamd/sogo/php-fpm/mariadb/redis/memcached/unbound/acme/olefy/watchdog/netfilter/ofelia/dockerapi/tlspol) — upgrade via update.sh only

Coverage Notes

  • Complete: all five PVE-9 nodes (data, utility, cloud, media, toc), edge2, cortex, pi-nas, and all discovered guests/containers.
  • Excluded: edge1 (the rebuilt ex-Contabo, mail-only) — mid-rebuild/maintenance at time of audit; re-audit once it settles.
  • Incomplete: WordPress (CT101) plugin and theme status — requires WP-CLI; not assessed.
  • Method: read-only throughout — SSH / pct exec, apt list --upgradable, docker manifest inspect for same-tag drift. Floating/pinned-tag caveats noted inline where drift could not be confirmed.