auto: docs sync 2026-07-13T12:00:23+00:00

Files changed: engine/.embcache.json engine/changelog.md engine/lint-report.md vault/.trash/2026-06-19.md vault/docs/hardware/environment.md vault/docs/hardware/ip-allocation.md vault/docs/matrix/archivist.md vault/docs/matrix/matrix_host.md vault/docs/matrix/mautrix_signal.md vault/docs/matrix/synapse.md vault/docs/matrix/synapse_retention_discovery.md vault/docs/navi/cc-rules.md vault/docs/navi/deployment.md vault/docs/navi/themes.md vault/docs/services/ots-setup.md vault/docs/services/services.md vault/docs/services/usenet.md vault/docs/software/authentik.md vault/docs/software/caddy.md vault/docs/software/central.md vault/docs/software/dns.md vault/docs/software/geo-tools.md vault/docs/software/navi.md vault/docs/software/recon.md vault/docs/software/searxng.md vault/glossary.md vault/notes/echo6-landing-page-data-export.md vault/notes/ia-download-queue.md vault/projects/advbbs-project.md vault/projects/argus.md vault/projects/deploy-livesync.md vault/projects/fleet-patch-audit.md vault/projects/fleet-platform-baseline.md vault/projects/matrix-synapse-deployment.md vault/projects/meshai-config-hot-apply.md vault/projects/meshai-region-routing-plan.md vault/projects/meshai.md vault/projects/meshcore-transport.md vault/projects/meshtastic-headscale-runbook.md vault/projects/mmud-project.md vault/projects/nominatim-v5-reimport.md vault/runbooks/add-peertube-channel.md vault/runbooks/authentik-access-groups.md vault/runbooks/authentik-create-invitation.md vault/runbooks/authentik-oidc-application.md vault/runbooks/authentik-upgrade.md vault/runbooks/central-deploy-cutover.md vault/runbooks/ct-runbook.md vault/runbooks/edge2-access-reference.md vault/runbooks/expose-service-contabo.md vault/runbooks/expose-service-edge2.md vault/runbooks/expose-service-home.md vault/runbooks/fleet-magicdns-resolved-migration.md vault/runbooks/headless-browser-page-verification.md vault/runbooks/headscale-oidc-boot-order.md vault/runbooks/headscale-onboard-node.md vault/runbooks/ia-cli-reference.md vault/runbooks/ia-download-mirror.md vault/runbooks/idahomesh-bridge-setup.md vault/runbooks/idahomesh-vpn-device-setup.md vault/runbooks/lxc-service-migration.md vault/runbooks/mailcow-create-mailbox.md vault/runbooks/meshai-prod-compose-override.md vault/runbooks/meshmonitor-password-reset.md vault/runbooks/meshtastic-sidecar-node.md vault/runbooks/meshtasticd-sim-nodes-runbook.md vault/runbooks/nordvpn-lxc.md vault/runbooks/peertube-remote-runner.md vault/runbooks/pg-backup.md vault/runbooks/pi-nas-omv-runbook.md vault/runbooks/pipeline-patterns.md vault/runbooks/proxmox-create-ubuntu-vm.md vault/runbooks/proxmox-onboard-node.md vault/runbooks/pymc-repeater-kiss-tnc-reenumeration.md vault/runbooks/recon-operations.md vault/runbooks/recon-service-integration.md vault/runbooks/syncthing-add-node.md vault/runbooks/toc-cortex-pve9.2-update.md vault/session-resume/SESSION-HANDOFF-meshai-test.md
This commit is contained in:
echo6-autocommit 2026-07-13 12:00:23 +00:00
commit ef8b1e0bd9
79 changed files with 448 additions and 360 deletions

View file

@ -3,9 +3,14 @@ title: Fleet Patch Audit — 2026-06-19
type: project
tags:
- proxmox
- ai
related: []
updated: 2026-06-22
aliases: []
related:
- [[fleet-platform-baseline]]
- [[lxc-service-migration]]
- [[caddy]]
- [[services]]
- [[ip-allocation]]
updated: 2026-07-13
status: complete
---
@ -13,7 +18,7 @@ status: complete
Read-only audit snapshot as of 2026-06-19. **Nothing has been applied — this is a planning document to build the patch plan from.**
**Topology note:** the old Contabo VPS has been rebuilt as **edge1 (mail-only)**; **edge2 is now the front door for everything else**. edge1 is excluded from this audit (mid-rebuild/maintenance). **Headscale:** edge2 CT107 is the main fleet tailnet (34 nodes, `vpn.echo6.co`, self-hosted Headscale 0.28.0); utility CT106 is a separate IdahoMesh sub-tailnet (`vpn.idahomesh.com`, 3 nodes, low-risk). No services route through old-Contabo. **Mailcow CT108:** destroyed 2026-06-20 (`pct destroy 108 --purge`); backup preserved durably on pi-nas (`…/contabo-prewipe-2026-06/mailcow/`, sha256-verified); live mail on edge1 (MX/A for mail.echo6.co → 5.189.158.149).
**Topology note:** the old Contabo VPS has been rebuilt as **edge1 (mail-only)**; **edge2 is now the front door for everything else**. edge1 is excluded from this audit (mid-rebuild/maintenance). **Headscale:** edge2 CT107 is the main fleet tailnet (34 nodes, `vpn.echo6.co`, self-hosted Headscale 0.28.0); utility CT106 is a separate IdahoMesh sub-tailnet (`vpn.idahomesh.com`, 3 nodes, low-risk). No [[services]] route through old-Contabo. **Mailcow CT108:** destroyed 2026-06-20 (`pct destroy 108 --purge`); backup preserved durably on pi-nas (`…/contabo-prewipe-2026-06/mailcow/`, sha256-verified); live mail on edge1 (MX/A for mail.echo6.co → 5.189.158.149).
---
@ -22,7 +27,7 @@ Read-only audit snapshot as of 2026-06-19. **Nothing has been applied — this i
**Phases 13 are fully done. The entire fleet is on the current platform.**
- **Phase 1** (guest/VM security apt) — COMPLETE 2026-06-20. 26 guests patched, ~600+ security packages cleared, zero data loss.
- **Phase 2** (app/container updates) — COMPLETE 2026-06-21. All app upgrades done: Authentik 2025.12.4→2026.5.3 (sequential), Forgejo 14→15, Headscale 0.28→0.29.1 (both instances), Immich 2.5.6→2.7.5, Nextcloud AIO→NC 33.0.5, media stack (Jellyfin/SABnzbd/arr), cortex AI stack (Ollama/TEI/Qdrant/Open-WebUI), and the low-urgency batch.
- **Phase 2** (app/container updates) — COMPLETE 2026-06-21. All app upgrades done: [[authentik]] 2025.12.4→2026.5.3 (sequential), Forgejo 14→15, Headscale 0.28→0.29.1 (both instances), Immich 2.5.6→2.7.5, Nextcloud AIO→NC 33.0.5, media stack (Jellyfin/SABnzbd/arr), cortex AI stack (Ollama/TEI/Qdrant/Open-WebUI), and the low-urgency batch.
- **Phase 3** (platform/reboot windows) — COMPLETE 2026-06-22. All 5 PVE nodes on 9.2.3/kernel 7.0.12-1-pve (including toc+cortex); pi-nas on OMV 8.4/kernel 6.18; cortex NVIDIA driver 580.167.08 + DKMS + nvidia-container-toolkit 1.19.1; GPU passthrough (vfio) survived the 7.0 kernel; cluster 5/5 quorate.
**Intentionally deferred / out of scope (not failures):**
@ -42,12 +47,12 @@ These containers have the highest raw security-update counts and have not been p
| Host | Guest | Upgradable / Security | Notes |
|------|-------|-----------------------|-------|
| utility | CT119 mesh-territory | 179 / 91 sec | Never patched |
| utility | CT108 meshai | 109 / 79 | — |
| utility | CT108 [[meshai]] | 109 / 79 | — |
| cloud | CT120 immich guest-OS | 191 / 101 | — |
| cloud | CT121 nextcloud guest-OS | 98 / 75 | — |
| media | CT110 peertube | 81 / 37 | — |
| utility | CT109 opentakserver | 34 / 31 | — |
| utility | CT104 central | 50 / 38 | Includes PostgreSQL 16.13 → 16.14 |
| utility | CT104 [[central]] | 50 / 38 | Includes PostgreSQL 16.13 → 16.14 |
### Tier 2 — App / Container Updates
@ -58,7 +63,7 @@ Updates where the application or its Docker images have drifted from current ups
| edge2 | CT105 | authentik 2025.12.4 → 2026.5.3 | **#1 security item** — 7 CVEs + 5 GHSAs in gap; sequential upgrade (min: 2025.12.6) |
| edge2 | CT107 | headscale 0.28.0 → 0.29.1 | Also a 2nd headscale on utility CT106 |
| edge2 | CT103 | forgejo 14.0.5 → 15.0.3 | **14.x EOL 2026-04-30** — migrate branch, not just patch |
| edge2 | CT106 | Synapse 1.155.0 / Element / MAS | Image drift + pending OS apt security updates |
| edge2 | CT106 | [[synapse]] 1.155.0 / Element / MAS | Image drift + pending OS apt security updates |
| edge2 | CT104 | livesync couchdb:3.4 | Docker image drift |
| edge2 | CT108 | ~~mailcow (18 containers)~~ | ✅ **Decommissioned 2026-06-20** — superseded by edge1; no longer an update target |
| cloud | CT120 | immich — server/ml/valkey:9/postgres(14-vectorchord) | 4 images drifted |
@ -98,10 +103,10 @@ Issues noted that are not package/image updates but warrant attention.
| Host / Guest | Flag | Detail |
|---|---|---|
| data | Disk 92% full | ~73 GB / 938 GB free; address before patching |
| utility CT118 archivist | rpcbind on 0.0.0.0:111 | No Tailscale client or firewall on this CT; exposed port |
| utility CT118 [[archivist]] | rpcbind on 0.0.0.0:111 | No Tailscale client or firewall on this CT; exposed port |
| media VM105 jellyseerr | Non-stable image | Running preview-OIDC tag, not a stable release |
| data VM1130 nominatim | Stale image (14 months) | nominatim:4.5, pinned; confirm intentional |
| edge2 CT106 matrix / CT107 headscale | No Tailscale client | Ingress via Caddy; verify internal routing before patching |
| edge2 CT106 matrix / CT107 headscale | No Tailscale client | Ingress via [[caddy]]; verify internal routing before patching |
---
@ -138,15 +143,15 @@ Running application version vs latest stable upstream, per app — the "is every
### Current / already past the fix (no action)
Vaultwarden 1.36.0 (edge2 CT102 — has the SSO-takeover/org-access CVE fixes) · PDM 1.1.4 (edge2 CT100 — past the RCE PSA) · WordPress 7.0 core + all plugins/themes (edge2 CT101) · Synapse 1.155.0 / Element / MAS (edge2 CT106 — current, only minor `:latest` digest drift) · obsidian-remote v1.12.7 (cortex) · PostgreSQL 16.14 (recon-vm).
Vaultwarden 1.36.0 (edge2 CT102 — has the SSO-takeover/org-access CVE fixes) · PDM 1.1.4 (edge2 CT100 — past the RCE PSA) · WordPress 7.0 core + all plugins/[[themes]] (edge2 CT101) · Synapse 1.155.0 / Element / MAS (edge2 CT106 — current, only minor `:latest` digest drift) · obsidian-remote v1.12.7 (cortex) · PostgreSQL 16.14 (recon-vm).
### Lower urgency
Mumble 1.5.517→1.5.901 · Caddy 2.10.2/2.11.3→2.11.4 · Qdrant 1.16.3→1.18.2 · TEI 1.7.4→1.9.3 · Valhalla 3.6.3→3.7.0 · Photon 1.1.0→1.2.0 · kiwix 3.7.0→3.8.2 · CouchDB 3.4.3→3.5.2 (livesync) · Navidrome 0.60.3→0.62.0 · Sonarr/Radarr/Prowlarr/Lidarr 12 versions · NATS 2.14.0→2.14.2 · PostgreSQL 16.12/16.13→16.14 · meshmonitor (~1 mo, exact ver undeterminable) · searxng (rolling, ~4.5 mo) + valkey-8 sidecar 8.1.5→8.1.8 · mautrix-signal v0.2603.0.
Mumble 1.5.517→1.5.901 · Caddy 2.10.2/2.11.3→2.11.4 · Qdrant 1.16.3→1.18.2 · TEI 1.7.4→1.9.3 · Valhalla 3.6.3→3.7.0 · Photon 1.1.0→1.2.0 · kiwix 3.7.0→3.8.2 · CouchDB 3.4.3→3.5.2 (livesync) · Navidrome 0.60.3→0.62.0 · Sonarr/Radarr/Prowlarr/Lidarr 12 versions · NATS 2.14.0→2.14.2 · PostgreSQL 16.12/16.13→16.14 · meshmonitor (~1 mo, exact ver undeterminable) · [[searxng]] (rolling, ~4.5 mo) + valkey-8 sidecar 8.1.5→8.1.8 · [[mautrix_signal]] v0.2603.0.
### Internal echo6 apps (no upstream to track)
central-*, meshai, archivist, meshwars, recon / recon-watchdog, navi-* — running; version = current git head.
central-*, meshai, archivist, meshwars, [[recon]] / recon-watchdog, navi-* — running; version = current git head.
---
@ -240,7 +245,7 @@ Lowest-risk changes first; everything reboot-bearing deferred to scheduled windo
### Incidental fixes made during Phase 1
- **(a) CT110 peertube — immutable `/etc/resolv.conf` blocked reboot.** The file had `chattr +i` set (intentional NordVPN DNS protection). Cleared the immutable flag to allow the reboot, then verified the flag was restored and DNS remained healthy after boot.
- **(a) CT110 peertube — immutable `/etc/resolv.conf` blocked reboot.** The file had `chattr +i` set (intentional NordVPN [[dns]] protection). Cleared the immutable flag to allow the reboot, then verified the flag was restored and DNS remained healthy after boot.
- **(b) CT111 mcc — DNS hijacked to unreachable MagicDNS.** Tailscale `accept-dns` was redirecting DNS to a MagicDNS address that was not reachable from this CT. Disabled `tailscale accept-dns`, set `1.1.1.1` / `8.8.8.8` persistently.
- **(c) PostgreSQL on central CT104 moved 16.13→16.14** as part of the security-pocket apt pass.
- **(d) Fleet-wide stale `/etc/hosts` fix** — see Critical Finding 2.
@ -268,7 +273,7 @@ Known minor leftover (cosmetic, non-blocking): CT107's OWN tailscale client node
### 🔴 Critical Finding 2 — fleet-wide stale /etc/hosts broke coordinator connectivity
7 fleet nodes — data, cloud, media, utility hosts, plus caddy CT101, cobalt CT112, and peertube CT110 — had a stale `5.189.158.149 vpn.echo6.co` line in `/etc/hosts` left over from before the 2026-06-19 headscale migration to edge2. This pinned `vpn.echo6.co` to edge1 (now mail-only), so tailscaled hit Mailcow's TLS cert and could never reach the real coordinator — affected nodes showed OFFLINE in headscale while coasting on persistent WireGuard tunnels (still SSH-reachable, masking the problem).
7 fleet nodes — data, cloud, media, utility hosts, plus caddy CT101, cobalt CT112, and peertube CT110 — had a stale `5.189.158.149 vpn.echo6.co` line in `/etc/hosts` left over from before the [[2026-06-19]] headscale migration to edge2. This pinned `vpn.echo6.co` to edge1 (now mail-only), so tailscaled hit Mailcow's TLS cert and could never reach the real coordinator — affected nodes showed OFFLINE in headscale while coasting on persistent WireGuard tunnels (still SSH-reachable, masking the problem).
**Fixed 2026-06-20:** removed the stale line and ran `tailscale up` on all 7; all confirmed ONLINE in the coordinator. `/etc/hosts.bak-20260620` backups left on each host.
@ -331,7 +336,7 @@ Empirically confirmed during the OTS update: updating OTS to 1.7.12 does **not**
### Incidental fixes and side-work during Phase 2
- **CT107 boot-survival fix** (applied earlier in the effort, during Phase 1 resolution) — rebound headscale/headplane ports to 10.10.10.25, dropped `tailscale-online.target` dependency, disabled `only_start_if_oidc_is_available` gate, repointed edge2 Caddy; proven by reboot self-heal in ~45 s. Also corrected CT107's own tailscale node ControlURL to `vpn.echo6.co` so it self-registers cleanly.
- **Utility node incident (resolved):** a batch delete of 9 LVM-thin snapshots triggered an SSD TRIM/discard storm that spiked I/O and load transiently; compounded by CT103 argus running hot (transcription + docker-compose build churn). Matt migrated argus to the cloud node, resolving the issue; utility load returned to normal. **LESSON: delete thin-pool snapshots one at a time — not in a batch — to avoid the discard storm.**
- **Utility node incident (resolved):** a batch delete of 9 LVM-thin snapshots triggered an SSD TRIM/discard storm that spiked I/O and load transiently; compounded by CT103 [[argus]] running hot (transcription + docker-compose build churn). Matt migrated argus to the cloud node, resolving the issue; utility load returned to normal. **LESSON: delete thin-pool snapshots one at a time — not in a batch — to avoid the discard storm.**
- **Nextcloud:** granted `matt@echo6.co` the NC admin role. (user_oidc has no group-claim sync, so this is durable across SSO logins.)
- **Radarr:** set up a `\\192.168.1.160\manual` SMB drop folder on the same NFS export as the library (atomic-move imports) for manual movie filing.
- **Snapshot hygiene:** all rollback snapshots cleaned up after validation — Phase 1 `presec-*`, Phase 2 `prewave2-*`, OTS `pre-ots-*` snapshots all removed.
@ -362,7 +367,7 @@ Empirically confirmed during the OTS update: updating OTS to 1.7.12 does **not**
### Separate deferred projects
- Nominatim v5 re-import — see [[nominatim-v5-reimport]]
- [[Nominatim v5 Re-import]] — see [[nominatim-v5-reimport]]
---