auto: docs sync 2026-06-22T06:00:09+00:00
Files changed: engine/lint-report.md vault/.obsidian/workspace.json vault/projects/fleet-patch-audit.md vault/runbooks/toc-cortex-pve9.2-update.md
This commit is contained in:
parent
277ac1194b
commit
cb50bce39a
4 changed files with 108 additions and 19 deletions
|
|
@ -5,7 +5,7 @@ tags:
|
|||
- proxmox
|
||||
- ai
|
||||
related: []
|
||||
updated: 2026-06-21
|
||||
updated: 2026-06-22
|
||||
status: active
|
||||
---
|
||||
|
||||
|
|
@ -63,8 +63,8 @@ Reboot-bearing updates. Coordinate maintenance windows carefully. cortex and toc
|
|||
| edge2 | PVE 8.4.19 | Already fully patched — no action needed |
|
||||
| cortex VM150 (PROTECTED) | NVIDIA driver 580.159 → 580.167 + DKMS | Reboot required |
|
||||
| cortex VM150 (PROTECTED) | nvidia-container-toolkit 1.18 → 1.19 | — |
|
||||
| pi-nas | Kernel 6.12 → 6.18 | Reboot required |
|
||||
| pi-nas | OMV 8.1 → 8.4 | — |
|
||||
| pi-nas | Kernel 6.12.62 → 6.18.34+rpt-rpi-2712 | ✅ DONE 2026-06-22 — operator present; `/boot` backup at `/root/boot-backup-pre6.18-20260622.tar.gz`; booted clean; rpi-eeprom held |
|
||||
| pi-nas | OMV 8.1 → 8.4 | ✅ DONE 2026-06-21 — 8.4.0-3 |
|
||||
|
||||
### Cross-Cutting (All / Most Hosts)
|
||||
|
||||
|
|
@ -151,7 +151,7 @@ Lowest-risk changes first; everything reboot-bearing deferred to scheduled windo
|
|||
| **1 — Guest/VM security apt** | utility CT100,101,102,103,104,106,107,108,109,112,118,119 · cloud CT120,121 · media VM105,CT110,CT111 · data VM1130 · edge2 CT100–107 | No | Lowest blast radius. Worst-first: CT119, CT108, immich/nextcloud guest-OS, peertube. |
|
||||
| **2 — Hypervisor host OS security** | data, utility, cloud, media host OSes (**not toc**) | No | One node at a time; restarts hypervisor-side daemons (smbd etc.). edge2 host already patched. |
|
||||
| **3 — App / container updates (Tier 2)** | per-app, native updater each | Per-app | **COMPLETE 2026-06-21** — all app upgrades done (security-critical, low-urgency batch, and cortex AI stack). See Phase 2 Execution Log. |
|
||||
| **4 — Reboot windows (Tier 3)** | PVE 9.2 + QEMU 11 + LXC 7 + kernel on the 5 PVE-9 nodes; pi-nas kernel + OMV; cortex NVIDIA/DKMS | **Yes** | Schedule deliberately; toc+cortex coordinated. |
|
||||
| **4 — Reboot windows (Tier 3)** | PVE 9.2 + QEMU 11 + LXC 7 + kernel on the 5 PVE-9 nodes; pi-nas kernel + OMV; cortex NVIDIA/DKMS | **Yes** | **IN PROGRESS (2026-06-21/22)** — 4-node cluster (media/data/utility/cloud) ✅ + pi-nas fully ✅ (OMV 8.4 + kernel 6.18); only toc+cortex remains (matt-desktop). See Phase 3 Execution Log. |
|
||||
| **Cross-cutting** | Tailscale 1.94→1.98 fleet-wide | No | Can ride along Phase 1/2. |
|
||||
|
||||
**Special handling — do NOT bulk-patch these; use the native updater**
|
||||
|
|
@ -350,6 +350,39 @@ Empirically confirmed during the OTS update: updating OTS to 1.7.12 does **not**
|
|||
|
||||
---
|
||||
|
||||
## Phase 3 — Execution Log (2026-06-21/22)
|
||||
|
||||
**Status: IN PROGRESS — 4-node cluster ✅ + pi-nas fully ✅ (OMV 8.4 + kernel 6.18); only toc+cortex remains (handed to matt-desktop).**
|
||||
|
||||
### Pre-flight findings
|
||||
|
||||
- Cluster `echo6-cluster` 5/5 quorate, quorum 3, NO HA configured (clean guest stop/start, no fencing).
|
||||
- BLOCKER found + fixed: data, cloud, media had NO Proxmox APT repo configured at all — added `pve-no-subscription` (trixie, matching utility/toc); all 5 then saw the 9.2 stack.
|
||||
- Note: PVE 9.2 ships **kernel 7.0 as the new default** (proxmox-default-kernel) — nodes boot 7.0.12-1-pve, not 6.17.13 as the audit predicted.
|
||||
|
||||
### Cluster nodes — ALL upgraded to PVE 9.2.3 / kernel 7.0.12-1-pve (QEMU 11, LXC 7)
|
||||
|
||||
One at a time; cluster stayed 5/5 quorate throughout.
|
||||
|
||||
- **media** ✅ — done first (canary). Incidental: CT110 peertube failed to auto-start (recurring `/etc/resolv.conf` immutable-flag vs LXC pre-start-hook conflict) → **permanently fixed**: the flag was an obsolete workaround (NordVPN `set dns` no longer overwrites resolv.conf), cleared it + enabled NordVPN auto-connect, so future reboots won't trip it.
|
||||
- **data** ✅ — recon-vm (VM1130) healthy; virtiofsd-{kiwix,library,nav} auto-recovered this time; nginx needed the one expected restart (pre-existing mesh-DNS startup race).
|
||||
- **utility** ✅ — all 11 CTs; central JetStream intact (12 streams); caddy proxying, mesh headscale (CT106, 3 nodes), mesh-bridge (CT107) both tailnets, OTS stack all healthy. CT102 searxng needed a manual `pct start` (transient auto-start miss, no persistent fault).
|
||||
- **cloud** ✅ — done last (per operator). immich (4 containers + API 200), nextcloud (12 AIO containers, occ healthy, v33.0.5), argus running. (Note: argus CT103 on cloud has argus-capture missing / argus-transcribe masked — operator's in-progress argus→cloud migration, not from the upgrade.)
|
||||
|
||||
### pi-nas ✅ FULLY DONE (OMV 8.4 + kernel 6.18)
|
||||
|
||||
**OMV (2026-06-21):** OMV 8.1.0-2 → 8.4.0-3 + Debian userspace (Docker, OpenSSL, salt, tailscale). All 5 NFS exports (arr/immich/nextcloud/peertube/data) serving; immich+nextcloud mounts confirmed OK.
|
||||
|
||||
**Kernel jump (2026-06-22):** 6.12.62 → 6.18.34+rpt-rpi-2712 — operator physically present. Pre-jump: full `/boot` backup at `/root/boot-backup-pre6.18-20260622.tar.gz` (154 MB) + old kernel left in place as fallback. Booted clean in ~15 s; all 5 NFS exports serving; immich+nextcloud mounts recovered; SD card healthy. Only package remaining HELD: `rpi-eeprom` (SPI bootloader — intentionally untouched; optional to update later).
|
||||
|
||||
**Important architecture note recorded:** pi-nas (2.8 TB RAID1) is the NFS storage backend for immich's photo library (644 GB), nextcloud files, peertube, and the *arr library — so a pi-nas reboot stalls those services' storage.
|
||||
|
||||
### Remaining (handed off)
|
||||
|
||||
- **toc + cortex** — deliberately NOT done by the cortex-based automation (rebooting toc drops cortex). Handed to a **matt-desktop** Claude Code session (Windows, now has Claude Code 2.1.185). Runbook + ready-to-paste prompt saved at `runbooks/toc-cortex-pve9.2-update.md`. Covers: cortex NVIDIA driver 580.159→580.167 + DKMS + container-toolkit + reboot, then toc PVE 9.2 + reboot. **Highest-risk check: GPU passthrough (vfio) surviving toc's new kernel 7.0.** Safe to run now (other 4 nodes up/quorate → toc reboot keeps 4/5).
|
||||
|
||||
---
|
||||
|
||||
## Execution Runbook (Meticulous)
|
||||
|
||||
Step-by-step plan to bring every application and package current. Per-app target versions live in **Application Currency** above; this section is the *how/when/order*. Execution model: scoped Sonnet task per target → Opus verifies output → report to Matt; **approval gate before every load-bearing or reboot step.**
|
||||
|
|
@ -576,7 +609,7 @@ Vaultwarden, PDM, WordPress, Synapse/Element/MAS, obsidian, recon-vm Postgres
|
|||
| 3 | cortex NVIDIA driver 580.159→580.167 + DKMS (reboot) | Driver + DKMS as discrete reversible step BEFORE toc reboot | `qm snapshot 150 pre-nvidia` + vzdump | `apt --only-upgrade nvidia-driver-580 nvidia-dkms-580`; DKMS rebuild; verify BEFORE toc reboot | `nvidia-smi`=580.167; `dkms status` installed; containers see GPU | restore snapshot; reinstall 580.159 + DKMS | Keep old driver pkg; after toolkit |
|
||||
| 3 | toc host PVE 9.2 + QEMU11 + LXC7 + kernel (reboot — LAST, drops cortex) | Platform + reboot; only guest is VM150 | vzdump/snapshot VM150 + record `pveversion` | snapshot VM150 → `dist-upgrade` toc → reboot → VM150 returns → cortex driver reboot | `pveversion`=9.2; toc rejoins (5/5 votes); VM150 + GPU healthy | restore VM150 from vzdump; GRUB prior kernel | LAST + alone; alternate control point TESTED; all 4 other nodes quorate |
|
||||
| 3 | cortex final stack health verification | Post-window full AI-stack + engine check | keep pre-window snaps until verified | verification only | driver 580.167; all 5 containers; ollama/tei/qdrant/open-webui; vault-tagger+TEI engines ok; TS online; Docker 29.6 | restore degraded component; worst case VM150 vzdump | LAST action of node group |
|
||||
| 3 | pi-nas OMV 8.1→8.4 (reboot #1, separate window) | OMV native update path (NOT raw apt) → reboot | off-box `config.xml` export + plugin/layout record | OMV UI Update Mgmt or `omv-upgrade`; reboot | OMV=8.4; shares/SMB/NFS/RAID/mergerfs healthy; client mounts | restore config.xml + `dpkg --set-selections` 8.1 + `omv-salt deploy` | Physical access; not mid-sync; after Phase 1 |
|
||||
| 3 | ~~pi-nas OMV 8.1→8.4 (reboot #1, separate window)~~ | ✅ **DONE 2026-06-21** — `apt full-upgrade` (kernel/firmware 6 pkgs held); OMV 8.4.0-3 installed; omv-salt deploy ran (14/0 OK); initramfs regenerated for 6.12.62 only; rebooted; uname -r still 6.12.62; all 5 NFS exports serving; immich-nfs-OK, nextcloud-nfs-OK; nfs-kernel-server + docker active; `degraded` = pre-existing quotaon "File exists" false alarm, unrelated to upgrade | N/A | N/A | N/A | Complete |
|
||||
| 3 | pi-nas kernel 6.12→6.18 (reboot #2, separate window) | Kernel bump → reboot; activates deferred security | back up `/boot`+`/boot/firmware`; keep 6.12 installed; fresh config.xml | `apt install linux-image-arm64` (unhold); update bootloader; reboot ON APPROVAL | `uname -r`=6.18; LAN+tailnet up; shares mount; Docker up | boot retained 6.12; restore `/boot`; on-site SD reflash | After OMV reboot verified; physical access |
|
||||
| 3 | edge2 host PVE 8.4 reboot — CONDITIONAL (CT107 boot-survival blocker CLEARED 2026-06-20) | ~~Previously blocked on CT107 headscale boot-survival~~ — cleared (self-heals in ~45s). Now gated only on Phase 0 kernel-CVE decision (Open Dec #5 = no reboot needed — host fully patched). If ever required: own no-failover window AFTER CT105/CT107 verified | vzdump all guests + `pveversion` | scoped kernel/security upgrade → reboot in approved window | `uname -r`=fixed; all CTs return; Caddy-fronted services reachable | GRUB prior kernel; vzdump restore | GATED on Phase 0; announce SSO+tailnet downtime; non-tailnet path |
|
||||
| — | Nominatim 4.5→5.3 + Photon 1.1→1.2 (SEPARATE PROJECT) | Full re-import into separate DB/instance; keep 4.5 until validated; Photon AFTER nominatim validated | `qm snapshot 1130` + retain 4.5 image/data | import to new DB; size flat-nodes vs free space (NOT on data); cut over after verify | 5.x geocode correct; Photon 1.2 reads v5 | re-pin nominatim:4.5 + old data | Wholly separate from patch campaign; after disk remediation |
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue