diff --git a/engine/changelog.md b/engine/changelog.md index 249b8f6..7424996 100644 --- a/engine/changelog.md +++ b/engine/changelog.md @@ -119,3 +119,5 @@ ## 2026-06-19T09:00:02Z — sweep deferred (competing GPU process: 2932 node /usr/bin/peertube-runner server --enable-job vod-hls-transcoding --enable-job vod-audio-merge-transcoding --enable-job live-rtmp-hls-transcoding --enable-job video-studio-transcoding --enable-job video-transcription) ## 2026-06-20T09:00:01Z — sweep deferred (competing GPU process: 2932 node /usr/bin/peertube-runner server --enable-job vod-hls-transcoding --enable-job vod-audio-merge-transcoding --enable-job live-rtmp-hls-transcoding --enable-job video-studio-transcoding --enable-job video-transcription) + +## 2026-06-21T09:00:01Z — sweep deferred (competing GPU process: 2932 node /usr/bin/peertube-runner server --enable-job vod-hls-transcoding --enable-job vod-audio-merge-transcoding --enable-job live-rtmp-hls-transcoding --enable-job video-studio-transcoding --enable-job video-transcription) diff --git a/engine/lint-report.md b/engine/lint-report.md index 4c50299..5e6c127 100644 --- a/engine/lint-report.md +++ b/engine/lint-report.md @@ -1,6 +1,6 @@ # Vault Lint Report -Generated: 2026-06-21T00:00:06Z | Docs scanned: 92 | Elapsed: 0.0s +Generated: 2026-06-21T06:00:05Z | Docs scanned: 92 | Elapsed: 0.0s ## Summary diff --git a/vault/.obsidian/workspace.json b/vault/.obsidian/workspace.json index 66870b8..f4586f1 100644 --- a/vault/.obsidian/workspace.json +++ b/vault/.obsidian/workspace.json @@ -199,17 +199,17 @@ }, "active": "8d53cdb6c257e685", "lastOpenFiles": [ + "projects/fleet-patch-audit.md.tmp.1493418.2fea79b9eb48", + "projects/fleet-patch-audit.md.tmp.1493418.f0355595161e", + "projects/fleet-patch-audit.md.tmp.1493418.00879ee0f4e5", + "projects/meshtastic-headscale-runbook.md.tmp.1493418.e92ea881edf6", + "projects/meshtastic-headscale-runbook.md.tmp.1493418.ac14d8b4b039", "projects/fleet-patch-audit.md.tmp.1493418.23b2caf55e12", "projects/fleet-patch-audit.md.tmp.1493418.f30558958079", "projects/fleet-patch-audit.md.tmp.1493418.5fe5d2d39856", "projects/fleet-patch-audit.md.tmp.1493418.6c2a45ea0ecd", "projects/fleet-patch-audit.md.tmp.1493418.9a000c8ac5aa", "projects/fleet-patch-audit.md.tmp.1493418.c361bd10be1b", - "projects/fleet-patch-audit.md.tmp.1493418.83a57813be85", - "projects/fleet-patch-audit.md.tmp.1493418.c4ba031d7df6", - "projects/fleet-patch-audit.md.tmp.1493418.d191aac142fc", - "projects/fleet-patch-audit.md.tmp.1493418.6777b1d5dff6", - "projects/fleet-patch-audit.md.tmp.1493418.731843387669", "projects/nominatim-v5-reimport.md", "2026-06-19.md", "Untitled.canvas", diff --git a/vault/projects/fleet-patch-audit.md b/vault/projects/fleet-patch-audit.md index 85e6ec5..7856897 100644 --- a/vault/projects/fleet-patch-audit.md +++ b/vault/projects/fleet-patch-audit.md @@ -5,7 +5,7 @@ tags: - proxmox - ai related: [] -updated: 2026-06-20 +updated: 2026-06-21 status: active --- @@ -150,7 +150,7 @@ Lowest-risk changes first; everything reboot-bearing deferred to scheduled windo |------|-------|---------|-------| | **1 — Guest/VM security apt** | utility CT100,101,102,103,104,106,107,108,109,112,118,119 · cloud CT120,121 · media VM105,CT110,CT111 · data VM1130 · edge2 CT100–107 | No | Lowest blast radius. Worst-first: CT119, CT108, immich/nextcloud guest-OS, peertube. | | **2 — Hypervisor host OS security** | data, utility, cloud, media host OSes (**not toc**) | No | One node at a time; restarts hypervisor-side daemons (smbd etc.). edge2 host already patched. | -| **3 — App / container updates (Tier 2)** | per-app, native updater each | Per-app | See "special handling" below — not a generic `docker pull`. | +| **3 — App / container updates (Tier 2)** | per-app, native updater each | Per-app | **Substantially complete 2026-06-21** — all security-critical apps done; low-urgency batch + cortex AI stack remain. See Phase 2 Execution Log and "special handling" below. | | **4 — Reboot windows (Tier 3)** | PVE 9.2 + QEMU 11 + LXC 7 + kernel on the 5 PVE-9 nodes; pi-nas kernel + OMV; cortex NVIDIA/DKMS | **Yes** | Schedule deliberately; toc+cortex coordinated. | | **Cross-cutting** | Tailscale 1.94→1.98 fleet-wide | No | Can ride along Phase 1/2. | @@ -262,6 +262,82 @@ Headscale stale-node cleanup: deleted dead nodes `mailcow` (destroyed CT108) and --- +## Phase 2 — Execution Log (2026-06-21) + +**Status: substantially complete; paused 2026-06-21.** All high-priority and security-critical app upgrades are done. Only the low-urgency batch items and the cortex AI stack (protected host, own window) remain. + +### Completed upgrades + +All targets below were validated healthy post-upgrade; rollback DB dumps retained on the respective CTs under `/root`. + +**Wave 1 — cloud apps** + +| App | Host | From → To | Notes | +|-----|------|-----------|-------| +| Immich | cloud CT120 | 2.5.6 → 2.7.5 | XSS + ACL-bypass fixes; stack includes Valkey 9 and vectorchord PG | +| PeerTube | media CT110 | 8.0.2 → 8.2.1 | — | +| Nextcloud AIO | cloud CT121 | NC 32.0.4 → **33.0.5** (major); mastercontainer updated | Major NC version; AIO orchestrated update path | + +**Wave 2 — media stack (media VM105)** + +| App | From → To | Notes | +|-----|-----------|-------| +| Jellyfin | 10.11.6 → 10.11.11 | — | +| Sonarr | → 4.0.17 | — | +| Radarr | → 6.2.1 | — | +| Prowlarr | → 2.4.0 | — | +| Navidrome | → 0.62.0 | — | +| SABnzbd | 4.5.5 → **5.0.4** (major) | Major upgrade; clean migration | +| Lidarr | DEFERRED | Image maintainer hasn't shipped v3 — needs image swap to linuxserver to go v3; left on current | +| Jellyseerr | INTENTIONALLY HELD on `preview-OIDC` tag | Stable lacks OIDC support; switching would break SSO login | + +**Authentik (edge2 CT105) 2025.12.4 → 2026.5.3** + +Done as 3 sequential hops: 12.4 → 12.6 → 2026.2.4 → 2026.5.3. `pg_dump` before each hop; migrations clean each time. **SSO verified end-to-end by Matt logging into navi.** Closes the May-2026 CVE waves. (ak app version 5.2.15.) + +**Forgejo (edge2 CT103) 14.0.5 → 15.0.3** + +EOL-branch migration; v15 schema migrations applied cleanly; web 200, API reports 15.0.3. Pre-15 DB dump at `/root/forgejo-db-pre15-20260621.sql`. + +**Headscale 0.28 → 0.29.1 (both instances)** + +- **CT106** (utility, IdahoMesh 3-node mesh) — upgraded first as rehearsal; native systemd binary. +- **CT107** (edge2, fleet coordinator, Docker) — 32/32 nodes reconnected post-upgrade; health 200; boot-survival fix preserved (ports remain on 10.10.10.25). 0.29 breaking changes documented in [[meshtastic-headscale-runbook]] (key changes: `randomize_client_port` removed — was a hard blocker; ephemeral key config nested; minimum Tailscale client 1.80.0; bare ACL `*` now tailnet-only). + +**OpenTAKServer (utility CT109) 1.7.10 → 1.7.12** + +Clean; all 9 TAK services healthy. + +### Key decision — RabbitMQ LEFT on 3.12.1 (EOL) + +Empirically confirmed during the OTS update: updating OTS to 1.7.12 does **not** touch RabbitMQ — OTS 1.7.12 runs fine against RabbitMQ 3.12.1/Erlang OTP 25. Forcing RabbitMQ → 4.x is the wrong move: OTS is built and tested against 3.12, and 4.x has breaking changes (queue-mirroring removal, feature-flag requirements) that would likely break OTS, which does not appear to support 4.x. AMQP/MQTT ports are localhost-bound, so exposure is low. The EOL 3.12 is a low-priority latent risk to revisit only if/when OTS officially supports RabbitMQ 4.x. **NOT a current action item.** + +### Incidental fixes and side-work during Phase 2 + +- **CT107 boot-survival fix** (applied earlier in the effort, during Phase 1 resolution) — rebound headscale/headplane ports to 10.10.10.25, dropped `tailscale-online.target` dependency, disabled `only_start_if_oidc_is_available` gate, repointed edge2 Caddy; proven by reboot self-heal in ~45 s. Also corrected CT107's own tailscale node ControlURL to `vpn.echo6.co` so it self-registers cleanly. +- **Utility node incident (resolved):** a batch delete of 9 LVM-thin snapshots triggered an SSD TRIM/discard storm that spiked I/O and load transiently; compounded by CT103 argus running hot (transcription + docker-compose build churn). Matt migrated argus to the cloud node, resolving the issue; utility load returned to normal. **LESSON: delete thin-pool snapshots one at a time — not in a batch — to avoid the discard storm.** +- **Nextcloud:** granted `matt@echo6.co` the NC admin role. (user_oidc has no group-claim sync, so this is durable across SSO logins.) +- **Radarr:** set up a `\\192.168.1.160\manual` SMB drop folder on the same NFS export as the library (atomic-move imports) for manual movie filing. +- **Snapshot hygiene:** all rollback snapshots cleaned up after validation — Phase 1 `presec-*`, Phase 2 `prewave2-*`, OTS `pre-ots-*` snapshots all removed. + +### Remaining work (next session) + +**Low-urgency batch:** +- Caddy (edge2 CT101) +- CouchDB 3.4 → 3.5 (livesync CT104 edge2) +- valkey-8 sidecar (searxng CT102) +- NATS 2.14.0 → 2.14.2 (central CT104 utility) +- MediaMTX 1.13 → 1.19 + Mumble on CT109 (OTS companions, separate from the OTS app) + +**Protected host — own deliberate window:** +- cortex AI stack: ollama 0.16 → 0.30, open-webui, qdrant, TEI + +**Separate deferred projects:** +- Nominatim v5 re-import — see [[nominatim-v5-reimport]] +- Optional: clean up retained `/root` DB dumps (authentik hop1/2/3, forgejo, headscale, OTS, immich, peertube, etc.) once upgrades are confirmed stable + +--- + ## Execution Runbook (Meticulous) Step-by-step plan to bring every application and package current. Per-app target versions live in **Application Currency** above; this section is the *how/when/order*. Execution model: scoped Sonnet task per target → Opus verifies output → report to Matt; **approval gate before every load-bearing or reboot step.** diff --git a/vault/projects/meshtastic-headscale-runbook.md b/vault/projects/meshtastic-headscale-runbook.md index 1a048cc..a6a1c80 100644 --- a/vault/projects/meshtastic-headscale-runbook.md +++ b/vault/projects/meshtastic-headscale-runbook.md @@ -10,7 +10,7 @@ related: - [[meshtastic-sidecar-node]] - [[headscale-onboard-node]] - [[caddy]] -updated: 2026-06-18 +updated: 2026-06-21 --- # IdahoMesh Tailnet Runbook @@ -572,4 +572,96 @@ ping 100.64.0.14 # cortex — should timeout/unreachable --- -*Last updated: 2026-02-11* +## Headscale 0.28 → 0.29 Upgrade Guide + +**Rehearsed on CT 106 (meshtastic-hs) 2026-06-21. Result: success. Use this as the playbook for CT 107 (mesh-bridge / Echo6 Headscale).** + +### Breaking changes that require action before starting + +#### 1. `randomize_client_port` removed from config.yaml (HARD BLOCKER) + +Headscale 0.29 **refuses to start** if `randomize_client_port` exists in `config.yaml`. Remove the line: + +```bash +sed -i '/^randomize_client_port:/d' /etc/headscale/config.yaml +``` + +If the feature is needed, move it to the ACL policy file (`/etc/headscale/acl.json`) as a top-level key: +```json +{ "randomizeClientPort": true } +``` + +Default is `false`; simply removing the key is equivalent to `randomize_client_port: false`. + +#### 2. `ephemeral_node_inactivity_timeout` removed from top-level config (deprecation → removed) + +The old top-level key is gone in 0.29. Replace with the new nested path: + +```yaml +# OLD (remove this): +ephemeral_node_inactivity_timeout: 30m + +# NEW (replace with): +node: + ephemeral: + inactivity_timeout: 30m +``` + +0.29 accepts the old key with a warning and uses the correct value, but the warning appears on every startup. Fix it to keep logs clean. + +#### 3. Strict version upgrade path enforced + +0.29 blocks skipping minor versions. If upgrading from 0.27 or earlier, you MUST go 0.27 → 0.28 → 0.29 one step at a time. We were on 0.28, so this is not a blocker for this fleet. + +#### 4. ACL wildcard `*` behavior changed (review your policy) + +In 0.29, bare `*` in ACL `src`/`dst` resolves to the CGNAT+ULA range (tailnet nodes only) instead of all IPs. If you had `"dst": ["*:*"]` meaning "any internet IP," you must switch to `autogroup:danger-all` as a source or explicit CIDRs. **The IdahoMesh ACL does not use bare `*` — no action needed for this tailnet.** + +#### 5. Minimum Tailscale client version raised to v1.80.0 + +Any tailscale client older than v1.80.0 will be rejected by 0.29. Verify all registered nodes are running a sufficiently recent tailscale before upgrading. + +### Other notable 0.29 changes (no action required, FYI) + +- GivenName collision suffix changed from random hash (`laptop-abc12xyz`) to monotonic numeric (`laptop`, `laptop-1`). MagicDNS names for nodes with old random suffixes will change on upgrade. The raw Hostname column is unchanged. +- `oidc.expiry` removed; use `node.expiry` for all registration methods. +- `headscale nodes register` deprecated in favour of `headscale auth register --auth-id`. +- `--namespace` flag removed from nodes commands (was already replaced by `--user`). +- New: `trusted_proxies` config option for True-Client-IP headers (previously honoured from any client). +- DB migration is automatic on first start — no manual migration command needed. + +### Rollback procedure (if upgrade fails) + +1. Stop the service: `systemctl stop headscale` +2. Reinstall the 0.28.0 binary: `cp /usr/local/bin/headscale.bak-v0.28.0-pre029 /usr/local/bin/headscale` +3. Restore config: `cp /etc/headscale/config.yaml.bak-pre029-20260621 /etc/headscale/config.yaml` +4. Restore DB: `cp /root/headscale-db-pre029-20260621.sqlite /var/lib/headscale/db.sqlite` +5. Start: `systemctl start headscale` + +> **Note:** 0.29 blocks downgrading once the DB migration has run. The DB backup from step 4 predates the migration, so this rollback is clean. + +### CT 106 (meshtastic-hs) upgrade record + +- **Date:** 2026-06-21 +- **From:** v0.28.0 → **To:** v0.29.1 +- **Deployment:** native systemd binary at `/usr/local/bin/headscale` +- **Config changes made:** + - Removed `randomize_client_port: false` + - Replaced `ephemeral_node_inactivity_timeout: 30m` with `node.ephemeral.inactivity_timeout: 30m` +- **DB migration:** automatic, clean, no errors in logs +- **Nodes after upgrade:** all 3 registered (burley-butte offline, mesh-bridge online, mt-isr offline) — same as before +- **Backups (pre-upgrade):** `/etc/headscale/config.yaml.bak-pre029-20260621`, `/root/headscale-db-pre029-20260621.sqlite`, `/usr/local/bin/headscale.bak-v0.28.0-pre029` + +### CT 107 (mesh-bridge / Echo6 Headscale) upgrade checklist + +CT 107 runs Docker (not native binary). Steps will differ: + +1. Back up: `docker cp headscale-vanilla:/var/lib/headscale/db.sqlite /root/headscale-db-pre029-$(date +%Y%m%d).sqlite` and `cp /path/to/config.yaml config.yaml.bak-pre029` +2. Apply the two config changes above (remove `randomize_client_port`, migrate ephemeral timeout key) +3. Change the image tag in `docker-compose.yml` from `headscale/headscale:0.28.x` to `headscale/headscale:0.29.1` +4. `docker compose up -d` +5. Validate: `docker exec headscale-vanilla headscale version`, check `docker logs headscale-vanilla` for migration success, `docker exec headscale-vanilla headscale nodes list` + +--- + +*Last updated: 2026-06-21*