auto: docs sync 2026-06-21T12:00:06+00:00

Files changed: engine/changelog.md engine/lint-report.md vault/.obsidian/workspace.json vault/projects/fleet-patch-audit.md vault/projects/meshtastic-headscale-runbook.md
This commit is contained in:
echo6-autocommit 2026-06-21 12:00:06 +00:00
commit 7484d2f162
5 changed files with 180 additions and 10 deletions

View file

@ -5,7 +5,7 @@ tags:
- proxmox
- ai
related: []
updated: 2026-06-20
updated: 2026-06-21
status: active
---
@ -150,7 +150,7 @@ Lowest-risk changes first; everything reboot-bearing deferred to scheduled windo
|------|-------|---------|-------|
| **1 — Guest/VM security apt** | utility CT100,101,102,103,104,106,107,108,109,112,118,119 · cloud CT120,121 · media VM105,CT110,CT111 · data VM1130 · edge2 CT100107 | No | Lowest blast radius. Worst-first: CT119, CT108, immich/nextcloud guest-OS, peertube. |
| **2 — Hypervisor host OS security** | data, utility, cloud, media host OSes (**not toc**) | No | One node at a time; restarts hypervisor-side daemons (smbd etc.). edge2 host already patched. |
| **3 — App / container updates (Tier 2)** | per-app, native updater each | Per-app | See "special handling" below — not a generic `docker pull`. |
| **3 — App / container updates (Tier 2)** | per-app, native updater each | Per-app | **Substantially complete 2026-06-21** — all security-critical apps done; low-urgency batch + cortex AI stack remain. See Phase 2 Execution Log and "special handling" below. |
| **4 — Reboot windows (Tier 3)** | PVE 9.2 + QEMU 11 + LXC 7 + kernel on the 5 PVE-9 nodes; pi-nas kernel + OMV; cortex NVIDIA/DKMS | **Yes** | Schedule deliberately; toc+cortex coordinated. |
| **Cross-cutting** | Tailscale 1.94→1.98 fleet-wide | No | Can ride along Phase 1/2. |
@ -262,6 +262,82 @@ Headscale stale-node cleanup: deleted dead nodes `mailcow` (destroyed CT108) and
---
## Phase 2 — Execution Log (2026-06-21)
**Status: substantially complete; paused 2026-06-21.** All high-priority and security-critical app upgrades are done. Only the low-urgency batch items and the cortex AI stack (protected host, own window) remain.
### Completed upgrades
All targets below were validated healthy post-upgrade; rollback DB dumps retained on the respective CTs under `/root`.
**Wave 1 — cloud apps**
| App | Host | From → To | Notes |
|-----|------|-----------|-------|
| Immich | cloud CT120 | 2.5.6 → 2.7.5 | XSS + ACL-bypass fixes; stack includes Valkey 9 and vectorchord PG |
| PeerTube | media CT110 | 8.0.2 → 8.2.1 | — |
| Nextcloud AIO | cloud CT121 | NC 32.0.4 → **33.0.5** (major); mastercontainer updated | Major NC version; AIO orchestrated update path |
**Wave 2 — media stack (media VM105)**
| App | From → To | Notes |
|-----|-----------|-------|
| Jellyfin | 10.11.6 → 10.11.11 | — |
| Sonarr | → 4.0.17 | — |
| Radarr | → 6.2.1 | — |
| Prowlarr | → 2.4.0 | — |
| Navidrome | → 0.62.0 | — |
| SABnzbd | 4.5.5 → **5.0.4** (major) | Major upgrade; clean migration |
| Lidarr | DEFERRED | Image maintainer hasn't shipped v3 — needs image swap to linuxserver to go v3; left on current |
| Jellyseerr | INTENTIONALLY HELD on `preview-OIDC` tag | Stable lacks OIDC support; switching would break SSO login |
**Authentik (edge2 CT105) 2025.12.4 → 2026.5.3**
Done as 3 sequential hops: 12.4 → 12.6 → 2026.2.4 → 2026.5.3. `pg_dump` before each hop; migrations clean each time. **SSO verified end-to-end by Matt logging into navi.** Closes the May-2026 CVE waves. (ak app version 5.2.15.)
**Forgejo (edge2 CT103) 14.0.5 → 15.0.3**
EOL-branch migration; v15 schema migrations applied cleanly; web 200, API reports 15.0.3. Pre-15 DB dump at `/root/forgejo-db-pre15-20260621.sql`.
**Headscale 0.28 → 0.29.1 (both instances)**
- **CT106** (utility, IdahoMesh 3-node mesh) — upgraded first as rehearsal; native systemd binary.
- **CT107** (edge2, fleet coordinator, Docker) — 32/32 nodes reconnected post-upgrade; health 200; boot-survival fix preserved (ports remain on 10.10.10.25). 0.29 breaking changes documented in [[meshtastic-headscale-runbook]] (key changes: `randomize_client_port` removed — was a hard blocker; ephemeral key config nested; minimum Tailscale client 1.80.0; bare ACL `*` now tailnet-only).
**OpenTAKServer (utility CT109) 1.7.10 → 1.7.12**
Clean; all 9 TAK services healthy.
### Key decision — RabbitMQ LEFT on 3.12.1 (EOL)
Empirically confirmed during the OTS update: updating OTS to 1.7.12 does **not** touch RabbitMQ — OTS 1.7.12 runs fine against RabbitMQ 3.12.1/Erlang OTP 25. Forcing RabbitMQ → 4.x is the wrong move: OTS is built and tested against 3.12, and 4.x has breaking changes (queue-mirroring removal, feature-flag requirements) that would likely break OTS, which does not appear to support 4.x. AMQP/MQTT ports are localhost-bound, so exposure is low. The EOL 3.12 is a low-priority latent risk to revisit only if/when OTS officially supports RabbitMQ 4.x. **NOT a current action item.**
### Incidental fixes and side-work during Phase 2
- **CT107 boot-survival fix** (applied earlier in the effort, during Phase 1 resolution) — rebound headscale/headplane ports to 10.10.10.25, dropped `tailscale-online.target` dependency, disabled `only_start_if_oidc_is_available` gate, repointed edge2 Caddy; proven by reboot self-heal in ~45 s. Also corrected CT107's own tailscale node ControlURL to `vpn.echo6.co` so it self-registers cleanly.
- **Utility node incident (resolved):** a batch delete of 9 LVM-thin snapshots triggered an SSD TRIM/discard storm that spiked I/O and load transiently; compounded by CT103 argus running hot (transcription + docker-compose build churn). Matt migrated argus to the cloud node, resolving the issue; utility load returned to normal. **LESSON: delete thin-pool snapshots one at a time — not in a batch — to avoid the discard storm.**
- **Nextcloud:** granted `matt@echo6.co` the NC admin role. (user_oidc has no group-claim sync, so this is durable across SSO logins.)
- **Radarr:** set up a `\\192.168.1.160\manual` SMB drop folder on the same NFS export as the library (atomic-move imports) for manual movie filing.
- **Snapshot hygiene:** all rollback snapshots cleaned up after validation — Phase 1 `presec-*`, Phase 2 `prewave2-*`, OTS `pre-ots-*` snapshots all removed.
### Remaining work (next session)
**Low-urgency batch:**
- Caddy (edge2 CT101)
- CouchDB 3.4 → 3.5 (livesync CT104 edge2)
- valkey-8 sidecar (searxng CT102)
- NATS 2.14.0 → 2.14.2 (central CT104 utility)
- MediaMTX 1.13 → 1.19 + Mumble on CT109 (OTS companions, separate from the OTS app)
**Protected host — own deliberate window:**
- cortex AI stack: ollama 0.16 → 0.30, open-webui, qdrant, TEI
**Separate deferred projects:**
- Nominatim v5 re-import — see [[nominatim-v5-reimport]]
- Optional: clean up retained `/root` DB dumps (authentik hop1/2/3, forgejo, headscale, OTS, immich, peertube, etc.) once upgrades are confirmed stable
---
## Execution Runbook (Meticulous)
Step-by-step plan to bring every application and package current. Per-app target versions live in **Application Currency** above; this section is the *how/when/order*. Execution model: scoped Sonnet task per target → Opus verifies output → report to Matt; **approval gate before every load-bearing or reboot step.**

View file

@ -10,7 +10,7 @@ related:
- [[meshtastic-sidecar-node]]
- [[headscale-onboard-node]]
- [[caddy]]
updated: 2026-06-18
updated: 2026-06-21
---
# IdahoMesh Tailnet Runbook
@ -572,4 +572,96 @@ ping 100.64.0.14 # cortex — should timeout/unreachable
---
*Last updated: 2026-02-11*
## Headscale 0.28 → 0.29 Upgrade Guide
**Rehearsed on CT 106 (meshtastic-hs) 2026-06-21. Result: success. Use this as the playbook for CT 107 (mesh-bridge / Echo6 Headscale).**
### Breaking changes that require action before starting
#### 1. `randomize_client_port` removed from config.yaml (HARD BLOCKER)
Headscale 0.29 **refuses to start** if `randomize_client_port` exists in `config.yaml`. Remove the line:
```bash
sed -i '/^randomize_client_port:/d' /etc/headscale/config.yaml
```
If the feature is needed, move it to the ACL policy file (`/etc/headscale/acl.json`) as a top-level key:
```json
{ "randomizeClientPort": true }
```
Default is `false`; simply removing the key is equivalent to `randomize_client_port: false`.
#### 2. `ephemeral_node_inactivity_timeout` removed from top-level config (deprecation → removed)
The old top-level key is gone in 0.29. Replace with the new nested path:
```yaml
# OLD (remove this):
ephemeral_node_inactivity_timeout: 30m
# NEW (replace with):
node:
ephemeral:
inactivity_timeout: 30m
```
0.29 accepts the old key with a warning and uses the correct value, but the warning appears on every startup. Fix it to keep logs clean.
#### 3. Strict version upgrade path enforced
0.29 blocks skipping minor versions. If upgrading from 0.27 or earlier, you MUST go 0.27 → 0.28 → 0.29 one step at a time. We were on 0.28, so this is not a blocker for this fleet.
#### 4. ACL wildcard `*` behavior changed (review your policy)
In 0.29, bare `*` in ACL `src`/`dst` resolves to the CGNAT+ULA range (tailnet nodes only) instead of all IPs. If you had `"dst": ["*:*"]` meaning "any internet IP," you must switch to `autogroup:danger-all` as a source or explicit CIDRs. **The IdahoMesh ACL does not use bare `*` — no action needed for this tailnet.**
#### 5. Minimum Tailscale client version raised to v1.80.0
Any tailscale client older than v1.80.0 will be rejected by 0.29. Verify all registered nodes are running a sufficiently recent tailscale before upgrading.
### Other notable 0.29 changes (no action required, FYI)
- GivenName collision suffix changed from random hash (`laptop-abc12xyz`) to monotonic numeric (`laptop`, `laptop-1`). MagicDNS names for nodes with old random suffixes will change on upgrade. The raw Hostname column is unchanged.
- `oidc.expiry` removed; use `node.expiry` for all registration methods.
- `headscale nodes register` deprecated in favour of `headscale auth register --auth-id`.
- `--namespace` flag removed from nodes commands (was already replaced by `--user`).
- New: `trusted_proxies` config option for True-Client-IP headers (previously honoured from any client).
- DB migration is automatic on first start — no manual migration command needed.
### Rollback procedure (if upgrade fails)
1. Stop the service: `systemctl stop headscale`
2. Reinstall the 0.28.0 binary: `cp /usr/local/bin/headscale.bak-v0.28.0-pre029 /usr/local/bin/headscale`
3. Restore config: `cp /etc/headscale/config.yaml.bak-pre029-20260621 /etc/headscale/config.yaml`
4. Restore DB: `cp /root/headscale-db-pre029-20260621.sqlite /var/lib/headscale/db.sqlite`
5. Start: `systemctl start headscale`
> **Note:** 0.29 blocks downgrading once the DB migration has run. The DB backup from step 4 predates the migration, so this rollback is clean.
### CT 106 (meshtastic-hs) upgrade record
- **Date:** 2026-06-21
- **From:** v0.28.0 → **To:** v0.29.1
- **Deployment:** native systemd binary at `/usr/local/bin/headscale`
- **Config changes made:**
- Removed `randomize_client_port: false`
- Replaced `ephemeral_node_inactivity_timeout: 30m` with `node.ephemeral.inactivity_timeout: 30m`
- **DB migration:** automatic, clean, no errors in logs
- **Nodes after upgrade:** all 3 registered (burley-butte offline, mesh-bridge online, mt-isr offline) — same as before
- **Backups (pre-upgrade):** `/etc/headscale/config.yaml.bak-pre029-20260621`, `/root/headscale-db-pre029-20260621.sqlite`, `/usr/local/bin/headscale.bak-v0.28.0-pre029`
### CT 107 (mesh-bridge / Echo6 Headscale) upgrade checklist
CT 107 runs Docker (not native binary). Steps will differ:
1. Back up: `docker cp headscale-vanilla:/var/lib/headscale/db.sqlite /root/headscale-db-pre029-$(date +%Y%m%d).sqlite` and `cp /path/to/config.yaml config.yaml.bak-pre029`
2. Apply the two config changes above (remove `randomize_client_port`, migrate ephemeral timeout key)
3. Change the image tag in `docker-compose.yml` from `headscale/headscale:0.28.x` to `headscale/headscale:0.29.1`
4. `docker compose up -d`
5. Validate: `docker exec headscale-vanilla headscale version`, check `docker logs headscale-vanilla` for migration success, `docker exec headscale-vanilla headscale nodes list`
---
*Last updated: 2026-06-21*