Files changed: engine/.embcache.json engine/changelog.md engine/lint-report.md vault/.obsidian/workspace.json vault/docs/hardware/environment.md vault/projects/fleet-storage-memory-upgrade.md vault/projects/navi-recon-separation.md vault/runbooks/corescope-ingest-stall-oom.md vault/runbooks/edge2-access-reference.md vault/runbooks/edge2-boot-recovery.md vault/runbooks/navi-lift-to-media.md vault/runbooks/peertube-sitemap-redis-oom.md
7.7 KiB
| title | type | tags | aliases | related | updated | ||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| edge2 Boot Recovery — compose-reconcile Tailscale Race | runbook |
|
|
2026-08-15 |
edge2 Boot Recovery — compose-reconcile Tailscale Race
edge2 reboots and most of its services never come back. Every vhost behind it returns 502 while all nine containers report running. Hit on 2026-08-14 for 2h05m. Root cause found and fixed the same day.
The June 2026 compose-reconcile.service was built to prevent exactly this and had never been tested by a real reboot. This was its first, and it lost.
Signature
auth 502 forge 502 notes 502 vault 502 matrix 502 element 502
vpn 200 <- headscale is not Docker, so it is unaffected
Home services through the utility caddy (ai, jellyfin, immich, nextcloud) stay healthy — the blast radius is edge2 only.
pct list shows every CT running, which is misleading. Look inside them:
ssh edge2 'sudo -n pct exec 105 -- docker ps -a --format "{{.Names}} | {{.Status}}"'
- CT 102 vaultwarden, CT 103 forgejo, CT 104 livesync: containers stuck in
Created, never started - CT 105 authentik, CT 106 matrix: app containers crash-looping on
Temporary failure in name resolutionfor their compose-sibling database hostnames (postgresql,matrix-postgres)
authentik being down 502s everything behind forward-auth, including services that are themselves perfectly healthy. On 2026-08-14 this is how the outage surfaced: mesh.echo6.co returned 502, but MeshMonitor was fine the entire time — only its login was broken. Check auth.echo6.co before diagnosing any individual service.
Cause
The reconcile unit fired in every container and failed in every one:
failed to bind host port 100.64.0.36:9000/tcp: cannot assign requested address (authentik)
failed to bind host port 100.64.0.33:8086/tcp: cannot assign requested address (vaultwarden)
The compose files publish ports on each container's Tailscale address. The unit's ExecStartPre waited only for docker info — not for tailscale0 to have its address. It ran about two minutes after boot, the address did not exist yet, the bind failed, and with no retry the oneshot exited 1 and gave up permanently.
ssh edge2 'sudo -n pct exec 105 -- systemctl status compose-reconcile.service'
ssh edge2 'sudo -n pct exec 105 -- journalctl -u compose-reconcile.service -b'
Fix (applied 2026-08-15)
Deployed to all six Docker CTs — 102 vaultwarden, 103 forgejo, 104 livesync, 105 authentik, 106 matrix, 107 headscale. Originals backed up in place as compose-reconcile.service.bak-<stamp>.
After=/Wants=now includetailscaled.service- A second
ExecStartPrewaits fortailscale0to hold an IPv4, bounded at 180 tries so it can never hang a boot ExecStartretriesdocker compose up -dsix times at ten seconds, giving roughly four minutes of total toleranceTimeoutStartSec=900
ExecStartPre=/bin/sh -c 'until docker info >/dev/null 2>&1; do sleep 1; done'
ExecStartPre=/bin/sh -c 'n=0; until ip -4 addr show tailscale0 2>/dev/null | grep -q "inet "; do n=$$((n+1)); if [ $$n -ge 180 ]; then break; fi; sleep 1; done'
ExecStart=/bin/sh -c 'docker compose down --remove-orphans || true; n=0; until docker compose up -d; do n=$$((n+1)); if [ $$n -ge 6 ]; then exit 1; fi; sleep 10; done'
Two traps worth knowing before editing this unit:
Restart=on-failure is not usable. systemd forbids Restart= on Type=oneshot. That is why the retry lives inside ExecStart instead. Do not "fix" it by adding Restart=.
A $ in a unit file must be written $$ to reach the shell literally. Get it wrong and the loop counter silently expands to nothing.
WorkingDirectory differs per container (/opt/authentik, /opt/vaultwarden, and so on) — preserve each one when rewriting.
Manual recovery
If it wedges again before the unit is trusted, re-run the reconcile. Once Tailscale is up the binds succeed and it heals in seconds.
ssh edge2 'sudo -n pct exec 105 -- systemctl restart compose-reconcile.service' # authentik FIRST
# then 102, 103, 104, 106
Do CT 105 first — clearing authentik lifts the 502 cascade off everything else.
The underlying heal is docker compose down --remove-orphans && docker compose up -d. A plain up -d does not work: compose sees the containers "Up" and the network "present" and concludes state has converged. Never use -v — it would wipe the CouchDB, Postgres and Vault volumes.
Testing this without rebooting the host
Reboot a single container. That exercises the real race — tailscaled against Docker at container boot — with no risk to the host.
ssh edge2 'sudo -n pct reboot 103'
ssh edge2 'sudo -n pct exec 103 -- systemctl show compose-reconcile.service -p Result -p ExecMainStatus'
Proven this way on 2026-08-15: Result=success, ExecMainStatus=0, containers up, forge.echo6.co 200. Never reboot the edge2 host to test this.
Verify
for u in auth forge notes vault matrix element vpn; do
printf '%-8s ' "$u"; curl -s -o /dev/null -m 15 -w 'HTTP %{http_code}\n' "https://$u.echo6.co/"
done
Healthy is auth 302 (login redirect), notes 401 (CouchDB auth), and 200 for the rest. A 302 from a forward-auth vhost is correct behaviour, not a fault.
What triggered it: the 2026-08-14 hard reset
edge2 died mid-operation at 01:12:45 UTC after 46 days of uptime and returned 26 seconds later. Nothing inside the guest asked for it or recorded a fault. Ruled out with evidence:
- No reboot issued — no shutdown record in
last -x; the PVE task log has nostopall, only a boot-timestartall - No clean shutdown — zero shutdown markers in the final five minutes; the journal stops mid-stream in routine sshd and tailscaled traffic
- No panic, oops, OOM, I/O error, hung task or soft lockup anywhere in the previous boot; the last kernel message was a routine hourly
drop_caches56 minutes earlier - Not the watchdog — edge2 is not clustered, so there is no HA and
/dev/watchdogwas never armed - No resource pressure — last full sample before death: load 3.38, CPU 15.9%, iowait 2.43%, memory 6.2 GB of 25 GB
- Not datacenter-wide — edge1, the other Contabo VPS, had been up since 2026-06-19 and sailed straight through
- No migration signature — identical CPU (AMD EPYC family 0x17 model 0x1 stepping 0x2) and memory map before and after
The boot was a full cold SeaBIOS re-init, so the VM was destroyed and recreated at hypervisor level. Contabo's status page listed no incident, but they do not publish single-node resets.
Honest caveat: a guest kernel panic cannot be fully excluded. kernel.panic = 10 is set and the 26-second gap fits panic then reset. A hard panic can lose the journal tail. Nothing supports it, and the only discriminator is console output, which only Contabo's panel retains.
One suggestive detail: about two minutes of external network flakiness immediately beforehand — tailscaled's IPv6 flapped at 01:10:52, a logtail upload failed at 01:11:26 and took 55 seconds to recover, then death at 01:12:45.
Closing it needs Contabo's server history, or a ticket asking why 184.174.35.153 was reset at 01:12:45 UTC on 2026-08-14.
Gap this exposed: edge2 has no sysstat or atop, so only 30-minute PVE RRD samples exist for that window — too coarse to catch a short spike. rrdtool CLI is not installed either; use the API instead, which needs no install:
ssh edge2 'sudo -n pvesh get /nodes/edge2/rrddata --timeframe day --output-format json'