echo6-docs/vault/runbooks/edge2-boot-recovery.md
echo6-autocommit 8f95c6e5c3 auto: docs sync 2026-08-15T12:00:04+00:00
Files changed: engine/.embcache.json engine/changelog.md engine/lint-report.md vault/.obsidian/workspace.json vault/docs/hardware/environment.md vault/projects/fleet-storage-memory-upgrade.md vault/projects/navi-recon-separation.md vault/runbooks/corescope-ingest-stall-oom.md vault/runbooks/edge2-access-reference.md vault/runbooks/edge2-boot-recovery.md vault/runbooks/navi-lift-to-media.md vault/runbooks/peertube-sitemap-redis-oom.md
2026-08-15 12:00:04 +00:00

152 lines
7.7 KiB
Markdown

---
title: edge2 Boot Recovery — compose-reconcile Tailscale Race
type: runbook
tags:
- mesh
aliases: []
related:
- [[caddy]]
- [[headscale-oidc-boot-order]]
- [[lxc-service-migration]]
- [[services]]
- [[edge2-access-reference]]
updated: 2026-08-15
---
# edge2 Boot Recovery — compose-reconcile Tailscale Race
edge2 reboots and most of its [[services]] never come back. Every vhost behind it returns 502 while all nine containers report `running`. Hit on 2026-08-14 for 2h05m. Root cause found and fixed the same day.
The June 2026 `compose-reconcile.service` was built to prevent exactly this and had never been tested by a real reboot. This was its first, and it lost.
---
## Signature
```
auth 502 forge 502 notes 502 vault 502 matrix 502 element 502
vpn 200 <- headscale is not Docker, so it is unaffected
```
Home services through the utility [[caddy]] (ai, jellyfin, immich, nextcloud) stay healthy — the blast radius is edge2 only.
`pct list` shows every CT `running`, which is misleading. Look inside them:
```bash
ssh edge2 'sudo -n pct exec 105 -- docker ps -a --format "{{.Names}} | {{.Status}}"'
```
- CT 102 vaultwarden, CT 103 forgejo, CT 104 livesync: containers stuck in **`Created`**, never started
- CT 105 [[authentik]], CT 106 matrix: app containers crash-looping on `Temporary failure in name resolution` for their compose-sibling database hostnames (`postgresql`, `matrix-postgres`)
**[[authentik]] being down 502s everything behind forward-auth**, including services that are themselves perfectly healthy. On 2026-08-14 this is how the outage surfaced: `mesh.echo6.co` returned 502, but MeshMonitor was fine the entire time — only its login was broken. Check `auth.echo6.co` before diagnosing any individual service.
---
## Cause
The reconcile unit fired in every container and failed in every one:
```
failed to bind host port 100.64.0.36:9000/tcp: cannot assign requested address (authentik)
failed to bind host port 100.64.0.33:8086/tcp: cannot assign requested address (vaultwarden)
```
The compose files publish ports on each container's **Tailscale** address. The unit's `ExecStartPre` waited only for `docker info` — not for `tailscale0` to have its address. It ran about two minutes after boot, the address did not exist yet, the bind failed, and with **no retry** the oneshot exited 1 and gave up permanently.
```bash
ssh edge2 'sudo -n pct exec 105 -- systemctl status compose-reconcile.service'
ssh edge2 'sudo -n pct exec 105 -- journalctl -u compose-reconcile.service -b'
```
---
## Fix (applied 2026-08-15)
Deployed to all six Docker CTs — 102 vaultwarden, 103 forgejo, 104 livesync, 105 authentik, 106 matrix, 107 headscale. Originals backed up in place as `compose-reconcile.service.bak-<stamp>`.
- `After=` / `Wants=` now include **`tailscaled.service`**
- A second `ExecStartPre` waits for `tailscale0` to hold an IPv4, **bounded at 180 tries** so it can never hang a boot
- `ExecStart` retries `docker compose up -d` **six times at ten seconds**, giving roughly four minutes of total tolerance
- `TimeoutStartSec=900`
```ini
ExecStartPre=/bin/sh -c 'until docker info >/dev/null 2>&1; do sleep 1; done'
ExecStartPre=/bin/sh -c 'n=0; until ip -4 addr show tailscale0 2>/dev/null | grep -q "inet "; do n=$$((n+1)); if [ $$n -ge 180 ]; then break; fi; sleep 1; done'
ExecStart=/bin/sh -c 'docker compose down --remove-orphans || true; n=0; until docker compose up -d; do n=$$((n+1)); if [ $$n -ge 6 ]; then exit 1; fi; sleep 10; done'
```
Two traps worth knowing before editing this unit:
**`Restart=on-failure` is not usable.** systemd forbids `Restart=` on `Type=oneshot`. That is why the retry lives inside `ExecStart` instead. Do not "fix" it by adding `Restart=`.
**A `$` in a unit file must be written `$$`** to reach the shell literally. Get it wrong and the loop counter silently expands to nothing.
`WorkingDirectory` differs per container (`/opt/authentik`, `/opt/vaultwarden`, and so on) — preserve each one when rewriting.
---
## Manual recovery
If it wedges again before the unit is trusted, re-run the reconcile. Once Tailscale is up the binds succeed and it heals in seconds.
```bash
ssh edge2 'sudo -n pct exec 105 -- systemctl restart compose-reconcile.service' # authentik FIRST
# then 102, 103, 104, 106
```
**Do CT 105 first** — clearing [[authentik]] lifts the 502 cascade off everything else.
The underlying heal is `docker compose down --remove-orphans && docker compose up -d`. A plain `up -d` does **not** work: compose sees the containers "Up" and the network "present" and concludes state has converged. **Never use `-v`** — it would wipe the CouchDB, Postgres and Vault volumes.
---
## Testing this without rebooting the host
Reboot a single container. That exercises the real race — `tailscaled` against Docker at container boot — with no risk to the host.
```bash
ssh edge2 'sudo -n pct reboot 103'
ssh edge2 'sudo -n pct exec 103 -- systemctl show compose-reconcile.service -p Result -p ExecMainStatus'
```
Proven this way on 2026-08-15: `Result=success`, `ExecMainStatus=0`, containers up, `forge.echo6.co` 200. **Never reboot the edge2 host to test this.**
---
## Verify
```bash
for u in auth forge notes vault matrix element vpn; do
printf '%-8s ' "$u"; curl -s -o /dev/null -m 15 -w 'HTTP %{http_code}\n' "https://$u.echo6.co/"
done
```
Healthy is `auth` 302 (login redirect), `notes` 401 (CouchDB auth), and 200 for the rest. A 302 from a forward-auth vhost is correct behaviour, not a fault.
---
## What triggered it: the 2026-08-14 hard reset
edge2 died mid-operation at **01:12:45 UTC** after 46 days of uptime and returned 26 seconds later. Nothing inside the guest asked for it or recorded a fault. Ruled out with evidence:
- **No reboot issued** — no shutdown record in `last -x`; the PVE task log has no `stopall`, only a boot-time `startall`
- **No clean shutdown** — zero shutdown markers in the final five minutes; the journal stops mid-stream in routine sshd and tailscaled traffic
- **No panic, oops, OOM, I/O error, hung task or soft lockup** anywhere in the previous boot; the last kernel message was a routine hourly `drop_caches` 56 minutes earlier
- **Not the watchdog** — edge2 is not clustered, so there is no HA and `/dev/watchdog` was never armed
- **No resource pressure** — last full sample before death: load 3.38, CPU 15.9%, iowait 2.43%, memory 6.2 GB of 25 GB
- **Not datacenter-wide** — edge1, the other Contabo VPS, had been up since [[2026-06-19]] and sailed straight through
- **No migration signature** — identical CPU (AMD EPYC family 0x17 model 0x1 stepping 0x2) and memory map before and after
The boot was a full cold SeaBIOS re-init, so the VM was destroyed and recreated at hypervisor level. Contabo's status page listed no incident, but they do not publish single-node resets.
Honest caveat: a guest kernel panic cannot be fully excluded. `kernel.panic = 10` is set and the 26-second gap fits panic then reset. A hard panic can lose the journal tail. Nothing supports it, and the only discriminator is console output, which only Contabo's panel retains.
One suggestive detail: about two minutes of external network flakiness immediately beforehand — tailscaled's IPv6 flapped at 01:10:52, a logtail upload failed at 01:11:26 and took 55 seconds to recover, then death at 01:12:45.
Closing it needs Contabo's server history, or a ticket asking why `184.174.35.153` was reset at 01:12:45 UTC on 2026-08-14.
**Gap this exposed:** edge2 has no sysstat or atop, so only 30-minute PVE RRD samples exist for that window — too coarse to catch a short spike. `rrdtool` CLI is not installed either; use the API instead, which needs no install:
```bash
ssh edge2 'sudo -n pvesh get /nodes/edge2/rrddata --timeframe day --output-format json'
```