Commit graph

1 commit

Author SHA1 Message Date
0a9252329a
fix(central): don't crash boot when Central is unreachable at startup (#26)
The Central NATS consumer's start() was called unguarded during boot, so if
Central was enabled but unreachable at startup, nats.connect() raised
NoServersError, propagated through bot.start(), and crashed the process —
crash-looping under Docker restart:unless-stopped.

Now _start_central_consumer_guarded() wraps start() in try/except: on failure
it logs a warning and continues booting (LLM bot, Meshtastic/MeshCore,
mesh-health, and native feeds all start), then a background retry loop
(30s->300s backoff) re-attempts the initial connect until it succeeds. Once
connected, NATS's own allow_reconnect handles runtime drops. The retry task is
cancelled cleanly on stop(). No retry is scheduled when nothing is
central-sourced.

Tests: +tests/test_central_boot_guard.py (11); 0 new failures.

Co-authored-by: Matt Johnson <mj@k7zvx.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 01:13:50 -06:00