Two live backlog-broadcast leaks traced to the in-memory first-poll seed:
(1) incremental-fetch adapters (wzdx: registry tick [0 events] then feeds
tick [many]) got marked _seeded on the EMPTY first tick, so the real batch
next tick all looked "new" and broadcast; (2) in-memory seed lost on restart.
Fix — durable baseline + guard:
- _seed_from_persistent() at store init: pre-load already-received item keys
from the persistent hazard tables into self._seen, so nothing ever received
can re-broadcast (immune to fetch staging + restart). Only sources whose
native emit key PROVABLY equals a persistent key are durably seeded:
wzdx (traffic_events.external_id) + usgs_quake (quake_events.event_id).
Resilient (per-table try/except; missing table -> skip).
- _seen_key() now namespaces by evt["source"] (matches persistent tables),
via shared _key_ext/_key_eid helpers used by both seed and live emit so
they can't drift.
- non-empty-seed guard: _ingest marks only sources that carried >=1 event
this poll as _seeded -> an empty first tick can never seed-then-leak. This
is the root-cause fix; covers all adapters (roads511/traffic fetch
atomically per tick, so the guard fully protects them).
- storage untouched (self._events populated for every event); Central
path/deciders untouched.
Live-DB verified: seed pre-loads 784 wzdx + 8 quake keys -> a live wzdx poll
of 784 known zones broadcasts 0. +6 tests (incremental staging, restart,
persistent-preseed, fresh-DB fallback); suite at 10-failure baseline.
Co-authored-by: Matt Johnson <mj@k7zvx.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Replace the native path's "scan accumulated state + suppress what we've
already broadcast" model with "broadcast only what newly arrived from the
API this poll." Storage is unchanged (self._events + firms_pixels etc. are
populated for EVERY received item, so the LLM/get_active backlog is intact);
only the BROADCAST decision changes.
- env/store.py: per-adapter in-memory seen-set (_seen) + _seeded. First
data-bearing poll for an adapter seeds keys and emits NOTHING (that batch
is pre-existing backlog); later polls emit only keys not seen before.
Restart => empty sets => next poll re-seeds silently. Structurally
impossible to broadcast backlog on cold start / restart / re-enable.
Key = external_id -> event_id -> content hash, namespaced per adapter.
self._events[key]=evt still runs unconditionally (storage preserved).
- Fixes the ~175 (roads511) / ~782 (wzdx) cold-start bursts AND the latent
quake/nws version (they only looked safe because Central pre-populated
their broadcast tables).
- env/satpass.py: broadcast on AOS IMMINENCE (now < aos <= now+lead,
broadcast_lead_seconds default 3600), future-only; window_hours still
governs prediction depth. Strict norad_ids post-filter + fixed
_parse_norad_ids char-iteration bug (cause of GOES/METEOR leak).
- Central path + broadcast-state tables untouched (native-only gate).
13 new tests; full suite at 10-failure baseline (1697 passed).
Co-authored-by: Matt Johnson <mj@k7zvx.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>