Skip to content

Epic: MeshMonitor Lore — peer knowledge sharing over the mesh #4959

Description

@Yeraze

Summary

MeshMonitor Lore lets MeshMonitor instances share what they learn about the mesh — route topology and unreachable nodes — over the mesh itself, on a private Meshtastic portnum. Peers that hear a Lore packet skip traceroutes the sender already ran, stop re-tracerouting dead nodes, and learn topology they never probed. Net effect: less traceroute traffic mesh-wide.

Feasibility study with full protocol research, airtime math, and live-data simulation: (see linked report / dev-notes doc from the design phase).

Why

Auto-traceroute is one of the most expensive things MeshMonitor does, and every instance on a mesh duplicates it. Real data from a live instance (Yeraze MQTT source, one 24 h window):

  • 5,449 traceroutes observed on the mesh; only 820 answered
  • 524 targets never answered — the worst three were traced 74, 73, and 55 times in 24 h with zero responses
  • Each failed attempt costs ~8 s of flood + want_ack retry airtime; the mesh burns most of its traceroute budget on the same dead nodes all day

One Lore packet (≤233 B, ~2.1 s per transmission, ~4–8 transmissions per flood after duplicate-cancellation) carries 10 unreachable nodes + 23 route edges. It pays for itself when one listening peer skips ~2 traceroutes.

Protocol facts (verified against firmware master + protobufs)

  • Portnum: develop on PRIVATE_APP = 256 (works with unmodified firmware); register MESHMONITOR_LORE_APP in the 64–127 range by PR once the format is stable.
  • Payload cap: 233 B (hard nanopb limit). No firmware fragmentation.
  • Routing: unknown portnums flood exactly like text messages. Exception: CORE_PORTNUMS_ONLY relays drop them entirely (neither relay nor deliver) — coverage is a subset of the mesh.
  • Broadcasts get no acks (firmware strips want_ack). Every packet must stand alone; loss just means fewer skips that cycle.
  • Never set want_response on the broadcast — no module handles the portnum, so every receiving node NAKs back: O(N) unicast storm. Also keep Data.bitfield bit 1 clear.
  • No firmware rate limit on custom portnums — all pacing is ours; gate every send on canTransmit() + isAutomationAirtimeGated().

Wire format

Packed binary inside a minimal protobuf envelope (bytes sections keep protobufjs decode + schema evolution; new sections = new fields, old receivers ignore them). First local proto in the repo: src/server/proto/meshmonitor.proto (NOT inside the protobufs/ submodule; add COPY lines to both Dockerfiles).

header       2 B   version + flags (bit0: per-edge SNR present)
unreachable  1 B count + 4 B/entry:  id24 (3 B) + tries<<4|ageHours (1 B)
edges        1 B count + 8 B/entry:  idA24 + idB24 + SNR int8 (raw quarter-dB) + age (15-min ticks)
  • 24-bit node IDs, resolved against the receiver's node DB. Live data settled this: 16-bit gave 238 ambiguous suffixes among 4,489 known nodes (~5%); 24-bit gives 12 (~0.3%). Ambiguous or unknown IDs are dropped — the data is advisory, so the failure mode is "one node doesn't get skipped."
  • Edges only, no full traceroutes. A fresh edge incident to node X proves X answered recently — the skip list falls out of the edge list. Paths are origin-relative and don't merge; edges do.
  • Extract edges from raw path pairs (fullRoute consecutive pairs), NOT from the route_segments table — that table is position-gated and would silently drop every node without GPS.
  • Filter 0 and 0xffffffff from route arrays (broadcast address really appears in stored routes).
  • A full packet is ~2.7× denser than naive protobuf messages (~228 B vs ~483 B for the same content).

Selection & scoring

More edges exist than fit (1,259 in a 24 h window vs 23 per packet), so selection is the core logic. Greedy pick under the byte budget:

score = wF·freshness + wC·coverageGain + wS·stability + wU·pathUse
  • freshness: exp decay, 6 h half-life
  • coverageGain: +1 per endpoint not yet covered in this packet (spreads picks across the topology)
  • stability: log(observation count), capped
  • pathUse: log(distinct targets routed through the edge)
  • rotation: soft penalty (~×0.2) for edges shared in the last N cycles — never hard exclusion, so live backbone edges resurface for late-joining peers

Weights come from the Selection Bias setting: Balanced (0.4/0.3/0.2/0.1), Prefer Rare Routes (boost coverage+novelty), Prefer Strong Routes (boost stability + SNR floor).

Unreachable list: targets with ≥2 attempts and 0 responses in the window, ranked by attempt count, capped at 10.

Simulation on live data: 8 rotated cycles cover ~100/298 topology nodes and 184/1,259 edges — the stale one-off edges age out of contention on their own.

Ingestion (receive side)

  • New case in the portnum switch (meshtasticManager.ts ~6297) + decode case in processPayload + packet-monitor preview branch (MESH_BEACON_APP is the closest precedent).
  • Persist via a dataEventEmitter subscriber service (the meshBeaconOfferService best-effort pattern) into a new peer_lore table — peer claims stay separate from local observations.
  • Skip logic: filter step in autoTracerouteSelectionService.selectNodeNeedingTraceroute() (between the AND and OR filter blocks, new dep on AutoTracerouteSelectionDeps): drop candidates covered by a fresh peer edge; deprioritize peer-reported-unreachable targets.
  • Peer roster is free: any received Lore packet's from header identifies a peer instance — no presence beacon. Roster = senders heard in the last N hours.
  • Work partitioning (later phase): instances that hear each other split auto-traceroute targets by hash(target) % N == myIndex. N instances → ~1 instance's worth of active traffic.

Trust model

Any channel member can forge these packets. Rules:

  • Peer data is advisory only — it delays local verification, never replaces it (peer-covered nodes still get traced at an extended expiry, e.g. 4× normal).
  • Rows are marked peer-sourced; never merged into local observations.
  • Per-peer ingest caps (rows/hour) so one bad actor can't flood the table.
  • Never act on peer data for anything security-relevant.

Settings & UI — "MeshMonitor Lore" panel

Per-source section in AutomationTab beside Auto-Traceroute. Keys via getSettingForSource; all added to VALID_SETTINGS_KEYS + the SettingsDraft/handleSave literal.

Setting Key Default Range
Enable broadcast loreEnabled off listen-only always on once shipped
Broadcast Interval loreIntervalHours 4 1–24
Channel loreChannelIndex 0 channel dropdown (broadcast + listen filter)
Minimum Routes loreMinRoutes 5 0–50 — skip the cycle below this many fresh, unshared edges
Selection Bias loreSelectionBias Balanced Balanced / Prefer Rare / Prefer Strong
Maximum Packet Size loreMaxPacketSize 233 60–233 B, live ToA readout beside it
Hop Limit loreHopLimit 3 1–7 — dominant airtime lever; warning text beside input

Mesh-impact rules baked in:

  • Broadcast off by default; a fresh install benefits without contributing until the user opts in.
  • Last-broadcast timestamp persisted to settings (autoAnnounceService pattern) — a settings save never re-arms the timer.
  • Every send gated on canTransmit() + airtime gate; skipped cycles log why.

Live simulation block

GET /api/sources/:id/lore/simulate?interval=&minRoutes=&bias=&maxSize=&hopLimit= runs the production selection code path against current DB data. Panel shows, live as knobs change (debounced):

  • Packet preview (unreachable + edges with names), packed size, ToA
  • Cost: ToA × estimated relay count → "~6.4 s mesh-wide per broadcast, ~38 s/day"
  • Benefit: est. traceroutes displaced/day per listening peer (own auto-traceroute rate + duplicate rate + unreachable retries)
  • Verdict chip: net-positive / marginal / net-negative

Recommended Values button

GET /api/sources/:id/lore/recommend derives: interval from edge-discovery rate (fill ~1 packet per cycle), min-routes from the noise floor, packet size from observed channel utilization, hop limit from the hop-distance distribution of known nodes. Returns the simulate payload so the panel shows why before the user applies.

Phases

Phase 1 — Listen-only (zero mesh cost, proves decode on real packets)

  • meshmonitor.proto + protobuf loader + Dockerfile COPY lines
  • Portnum constant mirrored in all four places (constants/meshtastic.ts, normalizePortNum, getPortNumName, packetFormat.ts)
  • Receive handler + packet-monitor preview + peer_lore table (migration via all three backends)
  • Ingest service (event-emitter subscriber, per-peer caps)

Phase 2 — Broadcast + skip logic (the headline benefit)

  • Generic createDataPacket() in meshtasticProtobufService + sendLorePacket() in the manager (modeled on sendTraceroute: guards, virtual-node mirror, packet log)
  • loreService scheduler (autoAnnounce pattern: persisted last-fire, airtime gate, clamp-in-setter)
  • Selection/scoring engine + packer (unit-tested against golden packets)
  • Skip filter in autoTracerouteSelectionService + extended-expiry re-verify
  • Settings keys + Lore panel UI (no sim block yet)

Phase 3 — Simulation & recommendations UI

  • /lore/simulate + /lore/recommend endpoints (share the production selection module)
  • Live-updating sim block + Recommended Values button + verdict chip

Phase 4 — Peer coordination

  • Peer roster from received packets; partition auto-traceroute targets by stable hash
  • Lore status view: peers heard, edges/unreachable ingested, traceroutes skipped (the payoff, made visible)

Phase 5 — Upstream registration

  • PR MESHMONITOR_LORE_APP into meshtastic/protobufs (64–127 range) once the format survives real use

Open questions

  • Confirm defaults above (interval 4 h, hop limit 3, min routes 5) — they're policy, not implementation
  • MQTT uplink policy for PRIVATE_APP is unverified — check whether Lore packets leak to public MQTT and whether we care
  • Should MeshCore sources get a Lore analog later, or is this Meshtastic-only by design?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions