Repository navigation
Conversation
`rpk discover` reads infrastructure and emits it as RackPeek YAML — one System for the machine it runs on, a Service per published Docker container, and a Server/System tree for a Proxmox cluster. Safe by default: prints unless --push, --dry-run previews, and imports are merge-only — discovery can add and update but never remove or rename what the user wrote. Identity: each discovered resource carries a hashed discoveryId (machine-id / platform UUID / Docker daemon id), so re-runs update renamed resources instead of duplicating them, hand-written entries are adopted, and cloned machine-ids are rejected with a fix hint. Remote Docker engines are identified via GET /info with a graceful fallback behind restricted socket proxies, and the merge preserves user-chosen runsOn links a remote collector cannot see. Persistence: discoveryId ships as schema v4 with a forward migration, per AGENTS.md §6; v3 files load and re-save as v4. The web server now eagerly loads the config at startup so the inventory API cannot merge against an empty collection, and every write path retries a failed load instead of overwriting the user's file. Tested by fixture-driven suites for the probes, parsers, mappers and id resolution, HTTP-level merge tests against a real server, real probe runs in CI on Linux and macOS, and E2E CLI coverage; verified live against a real Docker engine (local socket, TCP bridge, and a CONTAINERS=1 socket proxy) and a Proxmox fixture set. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Blazor Server prerenders every page before the circuit attaches, and input events sent to that static HTML are silently lost — so a fast test run could fill the add form, click add, and be told "name is required" because the server-side model never saw the keystrokes. CI runners lose that race intermittently (seen on OtherCardTests); a fast machine wins it, which is why it never reproduced locally. MainLayout now renders a hidden probe that flips to data-circuit-ready="true" on the first OnAfterRender — which never runs during prerendering, so the flip proves a live circuit. The add page object waits on the probe before typing. Reproduced deterministically first with the existing BlazorLatency helper (300ms per SignalR frame widens the attach window from milliseconds to seconds): fails on the old page object, passes with the wait. That repro stays as a regression test. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Add auto discovery: rpk discover system / docker / proxmox
RackPeek.Web now hosts a Model Context Protocol server over streamable HTTP at /mcp — no extra process, it runs whenever the server runs, and it sits behind the same X-Api-Key gate as /api/inventory (503 until RPK_API_KEY is set, so it is off by default). 24 consolidated tools in a new RackPeek.Mcp project (Domain-only deps, so the CLI binaries stay MCP-free): - Query: list/get/search resources, summary, containment tree, connections, subnets, and get_schema so agents learn the YAML format before writing. - Editing: upsert_resources (bulk YAML with dry-run diffs — the main write path), delete/rename/clone, tag/label edits, connection add/remove. - Exporters: ansible inventory, ssh config, hosts file, mermaid topology. - Git: status + commit(+push), active on the existing GIT_TOKEN opt-in. - Discovery: docker and proxmox run from the server with preview-then-apply; credentials come from server config, never from the conversation. Every tool reuses the existing use cases, so validation, conflict rules, discovery-id resolution and file locking behave exactly like the CLI and UI. Domain errors surface as actionable tool errors; anything unexpected stays generic. GlobalSearchService moves from Shared.Rcl to RackPeek.Domain/Search so the search tool can reach it without UI dependencies. Testing is end-to-end at two tiers: Tests.Mcp (70 tests) drives a real MCP client over WebApplicationFactory through every tool down to the YAML on disk, including fake docker/proxmox engines on real sockets and schema conformance on everything the tools emit; Tests.E2e gains McpE2eTests running a full session against the shipped Docker image over the network. New mcp-tests CI job and `just test-mcp` target. Also fixes a duplicate lounge-ap resource in the v3/v4 demo configs that crashed `rpk graph topology` (surfaced by the new export tool tests). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Laptop, Other and Ups were the only hardware kinds that could not describe their physical connectivity, which made it impossible to model USB-attached peripherals, IPTV decoders or a UPS monitoring link without misusing Desktop as a stand-in. Domain - Laptop, Other and Ups implement IPortResource (List<Port>? Ports). The open-generic registration of IAddPortUseCase<> / ISetPortUseCase<> / IRemovePortUseCase<> already covers the new types, so no use-case wiring was needed. - Add "usb" to Nic.ValidNicTypes. That single array drives port validation, CLI suggestions and both UI dropdowns. - Add PortSummaries.Describe and surface ports in the describe use cases: UpsDescription.PortSummary, OtherDescription.PortSummary and LaptopDescription.NicCount (matching DescribeDesktopUseCase). CLI - rpk ups port add|set|del and rpk other port add|set|del, following the --count / --index convention used by switch, router and firewall. - rpk laptops nic add|set|del, following the --ports convention used by servers and desktops. UI - PortGroupEditor on the laptop, other and UPS cards. Schema v4 - ports on the laptop, ups and other definitions; "usb" in the port type enum. All three published copies stay byte-identical.
CLI (Tests/): UpsPortWorkflowTests, OtherPortWorkflowTests and LaptopNicWorkflowTests exercise add/set/del/describe with exact YAML asserts, plus four error facts per kind (missing resource, invalid port type, invalid set index, invalid del index) and --help assertions for each new command branch. Playwright (Tests.E2e/): one card test per kind adds two port groups through PortGroupEditor, reloads, and asserts both groups and their individual ports survive the round trip. UpsCardPom, OtherCardPom and LaptopCardPom gained a Ports member and thin wrappers over the existing PortsPom, mirroring AccessPointCardPom. Every test pairs the new usb type with rj45 so the port summary has more than one group to fold.
Adds `rpk ups port`, `rpk other port` and `rpk laptops nic` (each with add/set/del) to cli-commands.md and cli-commands-index.md. Generated with generate-docs.sh. Two deviations were needed to run it outside CI, neither of which affects the output: - the publish step emits RackPeek.exe on Windows, so the `-x` probe on the extensionless path never matches - the script writes raw_docs/Commands.md and raw_docs/CommandIndex.md, but the files actually shipped and listed in docs-index.json are raw_docs/cli-commands.md and raw_docs/cli-commands-index.md The diff is purely additive, which confirms the rest of the output matches what is already committed.
Add ports to Laptop, Other and Ups, and a usb port type
…tems The fourth collector, for machines nothing else can describe — no agent, no API, just an address that answers. Pure .NET (no nmap): a bounded- parallel ICMP ping sweep with TCP connect fallback on a curated port list (a host is alive if either answers — plenty of gear drops ICMP), ARP for MAC identity, reverse DNS for names. Design follows the discovery philosophy: INetworkProbe is the thin IO seam; ArpTableParser, target enumeration and NetworkScanMapper are pure and fixture-tested. Identity is MAC-seeded (rpk1:net:<mac>), normalised across platforms because macOS prints unpadded MAC octets where Linux pads them; hosts with no ARP entry (routed segments) fall back to an IP-seeded id and the command says so. Scanned cards are deliberately sparse — ip, mac label, name, id, and nothing else — so a rescan can never overwrite the type/os/cores/ram a user or agent collector filled in on an adopted card. The v4 schema's System definition loses its required [type, os, cores, ram] to allow that; loosening validation is backwards-compatible and the emitters (Proxmox included) could already produce Systems without cores. --cidr defaults to the machine's own subnet and sweeps are capped at /16; --ports, --timeout and --parallel tune the sweep. Tests: 48 new in Tests.Discovery — both ARP formats parse to identical MACs, target enumeration edges (/16, /24, /30, /31, /32, top-of-space wrap), mapper identity contracts, scripted-probe scanner semantics (ping-only, TCP-only, first-answer short-circuit, ARP-after-sweep, concurrency cap), real-probe loopback and dead-block scans, and merge-through-the-real-server e2e: idempotent rescans, renames that survive, DHCP moves updating the same card, never stealing an agent-discovered host's identity, adopting a hand-written one. Plus 11 CLI validation tests that fail before any packet is sent. Proven against a real /24: 7 hosts found in ~10s, all identities stable across consecutive runs, the scanning machine finds itself. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Three integration fixes from reviewing how the scan sits with the rest of the platform: - Windows ARP support: arp.exe prints dash-separated MACs in a three- column table and only knows `arp -a` — the parser now reads that format (normalising to the same colon form as Linux/macOS, so the same machine hashes to the same id from any platform) and the probe falls back from `arp -an` to `arp -a`. Without this, every scan from the shipped win-x64 binary silently degraded to IP-seeded identity. - One MAC answering on several addresses (a gateway's VIPs/aliases) now folds the address into each card's id seed instead of emitting duplicate ids, which the import rejects with a misleading machine-id hint. Deterministic per (mac, ip). - The Ansible exporter now reads a System's own ip after the address labels, exactly like the ssh and hosts exporters already do — so discovered hosts are addressable in inventories without hand-adding labels, and an explicit ansible_host label still wins. Verified against a real /24 that all previously-emitted discovery ids are byte-identical after the mapper change. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
An adversarial review of the branch surfaced ten verified findings; all are addressed: - Identity stability (the review's top finding): hosts sharing a MAC now collapse into ONE card (lowest address as its ip, all addresses in an 'ips' label) instead of count-dependent id seeds — a VIP failing over or appearing can no longer move or duplicate a machine's identity. - The /16 sweep cap moved into NetworkScanner itself, so an auto-detected VPN/CGNAT /10 hits the same wall a typed --cidr does. - IpHelper.ToUInt32 range-checks octets: 192.168.256.0/24 is refused instead of silently wrapping into 192.169.0.0 and probing a network the user never named. - The ARP source chain trusts parse results, not exit codes: a source only wins if it yields usable entries, so arp.exe printing usage text with exit 0 can no longer cost Windows scans their MAC identity. - INetworkProbe.IsSupported guards the browser: the WASM viewer console now says scanning is unsupported instead of reporting a false-empty network. - LocalSubnet skips 169.254/16 self-assigned addresses when picking the subnet to auto-sweep. - Reverse DNS resolves in parallel under the same concurrency gate, with the per-lookup cap promoted from a buried constant to NetworkScanOptions.DnsTimeout — a PTR-dropping resolver now costs one timeout, not one per host. - Cidr.TryParse is the single definition of CIDR validity: the settings and ServiceSubnetsUseCase both use it, the command no longer re-parses on faith, and ResolvedPorts caches its parse instead of a null-forgive. - New drift guard: SchemaTests pins the wwwroot schema copies the server and viewer actually serve to the published schemas/ copies (the #310/ #311 failure mode), for every version. 20 new/updated tests; re-verified against the real /24 that every previously emitted discovery id is unchanged. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The same physical machine seen by two collectors used to become two cards, because their identity schemes cannot derive each other: the agent identifies by machine-id, a scan by MAC. Now they meet in the middle — one machine, one card, whichever collector ran first. The agent's side of the bridge: `rpk discover system` records the machine's physical-NIC MACs (loopback and virtual interfaces excluded) as a `macs` label, normalised by the same code that reads ARP tables so both sides always agree on the spelling. The local docker collector builds its host card through the same mapper, so it participates for free. The resolver's side: after an id lookup misses, a MAC shared with exactly one stored card of the same kind unifies onto that card. The stronger identity wins — an agent id replaces a scan id; a scan id is nulled before the merge so it can never downgrade one. Ambiguity never unifies: a MAC claimed by two stored cards identifies nothing, and two agent-grade identities sharing a MAC (cloned VMs) stay apart — that is what machine-ids are for. Names remain user-owned: the stored card keeps its name, so a scan-first card keeps its generated name until renamed once, after which every collector follows it. Proxmox remains outside the bridge (its API view carries no host MACs); the existing suffix protection still applies there, and the no-MAC-label case keeps its dedicated regression test. 18 new tests: resolver-level contracts (claim, enrich, ambiguity, cloned-VM refusal, kind mismatch, id-beats-MAC, spelling-independent matching, hand-written adoption), facts/mapper coverage, and both headline flows e2e through the real server. Proven live on this machine: scan-first created host-e64e5425, `rpk discover system` updated that same card to the sys identity with OS/cores/RAM, and a rescan reported "no changes". Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A guest's config already names the NIC MACs Proxmox assigned it — net0: virtio=BC:24:11:… for VMs, hwaddr=… for containers — in the very response the collector fetches for the OS and disks, so guests join the collector-unification bridge with zero extra API calls: the parser lifts every netN MAC (normalised by the shared ARP normaliser), the guest cards carry them as a macs label, and the resolver's existing rule does the rest. A VM found by a network sweep and the same guest reported by `rpk discover proxmox` are now one card, in either order, with the vmid identity winning over the scan's. Nodes stay outside the bridge (the API exposes no host MACs we read), and Proxmox-vs-agent-inside-the-guest remains two cards by design: vmid and machine-id are both agent-grade identities and MACs alone never unify those. 10 new tests: netN parsing across VM/container/dhcp configs and multi-NIC guests, non-MAC net lines yielding nothing, the macs label on mapped guest cards, and the scan↔proxmox unification e2e through the real server including the rescan round-trip. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The WebUI job on PR #339 failed on YamlImportTests: the "Apply" button stayed disabled for the full 15s timeout. Same prerender race b8f6d23 fixed for the add form — YamlImportPom.GotoAsync navigated and returned as soon as the textarea was visible, so a FillAsync that landed before the circuit attached was lost to the static prerendered HTML. The oninput diff then never ran and Apply never enabled. CI lost this race intermittently; a fast machine won it, which is why it didn't reproduce on every run. GotoAsync now waits for MainLayout's data-circuit-ready probe (the same signal the add form waits on) before returning, so every subsequent PasteAsync reaches a live circuit. Added Importing_Works_Before_The_Circuit_Has_Warmed_Up, mirroring AddResourceRaceTests: 300ms injected SignalR latency widens the attach window so the race loses every time. Verified it fails (16s timeout, matching CI) with the wait removed and passes with it in place. Not an MCP regression — a pre-existing flake in a test added on staging (97e7efd) that this PR's CI run happened to trigger. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Feature/network discovery
Timmoth
marked this pull request as draft
September 27, 2026 12:08
…ease guide
- Bump version to 2.2.0 in RpkConstants, AssemblyVersion, and the README badge
- Fix generate-docs.sh and the generate-docs workflow to write the real
cli-commands{,-index}.md files with /docs/cli-commands#... links that
resolve in the web docs viewer; regenerate the index
- Delete stale pre-rename docs/Commands.md + docs/CommandIndex.md
- Deduplicate versioning.md and correct the nightly branch (staging, not main)
- Refresh overview.md/README prose (drop beta wording, mention discovery + MCP)
- Update install-guide.md download URLs from the 0.0.3 release to 2.2.0
- Update publish workflow dispatch defaults to 2.2.0 / v2.2.0
- Add docs/development/release-guide.md documenting the end-to-end release flow
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ntly Two tests fail on staging, pinning the load-side contract of the issue: - A config truncated at a resource boundary (simulating an interrupted in-place save) is served as a plausible smaller inventory with no diagnostic — rpk summary shows 2 of 3 servers, exit 0. - A config cut mid-token is swallowed entirely: rpk summary reports an EMPTY inventory with exit 0 and no error naming the file. This is quieter than the YamlDotNet stack trace the issue observed on 2.x — the boot-load catch now hides the failure, so the next write would persist the empty inventory over the damaged-but-recoverable file. The writer-side fix (atomic + durable saves in PhysicalTextFileStore: temp file, flush, rename) is not black-box observable; these tests cover the belt-and-braces load-side behaviour that must accompany it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The guide told users to chown 1000:1000, but the image runs as 1654:1654 (APP_UID) — the exact confusion reported in #210. README was already corrected; this aligns the served install guide with it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Saving config.yaml was neither atomic nor durable, and a damaged file was then accepted without complaint on the next load. Together those could lose an inventory. Writes went through File.WriteAllTextAsync, which truncates the file to zero before writing a byte and never flushes — so an interrupted save could leave a truncated config, and a save that had returned successfully could still be lost to power failure. Migration backups shared that path, so the one existing safety net was written unsafely too. The store now writes a temp file, flushes it to disk, and renames over the destination; a failed write cleans up and leaves the original untouched. Reads were worse than the issue reported: on staging an unparseable config loaded as an EMPTY inventory with exit 0, because the boot-time catch for unreadable stores was swallowing parse failures too. A ConfigLoadException is now latched on the collection when the config exists but cannot be understood, and every read and write refuses while it is set — the CLI reports "Config error:" with exit 5, MCP forwards the message. Two cases are detected: the file fails to parse, and the file parses with no schema version (truncated before the version line, or not a RackPeek config). An empty file is still a legitimately empty inventory, and legacy versionless configs still migrate. Routes.razor gated the whole app on a successful load, so a damaged config left the Web UI stuck on "Loading…" — including the YAML editor that is how you fix it. Both hosts now render anyway, and the editor reports a still-broken save inline instead of tearing down the circuit. A save interrupted at a clean resource boundary leaves valid YAML that is indistinguishable from a smaller inventory; nothing at load time can detect that, which is why the atomic-write half is the primary remedy. Tests: 5 store tests (the concurrency one fails on the parent commit with a 12288-byte partial read), 6 CLI e2e tests, 2 Playwright tests covering repair through the YAML editor. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Make config saves atomic and refuse to read a damaged config (#337)
A guest on a hypervisor on a server read
nebula (server) / nebula-pve (system) / immich (system)
so the chain that exists to show the nesting hid what each layer actually
is. A System now reports its own type where it has one:
nebula (Server) / nebula-pve (Hypervisor) / immich (VM)
Kinds and types are stored lower-case, so they are title-cased for display
with the initialisms that would otherwise look wrong (VM, UPS) spelled
properly. The single-crumb view lower-cased its label separately; it now
shares the same description.
The Playwright test seeds a Server -> hypervisor -> vm chain and fails on
the parent commit.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Show a system's type in the breadcrumb, not its storage kind
The comment illustrated the bug with host names from a real network. The example reads the same with invented ones, and test fixtures should never carry identifiers from anyone's actual infrastructure. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A sweep of a multi-VLAN homelab produced 26 anonymous "host-<hash>" cards out of 27: ARP is link-local so remote subnets yield no MAC, and few home networks have PTR records. The sweep found everything and recognised nothing. Three sources close most of that gap without a new dependency. Service banners. Once a host is known alive it is asked what it is: a TLS certificate's common name (443/8006/8443), an HTTP page title or Server header, then an SSH greeting. This is the identification half of `nmap -sV` in about 150 lines of BCL sockets. Only open ports are asked — a handshake with a closed port buys nothing but a timeout. The three answers are different claims, so they are treated differently. An X.509 common name is a host name by construction, so a certificate names the machine. A page title names an application, and a host may run several, so each becomes a Service resource hanging off the card rather than renaming it. An SSH greeting names only the daemon — every Linux box on a subnet answers "OpenSSH" — so it annotates and never names. An appliance's own management page is not a service on itself, so a title matching the host's own name is skipped. Open ports. Liveness stops at the first answer, which left pingable hosts with no port evidence at all. A living host is now checked against a wider list (554 cameras, 1883 brokers, 8123, 9000, 3000, 32400 and friends) and the result lands in an open-ports label. These are observations, not conclusions: "554 is open" is a fact, "this is a camera" is an inference the reader is better placed to draw, and a port number is a convention rather than a guarantee — so nothing is named from one. They are, however, the best possible targets for the banner probes. MAC vendor. Where a MAC is known the card gains the organisation IEEE assigned that OUI to, from a curated 10,868-entry subset of the registry generated by generate-oui-table.py. A hypervisor's own prefix beats the locally-administered bit so a KVM guest reads as QEMU/KVM rather than anonymous, while an address a phone invented for itself is reported as randomised instead of attributed to whoever owns the block. Services are also anchored by address: one whose runsOn resolves to nothing takes the system at its own IP, which fixes the dangling link a docker collector leaves when it can only guess its host's name. Narrow by design — only an unresolvable link, only when exactly one system claims that address, and never for systems, where a shared address means the same machine rather than a parent. Cards stay sparse and identity stays put: nothing learned here writes an OS or a core count, and a name from a service never moves a discoveryId. --no-identify restores a pure liveness sweep. All names, addresses and MAC suffixes in tests and docs are invented; the OUI prefixes are public IEEE registry data. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Proxmox only records a guest's address when someone set one statically, so on a DHCP estate every guest arrived with no address at all. That cost more than the field: a network sweep of another subnet has no ARP entry to work from, so it identifies a host by address alone — and with the guests carrying no address there was nothing to match, leaving a second, emptier card beside every guest the hypervisor had already described in full. Guests are now asked where they are. A VM answers through its qemu-guest-agent, a container through its running interfaces; a stopped guest is never asked, since it has nothing to report and the call is one round trip per guest. Choosing among the answers is the hard part. A guest that runs containers reports several interfaces — docker0 and its per-network bridges, Home Assistant's hassio, any VPN tunnel — and recording 172.17.0.1 as the machine's address would be worse than recording nothing, because every container host on the estate reports the same one. The NIC MACs Proxmox assigned are the discriminator: an interface carrying one is a NIC the hypervisor gave the guest, anything else the guest invented. Where the MACs cannot be read, nothing is claimed. A statically configured address still wins, being what the administrator asked for. With addresses in hand a sweep's find can be matched to the guest it actually is, so the resolver gains an address bridge beside the MAC one. Narrow on purpose: only a scan-grade card that produced no MAC of its own, only against an agent-grade card of the same kind, and only when exactly one stored system claims that address. Services anchor to the stored card in preference to a sweep's stand-in for the same reason. On a live two-node estate this turned nine scanned cards on the guest subnet into six guests updated in place and two genuinely unknown hosts, with the discovered web services landing on the real VMs rather than on stand-ins. The read orchestration moves into the domain on the way past. The CLI and the MCP tool each had a copy, and the copies had already drifted — the MCP one dropped the guests' MACs, silently costing every guest its chance of unifying with a scan. Now there is one, and it is tested. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A Proxmox number written as a string took the whole run down. Its perl backend quotes numeric fields inconsistently — "vmid":"100" and "maxdisk":"512110190592" both turn up across versions — and the TryGetInt32 family does not return false for a string, it throws. The kind is now checked first and a quoted number parsed, which is what the caller meant either way. Before this, one such field produced "Unexpected error occurred" plus a stack trace on the CLI, and MCP forwarded the BCL's "requires an element of type 'Number'" as though the user had done something wrong. An ssh:// or npipe:// DOCKER_HOST crashed. `docker context` sets the first for a remote host and the second is the Windows default, so both are ordinary values to find in the environment. HttpClient accepts either URI and only throws NotSupportedException on the first request, which was past the command's catch list. They are now refused where the client is built, through the UriFormatException both front ends already turn into "not a usable Docker endpoint" — one phrasing for a bad endpoint rather than two — and the message says how to forward the socket instead. A remote engine with no IPv4 address had its services recorded at this machine's address. Services are recorded at their host's address and the inventory holds IPv4 only, so when the endpoint resolves to IPv6 alone the address is genuinely unknown; substituting whatever machine ran the command gave every service a confident, wrong address that then flowed into the ansible, ssh and hosts exports. There is no fallback now, and the run stops with an explanation. Each fix has tests that fail without it: quoted numbers across the guest list, both unreachable endpoint schemes, and an engine answering on an IPv6-only socket. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The collector for everything a sweep can see but not identify. ARP is link-local. Sweeping from one machine yields a MAC for that machine's own segment and nothing but an address for every other subnet — and an address alone cannot survive a DHCP re-lease or be matched against anything already documented. The firewall routes every subnet, so its neighbour table carries the MAC for all of them, the name it handed out, and which leg each machine answered on. Cards are seeded exactly as the network sweep seeds its own, keyed on the MAC, because both describe the same thing by the same evidence: a machine observed on the network rather than asked about itself. So a host the firewall knows and a host a sweep found are one card, whichever ran first, with no special case anywhere to say so. On a live estate this turned 21 resources into 13 updated in place and 8 machines no sweep had ever seen — devices that answer no port and no ping but sit in the firewall's table. What it leaves out, on purpose: the firewall's own addresses (every routed subnet contributes one and they are all the same box, which is a Firewall rather than the handful of Systems this would invent), entries that have aged out, broadcast and multicast, and neighbours on public addresses. That last one is the ISP's equipment on the WAN leg — not the user's infrastructure, and recording it would put a public address into a file people commit. --include-public asks for them. Tolerant of the endpoint rename in OPNsense 25.7, which moved these actions from camelCase to snake_case and broke integrations pinned to either spelling; both are tried. Both response shapes are read too, since the search wrapper wraps the same rows in an object. A key that lacks the Diagnostics: ARP Table privilege gets a redirect to the login page rather than a 401, so that is reported as the permission problem it is. Verified against a live two-firewall estate; the fixtures and tests use invented addresses and MAC suffixes throughout. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
An open port was landing in an `open-ports` label: a comma-separated string nothing could link to, filter on, or hang a note from. RackPeek already has a resource for "a thing listening on an address and a port", so each open port now becomes a Service with its own stable id, running on the host that serves it — nebula-ssh, nebula-https, nebula-proxmox. What a service said about itself still beats what its port number implies, because a port is a convention and an answer is evidence. Where nothing answered the name falls back to the host plus the port's usual service, and an unrecognised port keeps its number as nebula-tcp-9987 rather than guessing. An appliance whose management page only says its own name back is named for the port instead, so a firewall called opnsense gains an opnsense-https rather than a second card called opnsense. Two labels change with it. `vendor` becomes `nic-vendor`, because the OUI identifies whoever owns the network interface, which is not the same claim as who made the machine — a Proxmox guest's virtual NIC reads Proxmox while the box underneath it is a Dell. And `segment` goes: it described the firewall's wiring rather than the machine, and changed whenever anything was re-cabled. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A collector names a card from whatever it could see, and when it could see nothing the name falls back to a slug of the card's own id — host-1a2b3c4d says only that something is there. A later run, or a collector that can see more, often does know the machine's name: a firewall knows what it handed out over DHCP, a hypervisor knows what its guest is called. Those placeholders now give way to real names, and everything pointing at the old name follows: stored runsOn, stored connections, services named after their host, and the incoming payload's own references. A new `userNamed` flag draws the line discovery must not cross. Renaming sets it, every rename path goes through one use case, and nothing in discovery touches the name again. A resource with no discoveryId is user-named whatever the flag says — nothing but a person could have written it. The upgrade is one-way, placeholder to real, and one real name never replaces another, or two collectors that each knew a different name would rename a box back and forth on every run. The address bridge is now symmetric. It fired only when the incoming card was the scan and the stored one was agent-grade, so sweeping before running the hypervisor left both cards behind where the other order produced one. Which collector ran first is an accident of what someone typed and must not decide what the inventory holds. Two scan cards can bridge as well, but only when exactly one of them saw a MAC: a firewall's neighbour table names the interface answering at an address where a sweep of a subnet it does not sit on only knows something replied, and those are not two stand-ins. Cards of equal standing never bridge. Measured on a nine-subnet lab, sweeping first: fourteen machines were in the inventory twice, and are now in it once. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two bugs, one cause: three places each decided for themselves what a service's URL was, and all three were wrong. The card derived every link as http:// regardless of the port, so SSH on 22 rendered as http://10.0.0.5:22/ — a link that looks real, invites a click and cannot load. It was reading Network.Protocol as though it named a scheme, but discovery writes the transport there (TCP), which says nothing about what rides on top of it. The dependency trees had it worse: they fed NetworkString() into an href, and that was display text — "Ip: 10.0.0.5:3000", prefix and trailing space included — so the link was dead on arrival. Both rules now live in one place. A scheme is only inferred where the convention is genuinely a web one; ssh, smb, mqtt, dns and anything uncurated get no link at all, because a dead link is worse than none. A URL someone typed by hand still always wins, and an explicit http or https in the protocol field is still believed, which is the only thing that can answer on an uncurated port. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ices Registering RackPeek's services writes process-wide statics as a side effect, RpkConstants.HasGitServices among them. That is harmless in production, where a process registers once, and a race in a test run, where dozens of classes register with different configuration. Three groups were mutating it from three different xUnit collections — the Git tests, the YAML CLI host, and the API tests — and collections run in parallel with each other, so whichever ran last won. Git_Token_Set_Registers_... asserted the flag was true and read whatever an API test had just set, failing roughly one full-suite run in three while passing in isolation every time. They now share one collection with parallelisation off, named for the shared state rather than for the YAML CLI, because membership is decided by "does this touch the statics" rather than by what the test nominally exercises. Six consecutive full runs green, against a baseline that failed one run in two. Costs about three seconds on a fifteen-second suite. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
BrowsableUrl built its result with UriBuilder, which throws UriFormatException on a host it cannot parse — "a/b", a name with spaces, a bare run of dots. The address field holds whatever a person typed or a collector read off a device, and neither is obliged to produce something a URL can be built from. That throw was already latent on the service card, where it would have spoiled one card. It matters more now: the dependency trees call this for every service they render, so a single malformed address would blank the entire hardware or system page instead of costing one link. A missing link costs a click. Letting this escape costs the page. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Fix the Git service test racing every other test that registers services
The conflict check matched the resource being renamed, so a case-only rename such as Srv01 -> srv01 failed with "already exists". Only a different resource with the new name is now a conflict, so renaming onto another resource's case variant is still refused, including in a config that already holds both variants. Renaming a resource to exactly its current name now succeeds without writing anything, instead of failing with "already exists".
Discovery: services from ports, self-improving names, and symmetric identity
Allow renaming a resource to a different case of its own name (#328)
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.