RELOPS-2402: fleetbench PSU/firmware-throttle detector (hardware fleet) - #1401
Merged
Merged
Conversation
Split out of PR #1235 (combined RELOPS-2396 branch) into fleetbench-only. Per NUC13 hardware node, before worker-runner starts: - win_fleetbench module installs the version-pinned collector to C:\fleetbench plus per-hardware baselines, wired via the hardware_observability profile. - maintainsystem-hw.ps1 runs `fleetbench cpu --mode quick --duration 900s --json` once post-bootstrap then at most once per 72h; 900s self-warms the node so PSU/thermal throttling surfaces (a short cold-boot run can false-pass). - Evaluates GOOD/BAD/MARGINAL/UNKNOWN vs the locked per-hardware baseline, plus drift vs the node's first recorded run. - NSClient++ checks `fleetbench` and `fleetbench_variance` surface verdict + metrics to Marlin (Icinga2/Grafana). Hardware-only: gated to the datacenter maintain-system path + hardware_observability. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Invoke-FleetbenchCheck runs `fleetbench cpu --duration 900s` (~15 min) before worker-runner starts, so generic-worker is intentionally not up during the run. The hourly gw_exe_check task can fire mid-benchmark once uptime passes its 15-minute grace period, see no generic-worker process, and escalate to reboot/PXE reimage. maintainsystem-hw.ps1: set MOZ_FLEETBENCH_RUNNING (run's UTC start time) at Machine scope around the benchmark call, cleared in a finally so it is always removed even on error. gw_exe_check.ps1: read the marker live from the registry ([Environment]::GetEnvironmentVariable(..,'Machine'), not the possibly-stale process env block) and skip the check while it is fresh. The marker is a timestamp with a 30-min staleness cap, so a crash/reboot mid-run cannot silently disable the watchdog. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This was referenced Sep 17, 2026
The pin was v0.4.0, written when this work started on 2026-07-01; the collector has shipped through v0.4.6 since. Also verify the download against the checksum the release publishes in its own SHA256SUMS. archive's only guard was `creates => $binary`, so a truncated or corrupt download left a partial exe on disk that satisfied that guard on every later run - the collector would stay broken and puppet would never notice. checksum_verify makes that case fail loudly instead. The hash is pinned next to the version in hiera and the two move together. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
jwmossmoz
approved these changes
Sep 17, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Replaces #1263 and #1235, both of which were branched off stale bases. This is the same fleetbench work replayed onto current master (
47042c6b), so it can be tested on the alpha pools before it reaches production.Identical content to #1263: 11 files, +654/−0. Both of that PR's commits cherry-picked cleanly onto master —
maintainsystem-hw.ps1auto-merged against the changes #1297 and #1400 made to it, with no conflicts and nothing dropped.What it does
win_fleetbenchinstalls the pinned collector (fleetbench-v0.4.0-windows-x86_64.exe) from the fleetbench GitHub releases intoC:\fleetbench, plus therun_fleetbench.ps1wrapper and per-hardware-typefleetbench_baselines.json.roles_profiles::profiles::hardware_observability, which every Windows hw role already includes, so it lands on all hardware pools on a hash bump.win_nsclientgainscheck_fleetbench.ps1andcheck_fleetbench_variance.ps1plus the nsclient.ini entries, so results surface through the existing monitoring agent.maintainsystem-hw.ps1runs the benchmark, andgw_exe_check.ps1skips the generic-worker watchdog while it is running so the benchmark is not mistaken for a hung worker.Verified before opening
v0.4.0/fleetbench-v0.4.0-windows-x86_64.exereturns HTTP 200. Note the collector is now up to v0.4.6 — the version is one hiera line (windows.fleetbench.version) if we want the newer one for this test.win116424h2hwbakeexcludeshardware_observability, so nothing tries to install or run the collector inside the GPU-less bake guest. It is deploy-time only, which is what we want.Rollout plan
Alpha pools first — a worker-images PR points
win11-64-24h2-hw-alphaandwin11-64-24h2-hw-ref-alphaat this branch, and those nodes re-image and get watched. Production only after that, as a separate hash bump.🤖 Generated with Claude Code