Repository navigation
Production Benchmark Qualification & Reliability Hardening - #53
Merged
Merged
Conversation
Signed-off-by: 班扬 <xingjun.wxj@alibaba-inc.com>
Signed-off-by: 班扬 <xingjun.wxj@alibaba-inc.com>
Signed-off-by: 班扬 <xingjun.wxj@alibaba-inc.com>
…tions Natural-language input is only parsed as a command with an explicit leading slash; /host declares its actions once for usage, completion, and both runtimes; the host view reports health, failures, and recovery steps. Signed-off-by: 班扬 <xingjun.wxj@alibaba-inc.com>
Signed-off-by: 班扬 <xingjun.wxj@alibaba-inc.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Features
leap-bench): New standalone package — manifests, registry, async runner, unified metrics, JSONL evidence, and ready/conditional/blocked gates over tier0–tier4 (plus live-llm, production-sim, pre-hardware profiles), with aleap-benchCLI (list/doctor/run/resume/report/gate) and abenchmarkinstall extra.preflight_mode: zero_motion, per-device required checks, and a structured report assertingmotion_performed=false; any motion claim or missing evidence fails closed. Every run is isolated underbenchmarks/runs/<run_id>/<adapter>/, replacing temp output.DegradationCoordinatoris the single execution point for declared policies (halt/hold_position/safe_return/switch_sensor/continue_blind) across stream and health-monitor paths; latched states block writes until verified recovery or manual clear, with newhome_joint_targets,sensor_fallbacks, and boundedsafe_return_timeout_s.FailureFingerprintnormalizes volatile output so repeated identical failures terminate instead of being retried blindly.pid+window_id, malformed discovery is rejected, and bounded recovery plus a dependency circuit stop repeated RPCs on a known-bad driver — with MCP calls marshaled onto the loop that owns the stdio session./hostactions: Free text is only parsed as a command with a leading slash;/hostdeclares actions once for usage, completion, and both runtimes, reporting driver health, failures, and recovery steps.Enhancements
shell_runreturns structured failure codes with non-sensitive audit metadata; recovery classifies producer-declared terminal codes into a new non-recoverabletool_terminalcategory before text heuristics.src/benchmarks/tests/benchmarks; Tier0/Tier1 smoke runs wired into CI; manifests ship as package data; 14 new benchmark test modules plus coverage for slash routing, tool hardening, hardware degradation, and CUA mapping.Fixes
Refactor
DegradationCoordinator(no more forked halt/hold/safe-return logic); repetition guard switched to normalized structured fingerprints.