Stabilize live plugin validation - #59
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Executive Summary
This PR makes Claude Code live validation strict and repeatable across voice and host-bootstrap paths.
Description
Voice scenarios now sweep interrupted matching calls, snapshot both owners, apply a request-time lower bound, and require one stable driver/AUT pair before validation. Driver records prove scripted caller speech, while the exact AUT record must prove two-way conversation and agent-local output; teardown ends every post-snapshot matching record on both owners.
The hosted scenario hashes the full workflow run identity into three distinct speech-safe words and keeps its current-request, open-action, exact-recipient/body, receipt, and exactly-once checks. Dependency installation uses a bounded setup-only retry, while live messages, calls, model turns, and side effects remain single-attempt. Failure diagnostics expose only bounded state, counts, and error classes.
Reason
Recent live runs repeatedly evaluated two-way speech on the non-authoritative call record and sometimes lost a long spoken marker during hosted action extraction. Setup network reads and raw failure artifacts also made CI less stable and less safe to inspect publicly.
Decisions
Testing
.venv/bin/python -m pytest -q— 453 passed, 23 expected live-environment skips..venv/bin/python -m compileall -q inkbox_claude tests— passed.uvx ruff check --select F ...— passed.bash -n tests/ci/retry_install.shpassed, andgit diff --checkpassed.