CONFIRMED, two cases, shared root cause. In Tier-1 replay the agent produces requests and the cache is keyed on them; recorded request fields and timing metadata are never read back. So operators that mutate those can't reach the agent.
(a) ambiguous_user_turn forces a FALSE no_ship
operators.py:97-115 rewrites a recorded model_call prompt to "do the thing". At replay the agent regenerates its original prompt, whose cache key no longer exists → guaranteed ReplayMiss → faithfulness = 0.0 → any_zero rule forces no_ship. Verified: the bundled calc_agent (a fine agent) is blocked to no_ship solely by this scenario.
(b) inject_latency is an inert no-op
operators.py:80-94 only scales Step.latency_ms, which no replay/runtime path consumes and which the capture proxies overwrite at replay. Verified: all four metrics = 1.0, identical to baseline — a silent false PASS for the timeout_handling class.
Impact
(a) is a false negative that can block shipping real agents; (b) is a false positive that inflates the verdict. Only response-mutating operators (drop_tool_result, corrupt_field, prompt_injection) carry real signal.
Note
The fix changes what the reliability suite measures (verdict semantics) — a design decision. Options: mutate what the agent actually ingests (agent_input / a response it reads) instead of a recorded request; implement a replay-time delay/timeout hook for latency; or drop the claims until backed. Flagging for a founder call before any semantics-changing PR.
CONFIRMED, two cases, shared root cause. In Tier-1 replay the agent produces requests and the cache is keyed on them; recorded request fields and timing metadata are never read back. So operators that mutate those can't reach the agent.
(a)
ambiguous_user_turnforces a FALSEno_shipoperators.py:97-115 rewrites a recorded model_call prompt to
"do the thing". At replay the agent regenerates its original prompt, whose cache key no longer exists → guaranteedReplayMiss→faithfulness = 0.0→any_zerorule forcesno_ship. Verified: the bundledcalc_agent(a fine agent) is blocked tono_shipsolely by this scenario.(b)
inject_latencyis an inert no-opoperators.py:80-94 only scales
Step.latency_ms, which no replay/runtime path consumes and which the capture proxies overwrite at replay. Verified: all four metrics = 1.0, identical to baseline — a silent false PASS for thetimeout_handlingclass.Impact
(a) is a false negative that can block shipping real agents; (b) is a false positive that inflates the verdict. Only response-mutating operators (drop_tool_result, corrupt_field, prompt_injection) carry real signal.
Note
The fix changes what the reliability suite measures (verdict semantics) — a design decision. Options: mutate what the agent actually ingests (agent_input / a response it reads) instead of a recorded request; implement a replay-time delay/timeout hook for latency; or drop the claims until backed. Flagging for a founder call before any semantics-changing PR.