Repository navigation
fix(antigravity): read turn boundaries from the transcript, not file mtime - #592
Merged
garysheng merged 8 commits intoOct 4, 2026
Merged
Conversation
|
@JackFurton is attempting to deploy a commit to the Gary Sheng's projects Team on Vercel. A member of the Team first needs to authorize it. |
JackFurton
force-pushed
the
fix/antigravity-transcript-events
branch
5 times, most recently
from
September 6, 2026 00:13
189b44c to
775ea0a
Compare
…mtime Every tool call played "Job's done!" then "Right away!". Mtime inference read a long tool call as completion, and the write that followed it as a new prompt. Boundaries now come from transcript.jsonl records, and idle timeouts are off for sessions that have one. Sessions without a transcript keep the mtime fallback at 45s. Two things only real transcripts showed. A PLANNER_RESPONSE with no tool_calls is a turn end only if it also has prose: Antigravity emits empty ones beside every ERROR_MESSAGE while retrying a rate limit, giving 18 Stops in an 80-record session against 6 real ones. And transcript_full.jsonl is byte-identical to transcript.jsonl, so the transcript* glob would have doubled every event. Replaying 33 transcripts gives 124 acknowledgements for 124 prompts. Permission prompts needed four fixes that only a running machine exposed. cli.log is a symlink into log/cli-<timestamp>.log, retargeted at every launch, so watching the symlink's own directory never sees a write. Antigravity holds the log open and flushes in batches, so filesystem events arrived minutes late or never; the log is polled once a second instead, and is polled rather than observed so one offset has one reader. A flushed batch usually contains prompts already answered, so only a batch ending still unanswered sounds, once: replaying each line delivered a burst of notifications for prompts answered minutes earlier. Offsets are keyed by resolved path, the same file otherwise arriving under two names. Notifications named the wrong thing entirely. The watcher is a daemon, so its working directory is "/", and it passed that as every event's cwd; peon.sh cannot make a project name out of that and falls back to labelling anything non-Codex "claude", so Antigravity's banners claimed to be Claude Code's. The workspace now comes from the conversation's own database, which holds it as a file:// URI, with the last-conversation cache as a fallback. The database is what makes a resumed session resolve, the cache knowing only the newest conversation per workspace. Across 46 real conversations that covers 37; the rest are default-cli-project sessions with no workspace recorded anywhere, and those fall back to home rather than to the "claude" label. Command failures play task.error, which needed tool_name and error plumbed through the watcher protocol since that is what peon.sh gates on. Permission prompts are attributed to the most recently active conversation, the surfacing line carrying no conversation id, seeded from the newest transcript at startup. Two concurrent sessions can misattribute one, worst case a sound against the wrong tab. resource.limit stays unreachable, peon.sh routing it only from PreCompact, so rate limits land on task.error. Subagent events are unwired: peon.sh tracks subagents by session-id relationships the watcher does not model. Also: gemini.sh and gemini.ps1 mapped AfterTool exit 0 to Stop, and both the bats and Pester tests asserting that are updated; install.sh never shipped antigravity-watcher.py, so the adapter died on its own preflight; peon status --verbose now lists Antigravity; the watcher's stderr appends to the adapter log rather than truncating it, the LaunchAgent having the same file open. The dispatch loop in antigravity-py.sh reads tab-delimited fields on one line. Field-per-line aborts under set -e, because $() strips trailing newlines and the read for an absent tool_name hits EOF.
JackFurton
force-pushed
the
fix/antigravity-transcript-events
branch
from
September 6, 2026 00:32
775ea0a to
82957d6
Compare
Legacy mtime completions bypassed the per-conversation cwd resolver, so the prompt named the agent workspace and the final notification named the daemon workspace. Emit Stop through the same path as transcript events. Add a real emitted-payload regression, sync the Japanese adapter description, and correct the status comment to acknowledge upstream native hooks. Validation: reproduced failing regression before the change; 74 Python tests and 21 focused Antigravity/Gemini BATS pass. Native watchdog integration emits prompt, permission, error, and completion with the resolved cwd and stays silent during a tool pause. Local clone install copies the Python watcher byte for byte. Bash syntax and diff whitespace checks pass. Human prose updates require no test.
Read ToolResult.error from the native tool_response object in both adapters. Prefer its message over the legacy top-level exit_code/stderr contract, which remains supported for compatibility with existing inputs. Successful AfterTool calls stay silent because AfterAgent owns completion. Build Bash payloads from parsed JSON rather than interpolated Python source so quoted session ids and project paths survive the adapter boundary. Regression tests execute the adapters with native failure, native success, legacy input, message precedence, and empty-error payloads. The Bash suite also verifies the actual task.error sound through peon.sh. Synchronize the English, Chinese, and Japanese event mapping documentation. Validation: 9 Bash tests, 9 functional PowerShell tests using portable PowerShell 7.6.6, 12 outside-repository CLI probes, shell syntax, ShellCheck, Python quoting lint, PowerShell parser, and git diff --check. Windows PowerShell 5.1 execution remains for Windows CI.
The approved maintenance batch excludes releases. Preserve the feature notes under Unreleased and retain the maintained 2.37.0 version until a separately prepared containing release. These are metadata-only changes; the VERSION bytes match maintained main and the remaining changelog content is unchanged.
Pure documentation change. Checked the native error and silent-success descriptions against both adapter behavior and the three synchronized README sections.
…tegration-20261004
Retain all Antigravity changes and the separately approved native Gemini mappings and regressions. The four conflict files match Gemini candidate 1d2b741. Candidate whitespace check against maintained main passes; two inherited blank lines present in maintained main were flagged only by the staged merge-parent diff.
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Every tool call played "Job's done!" then "Right away!". Mtime inference read a long tool call as completion, and the write that followed it as a new prompt.
Boundaries now come from
transcript.jsonlrecords:USER_INPUT/USER_EXPLICITis a prompt,PLANNER_RESPONSEwithtool_callsis work in flight, and one with prose and none is the final answer. Idle timeouts are off for sessions that have a transcript. Sessions without one keep the mtime fallback, now 45s.Two things only real transcripts showed:
PLANNER_RESPONSEwith notool_callsis a turn end only if it also has prose. Antigravity emits empty ones beside everyERROR_MESSAGEwhile retrying a rate limit, giving 18 Stops in an 80-record session against 6 real ones.transcript_full.jsonlis byte-identical totranscript.jsonl, so thetranscript*glob would have doubled every event once records drove them.Replaying 33 transcripts gives 124 acknowledgements for 124 prompts.
Permission prompts
Four fixes, none of which unit tests could have found. Running this against live Antigravity is what surfaced them:
cli.logis a symlink intolog/cli-<timestamp>.log, retargeted at every launch, so watching the symlink's own directory never sees a write. Prompts never fired at all.Surfacingline delivered a burst of notifications for prompts answered minutes earlier.Attribution is the most recently active conversation, the surfacing line carrying no conversation id, seeded from the newest transcript at startup so a watcher started mid-session is not silent until the user types. Two concurrent sessions can misattribute one, worst case a sound against the wrong tab.
Notifications claimed to be Claude Code
The watcher is a daemon, so its working directory is
/, and it passed that as every event'scwd. peon.sh cannot make a project name out of/and falls back to labelling anything non-Codexclaude, so every Antigravity banner was titled that, with no way to tell which workspace it came from.The workspace now comes from the conversation's own
conversations/<guid>.db, which stores it as afile://URI intrajectory_metadata_blob, falling back tocache/last_conversations.json. Reading the database is what makes a resumed session resolve: the cache holds only the newest conversation per workspace, so anything older missed. Across 46 real conversations this resolves 37. The 9 it does not aredefault-cli-projectsessions started outside a workspace, which record no path anywhere; those fall back to home rather than to the misleadingclaudelabel.Claude Code has no equivalent problem because it passes
cwdin the hook payload. Antigravity has no hook API, so the watcher has to recover it.The rest
Command failures play
task.error, which neededtool_nameanderrorplumbed through the watcher protocol since that is what peon.sh gates on.resource.limitstays unreachable, peon.sh routing it only fromPreCompact, so rate limits land ontask.error. Subagent events are unwired: peon.sh tracks subagents by session-id relationships the watcher does not model, andINVOKE_SUBAGENTappears once across 33 transcripts. Both felt like the right call, but I would take either on if you disagree.Also:
gemini.shandgemini.ps1mappedAfterToolexit 0 toStop, with the bats and Pester tests asserting that updated;install.shnever shippedantigravity-watcher.py, so the adapter died on its own preflight;peon status --verbosenow lists Antigravity; the watcher's stderr appends to the adapter log rather than truncating it, the LaunchAgent having the same file open.The dispatch loop in
antigravity-py.shis a function so it can be tested, reading tab-delimited fields on one line. Field-per-line aborts underset -e, because$()strips trailing newlines and the read for an absenttool_namehits EOF.Verified on a real machine
Installed on macOS as a LaunchAgent against live Antigravity. A turn with one failing command:
A flushed batch of three prompts where two were already answered fires exactly one
PermissionRequest; the same batch arriving while one is already pending fires none.Those 6 pass on CI and fail only on my machine, where a real
~/.codexand terminal state leak into them.Version bumped to 2.38.0 with a CHANGELOG entry. No tag pushed.