Skip to content

fix(antigravity): read turn boundaries from the transcript, not file mtime - #592

Merged
garysheng merged 8 commits into
PeonPing:mainfrom
JackFurton:fix/antigravity-transcript-events
Oct 4, 2026
Merged

garysheng merged 8 commits into
PeonPing:mainfrom
JackFurton:fix/antigravity-transcript-events

Conversation

@JackFurton

@JackFurton JackFurton commented Sep 5, 2026 •

Copy link
Copy Markdown
Contributor

Every tool call played "Job's done!" then "Right away!". Mtime inference read a long tool call as completion, and the write that followed it as a new prompt.

Boundaries now come from transcript.jsonl records: USER_INPUT/USER_EXPLICIT is a prompt, PLANNER_RESPONSE with tool_calls is work in flight, and one with prose and none is the final answer. Idle timeouts are off for sessions that have a transcript. Sessions without one keep the mtime fallback, now 45s.

Two things only real transcripts showed:

  • A PLANNER_RESPONSE with no tool_calls is a turn end only if it also has prose. Antigravity emits empty ones beside every ERROR_MESSAGE while retrying a rate limit, giving 18 Stops in an 80-record session against 6 real ones.
  • transcript_full.jsonl is byte-identical to transcript.jsonl, so the transcript* glob would have doubled every event once records drove them.

Replaying 33 transcripts gives 124 acknowledgements for 124 prompts.

Permission prompts

Four fixes, none of which unit tests could have found. Running this against live Antigravity is what surfaced them:

  • cli.log is a symlink into log/cli-<timestamp>.log, retargeted at every launch, so watching the symlink's own directory never sees a write. Prompts never fired at all.
  • Antigravity holds the log open and flushes in batches. Filesystem events arrived minutes late or never, so the log is polled once a second, and polled rather than observed so one offset has one reader.
  • A flushed batch usually contains prompts already answered. Only a batch that ends still unanswered sounds, and once. Replaying each Surfacing line delivered a burst of notifications for prompts answered minutes earlier.
  • Offsets are keyed by resolved path; the same file otherwise arrives under two names.

Attribution is the most recently active conversation, the surfacing line carrying no conversation id, seeded from the newest transcript at startup so a watcher started mid-session is not silent until the user types. Two concurrent sessions can misattribute one, worst case a sound against the wrong tab.

Notifications claimed to be Claude Code

The watcher is a daemon, so its working directory is /, and it passed that as every event's cwd. peon.sh cannot make a project name out of / and falls back to labelling anything non-Codex claude, so every Antigravity banner was titled that, with no way to tell which workspace it came from.

The workspace now comes from the conversation's own conversations/<guid>.db, which stores it as a file:// URI in trajectory_metadata_blob, falling back to cache/last_conversations.json. Reading the database is what makes a resumed session resolve: the cache holds only the newest conversation per workspace, so anything older missed. Across 46 real conversations this resolves 37. The 9 it does not are default-cli-project sessions started outside a workspace, which record no path anywhere; those fall back to home rather than to the misleading claude label.

Claude Code has no equivalent problem because it passes cwd in the hook payload. Antigravity has no hook API, so the watcher has to recover it.

The rest

Command failures play task.error, which needed tool_name and error plumbed through the watcher protocol since that is what peon.sh gates on.

resource.limit stays unreachable, peon.sh routing it only from PreCompact, so rate limits land on task.error. Subagent events are unwired: peon.sh tracks subagents by session-id relationships the watcher does not model, and INVOKE_SUBAGENT appears once across 33 transcripts. Both felt like the right call, but I would take either on if you disagree.

Also: gemini.sh and gemini.ps1 mapped AfterTool exit 0 to Stop, with the bats and Pester tests asserting that updated; install.sh never shipped antigravity-watcher.py, so the adapter died on its own preflight; peon status --verbose now lists Antigravity; the watcher's stderr appends to the adapter log rather than truncating it, the LaunchAgent having the same file open.

The dispatch loop in antigravity-py.sh is a function so it can be tested, reading tab-delimited fields on one line. Field-per-line aborts under set -e, because $() strips trailing newlines and the read for an absent tool_name hits EOF.

Verified on a real machine

Installed on macOS as a LaunchAgent against live Antigravity. A turn with one failing command:

New session: f227406a
User prompt: f227406a
Command failed (127): f227406a

A flushed batch of three prompts where two were already answered fires exactly one PermissionRequest; the same batch arriving while one is already pending fires none.

python3 -m unittest tests.test_antigravity_watcher            73 passed (was 14)
bats tests/antigravity.bats antigravity-py.bats gemini.bats   32 passed, 0 failed
bats tests/peon.bats                                         455 passed, 6 failed locally

Those 6 pass on CI and fail only on my machine, where a real ~/.codex and terminal state leak into them.

Version bumped to 2.38.0 with a CHANGELOG entry. No tag pushed.

@vercel

vercel Bot commented Sep 5, 2026

Copy link
Copy Markdown

@JackFurton is attempting to deploy a commit to the Gary Sheng's projects Team on Vercel.

A member of the Team first needs to authorize it.

@JackFurton
JackFurton force-pushed the fix/antigravity-transcript-events branch 5 times, most recently from 189b44c to 775ea0a Compare September 6, 2026 00:13
…mtime

Every tool call played "Job's done!" then "Right away!". Mtime inference read a
long tool call as completion, and the write that followed it as a new prompt.
Boundaries now come from transcript.jsonl records, and idle timeouts are off for
sessions that have one. Sessions without a transcript keep the mtime fallback at
45s.

Two things only real transcripts showed. A PLANNER_RESPONSE with no tool_calls
is a turn end only if it also has prose: Antigravity emits empty ones beside
every ERROR_MESSAGE while retrying a rate limit, giving 18 Stops in an 80-record
session against 6 real ones. And transcript_full.jsonl is byte-identical to
transcript.jsonl, so the transcript* glob would have doubled every event.
Replaying 33 transcripts gives 124 acknowledgements for 124 prompts.

Permission prompts needed four fixes that only a running machine exposed.
cli.log is a symlink into log/cli-<timestamp>.log, retargeted at every launch,
so watching the symlink's own directory never sees a write. Antigravity holds
the log open and flushes in batches, so filesystem events arrived minutes late
or never; the log is polled once a second instead, and is polled rather than
observed so one offset has one reader. A flushed batch usually contains prompts
already answered, so only a batch ending still unanswered sounds, once: replaying
each line delivered a burst of notifications for prompts answered minutes
earlier. Offsets are keyed by resolved path, the same file otherwise arriving
under two names.

Notifications named the wrong thing entirely. The watcher is a daemon, so its
working directory is "/", and it passed that as every event's cwd; peon.sh
cannot make a project name out of that and falls back to labelling anything
non-Codex "claude", so Antigravity's banners claimed to be Claude Code's. The
workspace now comes from the conversation's own database, which holds it as a
file:// URI, with the last-conversation cache as a fallback. The database is
what makes a resumed session resolve, the cache knowing only the newest
conversation per workspace. Across 46 real conversations that covers 37; the
rest are default-cli-project sessions with no workspace recorded anywhere, and
those fall back to home rather than to the "claude" label.

Command failures play task.error, which needed tool_name and error plumbed
through the watcher protocol since that is what peon.sh gates on. Permission
prompts are attributed to the most recently active conversation, the surfacing
line carrying no conversation id, seeded from the newest transcript at startup.
Two concurrent sessions can misattribute one, worst case a sound against the
wrong tab.

resource.limit stays unreachable, peon.sh routing it only from PreCompact, so
rate limits land on task.error. Subagent events are unwired: peon.sh tracks
subagents by session-id relationships the watcher does not model.

Also: gemini.sh and gemini.ps1 mapped AfterTool exit 0 to Stop, and both the
bats and Pester tests asserting that are updated; install.sh never shipped
antigravity-watcher.py, so the adapter died on its own preflight; peon status
--verbose now lists Antigravity; the watcher's stderr appends to the adapter log
rather than truncating it, the LaunchAgent having the same file open.

The dispatch loop in antigravity-py.sh reads tab-delimited fields on one line.
Field-per-line aborts under set -e, because $() strips trailing newlines and the
read for an absent tool_name hits EOF.
@JackFurton
JackFurton force-pushed the fix/antigravity-transcript-events branch from 775ea0a to 82957d6 Compare September 6, 2026 00:32
Local review only. Clean integration of 82957d6 onto 8ef3766. No remote mutation.
Legacy mtime completions bypassed the per-conversation cwd resolver, so the prompt named the agent workspace and the final notification named the daemon workspace. Emit Stop through the same path as transcript events. Add a real emitted-payload regression, sync the Japanese adapter description, and correct the status comment to acknowledge upstream native hooks.

Validation: reproduced failing regression before the change; 74 Python tests and 21 focused Antigravity/Gemini BATS pass. Native watchdog integration emits prompt, permission, error, and completion with the resolved cwd and stays silent during a tool pause. Local clone install copies the Python watcher byte for byte. Bash syntax and diff whitespace checks pass. Human prose updates require no test.
Read ToolResult.error from the native tool_response object in both
adapters. Prefer its message over the legacy top-level exit_code/stderr
contract, which remains supported for compatibility with existing inputs.
Successful AfterTool calls stay silent because AfterAgent owns completion.

Build Bash payloads from parsed JSON rather than interpolated Python source
so quoted session ids and project paths survive the adapter boundary.

Regression tests execute the adapters with native failure, native success,
legacy input, message precedence, and empty-error payloads. The Bash suite
also verifies the actual task.error sound through peon.sh. Synchronize the
English, Chinese, and Japanese event mapping documentation.

Validation: 9 Bash tests, 9 functional PowerShell tests using portable
PowerShell 7.6.6, 12 outside-repository CLI probes, shell syntax,
ShellCheck, Python quoting lint, PowerShell parser, and git diff --check.
Windows PowerShell 5.1 execution remains for Windows CI.
The approved maintenance batch excludes releases. Preserve the feature notes under Unreleased and retain the maintained 2.37.0 version until a separately prepared containing release. These are metadata-only changes; the VERSION bytes match maintained main and the remaining changelog content is unchanged.
Pure documentation change. Checked the native error and silent-success descriptions against both adapter behavior and the three synchronized README sections.
Retain all Antigravity changes and the separately approved native Gemini mappings and regressions. The four conflict files match Gemini candidate 1d2b741. Candidate whitespace check against maintained main passes; two inherited blank lines present in maintained main were flagged only by the staged merge-parent diff.
@garysheng
garysheng merged commit cc33dbe into PeonPing:main Oct 4, 2026
4 checks passed

This branch was successfully deployed

1 active deployment
Preview — 82df9b49 Deployed Oct 4, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants