Skip to content

fix(lima): reap orphaned host agents on stop - #50

Merged
olamide226 merged 2 commits into
mainfrom
fix/lima-orphaned-host-agents
Aug 7, 2026
Merged

olamide226 merged 2 commits into
mainfrom
fix/lima-orphaned-host-agents

Conversation

@olamide226

@olamide226 olamide226 commented Aug 7, 2026 •

Copy link
Copy Markdown
Owner

Summary

Fixes the lifecycle gap where Lima reports an Avar environment as stopped while orphaned limactl hostagent processes continue consuming CPU. Both avr stop and avr stop --all now converge already-stopped environments through the provider, which removes only exact host-agent matches for the owned machine and uses a bounded TERM-to-KILL escalation for wedged agents.

A real-Lima E2E run also exposed that a spinning host agent can keep limactl stop itself alive indefinitely, preventing orphan cleanup from running. Graceful shutdown is now bounded at 30 seconds before Avar enters its existing warned force-stop path.

Regression coverage locks in the previously missed command path, the stopped-state provider cleanup call, exact process matching, and the hung graceful-shutdown path.

Spec task: tasks.md task 10 — Implement avr status and avr stop (corrective bug fix)

Requirements covered

Requirement Acceptance criterion How this PR satisfies it
REQ-5.2 When a user runs avr stop, stop the selected machine; avr stop --all stops all Avar-managed machines. Stop now converges both the reported VM state and stale Lima host-agent processes. A hung graceful shutdown is bounded and escalated instead of blocking forever.
REQ-5.4 Avar only manages machines it created and never modifies or stops other Lima machines. Cleanup remains behind the ownership gate and matches the exact limactl executable plus the exact Avar machine argument.

Correctness properties

Property Where it is proven
PROP-6 TestParseOrphanHostAgentPIDs_MatchesOnlyTheExactLimaMachine, existing provider ownership tests, and TestStop_AllStopsEveryOwnedMachineAndNothingElse_REQ_5_2

Deliberately not covered

None. This is a corrective lifecycle fix for the already-completed stop task; no tasks.md checkbox changes are needed.

Verification

  • make lint
  • make tidy-check
  • make release-version-test
  • make build
  • make test
  • make e2e — passed against real Lima on the affected Mac: ok github.com/olamide226/avar/e2e 390.719s
  • Post-E2E ./bin/avr stop --all — stopped three running test environments; all six Lima environments report Stopped and no Avar limactl hostagent process remains.
  • git diff --check

Notes for review

The first real-Lima E2E attempt reproduced a second part of the failure: limactl stop avr-ubuntu-24.04-arm64 remained alive for more than four minutes while a child host agent consumed roughly 95% CPU. That run was interrupted, the three test VMs were force-stopped directly through Lima, and commit 2750179 added the bounded graceful-stop path plus TestStop_EndsAMachineWhenGracefulShutdownHangs_REQ_5_2. The complete E2E rerun then passed.

Lima 2.2.0 can lose track of stale host agents once its pid/socket files have been replaced or removed; another limactl stop --force then reports the agent as already stopped and does not reap those processes. The cleanup therefore inspects the macOS process table after provider stop convergence. It sends TERM first, waits 200 ms, and uses KILL only for matching agents that remain. Orphan cleanup is bounded by two seconds.

Lima can report an environment as stopped while stale hostagent processes continue consuming CPU. Converge stopped machines through the provider, terminate only exact Avar-owned hostagent matches, and escalate briefly when an agent does not exit cleanly.\n\nRegression tests cover the command path, stopped-state provider cleanup, and exact process matching.\n\nRequirements: REQ-5.2, REQ-5.4
A real-Lima E2E run showed that a spinning hostagent can keep limactl stop alive indefinitely, preventing the orphan cleanup from running. Bound graceful shutdown to 30 seconds, preserve caller cancellation, and escalate through the existing warned force-stop path when the internal deadline expires.\n\nRequirements: REQ-5.2
@olamide226
olamide226 merged commit b8e3e84 into main Aug 7, 2026
2 checks passed
@olamide226
olamide226 deleted the fix/lima-orphaned-host-agents branch August 7, 2026 13:46
olamide226 added a commit that referenced this pull request Sep 3, 2026
PR #50 made `avr stop` reap the orphaned Lima host agents a stopped
instance can leave behind. The detector is right and was earned from a
real leak; three defects sit around it, found by reading the merged code
rather than by a failure.

signalPID treats /bin/kill's exit status as the verdict, and killing a
pid that has already exited exits 1. reapHostAgents returns that error
and stopMachine wraps it into a failed stop, so an agent exiting on its
own between the scan and the signal turns a successful stop into a
reported failure. The escalation pass makes it likely rather than rare:
it sends KILL to pids that received TERM 200ms earlier, which are
exactly the ones exiting.

avr stop is also the only command that passes types.DiscardProgress, at
all four of its call sites, while every other command routes progress to
stderr. The force-stop warning that unsaved work may be lost therefore
reaches nobody. And idle auto-stop skips machines that are not running,
which is the state a leaked agent leaves behind, so only an explicit
avr stop ever self-heals.

The task records the constraint on the obvious fix as well: this package
compiles on Windows and hostagent.go carries no build tag, so
syscall.Kill would reintroduce exactly what task 38 repaired.

One item is deliberately written as a question rather than a finding —
whether killing the agent takes an emulated instance's qemu process with
it is unverified, and the task says to check it on a real machine before
adding a predicate on suspicion.

Requirements: 5.2, 5.5, 1.5
olamide226 added a commit that referenced this pull request Sep 3, 2026
PR #50 made `avr stop` reap the orphaned Lima host agents a stopped
instance can leave behind. The detector is right and was earned from a
real leak; three defects sit around it, found by reading the merged code
rather than by a failure.

signalPID treats /bin/kill's exit status as the verdict, and killing a
pid that has already exited exits 1. reapHostAgents returns that error
and stopMachine wraps it into a failed stop, so an agent exiting on its
own between the scan and the signal turns a successful stop into a
reported failure. The escalation pass makes it likely rather than rare:
it sends KILL to pids that received TERM 200ms earlier, which are
exactly the ones exiting.

avr stop is also the only command that passes types.DiscardProgress, at
all four of its call sites, while every other command routes progress to
stderr. The force-stop warning that unsaved work may be lost therefore
reaches nobody. And idle auto-stop skips machines that are not running,
which is the state a leaked agent leaves behind, so only an explicit
avr stop ever self-heals.

The task records the constraint on the obvious fix as well: this package
compiles on Windows and hostagent.go carries no build tag, so
syscall.Kill would reintroduce exactly what task 38 repaired.

One item is deliberately written as a question rather than a finding —
whether killing the agent takes an emulated instance's qemu process with
it is unverified, and the task says to check it on a real machine before
adding a predicate on suspicion.

Requirements: 5.2, 5.5, 1.5
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant