Skip to content

Feature: host-aware resource budget for seats (build parallelism, priority, transcript capture cost) #80

Description

@kainne44

Summary

OpenRig knows how many seats are live on a host, but seats get no signal about how much of the host they should use. With several builder seats on one machine, each toolchain defaults to "all cores". The host oversubscribes, and interactive work (the TUI, the operator terminal) lags. Feature request: a host resource budget that the daemon computes from active seats and exports to each seat.

Observed (OpenRig 0.5.16, macOS, Apple M4 Pro, 14 cores / 48 GB)

  • Rigs: one 6-seat build rig (Rust workspace, 3 builders using separate git worktrees), a 2-seat rig and the kernel. That is 10 tmux sessions in total.
  • Load average reached ~58–67 on 14 cores, and the operator's terminal UI lagged noticeably.
  • Contributors:
    1. Concurrent cargo build / cargo test from several seats. Each defaults to jobs = ncpu and test threads = ncpu, so 3 builders ask for about 3× the machine.
    2. Transcript capture: tmux capture-pane -p -S -1000 of every pane every transcripts.poll_interval_seconds (default 2 s). With 10 panes, the tmux server and capture clients were among the top CPU consumers (about 40–90% each in bursts). Raising the interval to 10 s (a daemon restart is needed; the setting is read at boot) dropped load from 67 to 39 within a minute.

Local workaround (works, but it's host-specific and Rust-only)

  • A cargo wrapper on PATH: jobs = budget / (running top-level cargo builds + 1), minimum 2, RUST_TEST_THREADS to match, and nice -n 10.
  • sccache shared across worktrees.
  • Limits: the split is fixed at build start; it covers only cargo (not Godot, npm, pytest, etc.); every OpenRig user would have to reinvent it.

Proposal

  1. Daemon-computed budget per seat. From live seats on the host (optionally weighted by role, e.g. builders vs. orchestrators), export env such as OPENRIG_CPU_BUDGET, plus common tool knobs (CARGO_BUILD_JOBS, RUST_TEST_THREADS, MAKEFLAGS=-jN, CMAKE_BUILD_PARALLEL_LEVEL, NPM_CONFIG_JOBS/UV_CONCURRENT_BUILDS, …) at launch or relaunch. Let rig specs declare per-member weights and opt-outs.
  2. Optionally, a shared jobserver: a named-FIFO GNU make jobserver (MAKEFLAGS=--jobserver-auth=fifo:PATH, which cargo's jobserver client understands) owned by the daemon. Concurrent builds across seats then share one token pool dynamically, rather than a split taken at start.
  3. Scheduling priority for build-heavy seats (nice or QoS), so the operator TUI and the human's terminal stay responsive.
  4. Transcript capture cost: a smarter default, or adaptive polling. For example, capture only panes with activity since the last poll (the activity hooks already exist), capture incrementally instead of the last 1000 lines each time, or scale the interval with the number of live panes. Also make transcripts.poll_interval_seconds hot-reloadable.
  5. Surface host load in rig ps / the TUI (load average vs. cores, top seat consumers), so "why is everything slow" is diagnosable from OpenRig itself.

Why it matters

OpenRig's value is running many seats in parallel. Without host awareness, the parallelism a user asks for degrades the host, and the UI the human uses to supervise it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions