Summary
OpenRig knows how many seats are live on a host, but seats get no signal about how much of the host they should use. With several builder seats on one machine, each toolchain defaults to "all cores". The host oversubscribes, and interactive work (the TUI, the operator terminal) lags. Feature request: a host resource budget that the daemon computes from active seats and exports to each seat.
Observed (OpenRig 0.5.16, macOS, Apple M4 Pro, 14 cores / 48 GB)
- Rigs: one 6-seat build rig (Rust workspace, 3 builders using separate git worktrees), a 2-seat rig and the kernel. That is 10 tmux sessions in total.
- Load average reached ~58–67 on 14 cores, and the operator's terminal UI lagged noticeably.
- Contributors:
- Concurrent
cargo build / cargo test from several seats. Each defaults to jobs = ncpu and test threads = ncpu, so 3 builders ask for about 3× the machine.
- Transcript capture:
tmux capture-pane -p -S -1000 of every pane every transcripts.poll_interval_seconds (default 2 s). With 10 panes, the tmux server and capture clients were among the top CPU consumers (about 40–90% each in bursts). Raising the interval to 10 s (a daemon restart is needed; the setting is read at boot) dropped load from 67 to 39 within a minute.
Local workaround (works, but it's host-specific and Rust-only)
- A
cargo wrapper on PATH: jobs = budget / (running top-level cargo builds + 1), minimum 2, RUST_TEST_THREADS to match, and nice -n 10.
- sccache shared across worktrees.
- Limits: the split is fixed at build start; it covers only cargo (not Godot, npm, pytest, etc.); every OpenRig user would have to reinvent it.
Proposal
- Daemon-computed budget per seat. From live seats on the host (optionally weighted by role, e.g. builders vs. orchestrators), export env such as
OPENRIG_CPU_BUDGET, plus common tool knobs (CARGO_BUILD_JOBS, RUST_TEST_THREADS, MAKEFLAGS=-jN, CMAKE_BUILD_PARALLEL_LEVEL, NPM_CONFIG_JOBS/UV_CONCURRENT_BUILDS, …) at launch or relaunch. Let rig specs declare per-member weights and opt-outs.
- Optionally, a shared jobserver: a named-FIFO GNU make jobserver (
MAKEFLAGS=--jobserver-auth=fifo:PATH, which cargo's jobserver client understands) owned by the daemon. Concurrent builds across seats then share one token pool dynamically, rather than a split taken at start.
- Scheduling priority for build-heavy seats (nice or QoS), so the operator TUI and the human's terminal stay responsive.
- Transcript capture cost: a smarter default, or adaptive polling. For example, capture only panes with activity since the last poll (the activity hooks already exist), capture incrementally instead of the last 1000 lines each time, or scale the interval with the number of live panes. Also make
transcripts.poll_interval_seconds hot-reloadable.
- Surface host load in
rig ps / the TUI (load average vs. cores, top seat consumers), so "why is everything slow" is diagnosable from OpenRig itself.
Why it matters
OpenRig's value is running many seats in parallel. Without host awareness, the parallelism a user asks for degrades the host, and the UI the human uses to supervise it.
Summary
OpenRig knows how many seats are live on a host, but seats get no signal about how much of the host they should use. With several builder seats on one machine, each toolchain defaults to "all cores". The host oversubscribes, and interactive work (the TUI, the operator terminal) lags. Feature request: a host resource budget that the daemon computes from active seats and exports to each seat.
Observed (OpenRig 0.5.16, macOS, Apple M4 Pro, 14 cores / 48 GB)
cargo build/cargo testfrom several seats. Each defaults tojobs = ncpuand test threads = ncpu, so 3 builders ask for about 3× the machine.tmux capture-pane -p -S -1000of every pane everytranscripts.poll_interval_seconds(default 2 s). With 10 panes, the tmux server and capture clients were among the top CPU consumers (about 40–90% each in bursts). Raising the interval to 10 s (a daemon restart is needed; the setting is read at boot) dropped load from 67 to 39 within a minute.Local workaround (works, but it's host-specific and Rust-only)
cargowrapper on PATH:jobs = budget / (running top-level cargo builds + 1), minimum 2,RUST_TEST_THREADSto match, andnice -n 10.Proposal
OPENRIG_CPU_BUDGET, plus common tool knobs (CARGO_BUILD_JOBS,RUST_TEST_THREADS,MAKEFLAGS=-jN,CMAKE_BUILD_PARALLEL_LEVEL,NPM_CONFIG_JOBS/UV_CONCURRENT_BUILDS, …) at launch or relaunch. Let rig specs declare per-member weights and opt-outs.MAKEFLAGS=--jobserver-auth=fifo:PATH, which cargo's jobserver client understands) owned by the daemon. Concurrent builds across seats then share one token pool dynamically, rather than a split taken at start.transcripts.poll_interval_secondshot-reloadable.rig ps/ the TUI (load average vs. cores, top seat consumers), so "why is everything slow" is diagnosable from OpenRig itself.Why it matters
OpenRig's value is running many seats in parallel. Without host awareness, the parallelism a user asks for degrades the host, and the UI the human uses to supervise it.