Skip to content

Add llama.cpp as a PAIR-managed engine - #148

Merged
Noah-Tervalon-Nvidia merged 70 commits into
developfrom
sherief/llamacpp-downloads
Oct 8, 2026
Merged

Noah-Tervalon-Nvidia merged 70 commits into
developfrom
sherief/llamacpp-downloads

Conversation

@Noah-Tervalon-Nvidia

Copy link
Copy Markdown
Collaborator

Changelog title

llama.cpp is a PAIR-managed engine

Changelog body

  • llama.cpp can be installed, started, stopped, and updated from PAIR, alongside Ollama and LM Studio.
  • Browse and pull llama.cpp models from the model hub, sourced live from approved Hugging Face publishers, with a bounded search across public publishers.
  • Delete a downloaded llama.cpp model from the engine's model list.
  • Inference routed to llama.cpp goes through the local proxy, so the engine participates in cluster routing like the others.
  • On Windows, PAIR installs the Visual C++ runtime llama.cpp requires.

Bumps

  • services: minor
  • nvpair-cluster-manager: patch
  • nvpair-engine-manager: minor
  • nvpair-errors: patch
  • nvpair-job-scheduler: minor
  • nvpair-manual-nodes: patch
  • nvpair-node-info: patch
  • nvpair-node-scanner: patch
  • nvpair-node-settings: patch
  • nvpair-proxy: minor
  • nvpair-tui: minor
  • nvpair-ui-broker: minor
  • nvpair-workload-manager: minor

Summary

Adds llama.cpp as a third PAIR-managed engine: installed, supervised, proxied,
and routed the same way Ollama and LM Studio already are.

This pull request is the base of a stack. It targets fix/macos-app-name
rather than develop so the macOS rename lands beneath it. Review the diff
here; merge the base first.

The work spans the layers an engine touches:

  • nvpair-engine-manager — a llama.cpp manifest with verified
    multi-artifact installs, SSE-backed pulls, idle model sleep, JSON identity
    probes, and model deletion.
  • nvpair-proxy — a llama.cpp facade, model-list path remapping, and
    aggregation across both list routes.
  • nvpair-ui-broker — prepositioned process facades, generalized facade
    status subscriptions, and proxy port validation through settings for every
    engine rather than the two that were special-cased.
  • shared/noderec — a new lc discovery key so a node can advertise a
    llama.cpp endpoint. Every binary links this package, which is why components
    that are otherwise untouched carry a patch.
  • Desktop — the engine declaration, service-state bridge, model hub
    catalog, and the Inference Demo route.

The model catalog is sourced live from approved Hugging Face publishers and is
the one searchable source; every result is an exact owner/repository:Q4_K_M
pull id.

Known follow-ups, handled above this in the stack

  • The macOS and Linux install layout is corrected in a later branch — archives
    wrap their contents in a top-level directory, which this manifest does not
    yet strip.
  • The catalog moves out of the desktop and into engine:catalog so both front
    ends share one implementation, which supersedes the desktop-side llama.cpp
    catalog tests added here.

Test plan

  • go test ./... in each touched services/ component and in services/tests.
  • npm run typecheck, npm run lint, npm run dead-code:check,
    npm run test:unit, npm run service-contracts:check from desktop/.
  • Installed llama.cpp through PAIR on macOS arm64 and Windows x64; pulled a
    model from the hub, started the engine, and confirmed inference through the
    proxy.

@kjlubick kjlubick left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM after all our internal reviews

Base automatically changed from fix/macos-app-name to develop October 8, 2026 17:01
@Noah-Tervalon-Nvidia
Noah-Tervalon-Nvidia force-pushed the sherief/llamacpp-downloads branch from 04c76e2 to 39a5c44 Compare October 8, 2026 17:01
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
sherief-nv and others added 20 commits October 8, 2026 13:02
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
…ngine

Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
Signed-off-by: Sherief Farouk <sfarouk@nvidia.com>
The Visual C++ Runtime install failures were the only user-visible
installer messages still saying "Personal AI Router". The standalone
services installer now uses PRODUCT_DISPLAY_NAME, which its own comment
reserves for user-visible text, rather than PRODUCT_NAME, which has to
stay exact because the uninstall registry key is keyed on it.

Signed-off-by: Terve <ntervalon@nvidia.com>
Installing llama.cpp failed with `engine "llamacpp" was not detected
after install` on every macOS and Linux machine. The download verified
and the archive extracted; the manifest was looking in the wrong place.

detect and runtime.bin named {install_dir}/build/bin/llama-server, which
is where a local cmake build puts its output, not where the published
release archives unpack. Every llama.cpp release tarball wraps its
contents in a directory named for the build tag, so the binary landed at
{install_dir}/llama-b11146/llama-server and nothing matched.

Each tar extraction now strips that wrapper and the paths name
{install_dir}/llama-server, which converges macOS and Linux on the flat
layout Windows already had -- its zip is not wrapped, which is why
Windows was the one platform that worked. Linux strips both archives,
whose wrapper directories are named differently, and its
LD_LIBRARY_PATH follows the binary.

Stripping rather than naming the wrapper keeps the release tag in one
place. Spelling it in detect would mean editing four blocks in lockstep
with every URL bump, which is the shape of mistake this already was.

TestBundledInstallLayout downloads each archive, extracts it the way the
manifest says to, and checks the detect path appears -- the comparison
nobody was making. It is gated on NVPAIR_LIVE_LAYOUT because the
archives are gigabytes, by environment variable rather than the `live`
build tag, which does not currently compile in this package.

TestBundledRuntimeBinIsDetected adds the part that needs no download:
runtime.bin has to be one of the detect paths, so an engine cannot report
installed and then launch something else. It asserts nothing about how
deep a path may be -- llama.cpp wraps and Ollama's Linux archive is
already bin/ and lib/, so there is no shared convention to pin.

Signed-off-by: Terve <ntervalon@nvidia.com>
`npm run stage:vc-redist` asked PowerShell for the staged Visual C++
Redistributable's signature, so it only ran on Windows. The public
`build:electron:win:x64` and `:arm64` scripts cross-build the Windows
installers from Linux and macOS, where the staging step could not run at all,
leaving the packaging precondition unsatisfiable off Windows.

Read the signature out of the PE instead on those platforms: the certificate
table's PKCS#7 SignedData, the image digest it covers, the embedded chain, and
the version resource. Staging now fails unless the digest computed over these
bytes matches the digest Microsoft signed, and unless the signing certificate
chains by issuance and signature to a pinned Microsoft certificate authority.
Windows still defers to the operating system, and both paths meet one policy in
`validateAuthenticodeMetadata()`, so the recorded provenance is identical
whichever host staged the package.

Certificate validity windows, revocation, and the timestamp countersignature
are deliberately not checked; a signing certificate that has expired since it
signed is normal.

Signed-off-by: Terve <ntervalon@nvidia.com>
`stage:vc-redist` failed on both Windows runners with "The
'Get-AuthenticodeSignature' command was found in the module
'Microsoft.PowerShell.Security', but the module could not be loaded".

The build step runs under PowerShell 7, the windows-latest default. npm
and node sit between it and the powershell.exe this script starts, so
PowerShell 7 never strips its own entries from PSModulePath for the
child, and Windows PowerShell inherits a module path that finds 7's
Microsoft.PowerShell.Security first, which it cannot load
(PowerShell/PowerShell#18530).

Start Windows PowerShell without PSModulePath so it builds its own
default. The same failure met anyone running build:win:* from a pwsh
terminal, not only CI.

Signed-off-by: Terve <ntervalon@nvidia.com>
@Noah-Tervalon-Nvidia
Noah-Tervalon-Nvidia force-pushed the sherief/llamacpp-downloads branch from 39a5c44 to 0fa9866 Compare October 8, 2026 17:03
@Noah-Tervalon-Nvidia
Noah-Tervalon-Nvidia merged commit 6d2d67f into develop Oct 8, 2026
1 check passed
pair-release-intent Bot added a commit that referenced this pull request Oct 8, 2026
Apply release intent from PR #148.

Applies-PR: #148
@Noah-Tervalon-Nvidia
Noah-Tervalon-Nvidia deleted the sherief/llamacpp-downloads branch October 8, 2026 17:05
Noah-Tervalon-Nvidia added a commit that referenced this pull request Oct 8, 2026
## Stack

Targets `sherief/llamacpp-downloads` (#148), which targets
`fix/macos-app-name` (#147). Merge those first; this diff is only what
sits on top of them.

Rebasing onto the llama.cpp branch added fourteen commits to the
sixty-nine here:

- **Return the terminal interface to develop's** — removes #148's
additions to the old interface, which the rebuild deletes or rewrites,
so each rebuild commit lands on the tree it was written against. The
broker still fronts every engine by default.
- **llama.cpp in the rebuilt interface and in `engine:catalog`** (nine
commits) — the engine manager serves llama.cpp's Hugging Face catalogue
and its search, the desktop relays the query, and the terminal interface
gains llama.cpp's model actions and an upstream search.
- **Four fixes from review** — a declared pipeline decides whether a
model can chat; GGUF draft heads llama.cpp will not pull are skipped; an
over-long search is refused rather than shortened; the model browser no
longer downloads a row it has stopped showing.

## Description

Rebuilds `nvpair-tui` around a node-first tab layout and closes the gaps
against the desktop application, then moves the model catalogue into the
backend so both front ends browse one implementation.

A node is the unit an operator reasons about, so the tabs become Nodes,
Jobs, Service, Errors, and Logs, and everything specific to one machine
hangs off its row. The ten tabs this replaces spread one machine's facts
across four of them: its engines in one, its models in another, its
ports in a third, its cluster standing in a fourth.

Much of the rest is about claiming no more than the backend guarantees.
A port change reports the port actually bound rather than the one
requested. Clearing an error is only offered where clearing sticks. The
cluster name is labelled as this machine's own. Node presence follows
the broker's snapshot instead of a timestamp that never advanced.

## How to review this

Seventeen commits, in dependency order. Four are prerequisites, one is
the rewrite, and the rest are separable features and fixes. **The
rewrite (`5381e2f1`) is 51 files and is best read in the order below** —
it follows the data rather than the alphabet.

The five views only exist together: the shell sizes them, they share its
status line and frame budget, and the tab set *is* the change. Splitting
them further would have meant authoring intermediate versions of
`nodedetail.go` and `jobs.go` that never existed, so the guided read is
offered instead.

### Read the commits in this order

1. `6b377e78` frame cap single-sourced across four hops — 5 files
2. `e99efb05` `engine:catalog` in the engine manager — the `+109k` is
the committed Ollama list; review `catalog.go` and `catalog_test.go`
3. `001b0e4a` desktop served from that catalogue — the `−109k` is the
same list leaving Electron
4. `c414dbf0` one `inference-dispatcher`, from `cli-bin`
5. `9a61f3de` broker stream errors reported as disconnects
6. `e70fbd21` shared frame primitives (table, status line, row budget)
7. `5381e2f1` **the rewrite** — see the file order below
8. `ec7a6c84` release version stamped for the update notice
9. `144e472f` engine launch-arguments editor, ports onto the same write
10. `32738dcd` documentation and the architecture rule
11. `953a9ab0` CI's staged-binary check learns about the demo client
12. `eaedf559` light-terminal readability
13. `7f5610da` the `requestId` a settings commit requires
14. `67ca9630` background detected before Bubble Tea takes stdin
15. `fbb37e0a` local settings snapshots kept current
16. `94d041d2` leaving the Nodes tab returns to the list
17. `33b5470c` settings verdict taken from the backend

Commits 12–17 came from running the build on real hardware. Four of them
fix defects in commit 9; each says what the backend actually does, which
is worth reading even where the diff is small.

### Read the rewrite's files in this order

**Start with the shape** — 400 lines, and the rest follows from it:

1. `ui/ui.go` — the tab set. The whole change in 27 lines
2. `ui/model.go` — the shell: frame budget, tab switching, the notice
row
3. `ui/keys.go` — why the bindings are what they are
4. `ui/view.go` *(unchanged)* — the contract the five views implement

**The data, before the screens that render it:**

5. `ui/nodeswire.go` — the discovery, cluster, and manual wire shapes
6. `ui/nodesmodel.go` — merging those three feeds onto one key. The
heart of the Nodes tab
7. `ui/nodenames.go` — resolving a node id to a name
8. `ui/engineswire.go` — engine status and model inventory shapes

**The tabs, simplest first:**

9. `ui/logs.go`, `ui/errors.go` — smallest, and they establish the view
idioms
10. `ui/service.go` — worker liveness, log level, data reset
11. `ui/proxystatus.go` — per-engine facade readiness and ports
12. `ui/jobs.go` — the endpoints and live inference work
13. `ui/nodes.go` — the merged node table, filtering, pairing keys
14. `ui/invite.go` — pairing events scoped to their session

**The drill-down and its two satellites** — the largest file, read last:

15. `ui/nodedetail.go` — engines and models for one machine
16. `ui/nodetelemetry.go` — the direct `/v1/node-info` poll
17. `ui/catalog.go` — the downloadable-model browser
18. `ui/demo.go`, `ui/demoschedule.go` — the Inference Demo

**Then the tests**, which read as the specification:

19. `ui/framebudget_test.go` — no view overflows its frame at any
supported size
20. `ui/table_test.go` — every table has columns at construction
21. `ui/nodesmodel_test.go`, `ui/nodedetail_test.go`,
`ui/service_test.go` — the per-area behaviour
22. the remaining `_test.go` files

**Deletions need no reading**: `cluster.go`, `engines.go`, `health.go`,
`manualnodes.go`, `proxies.go`, `settings.go`, `workloads.go` are the
tabs the five replace.

## Scope

Included: the terminal interface rewrite, its Inference Demo and update
notice; `engine:catalog` replacing the Electron-only model hub; one
`inference-dispatcher` in `cli-bin`; an editor for an engine's launch
arguments with the two port fields moved onto the same revision-checked
write; documentation.

Excluded: no change to routing, scheduling, or proxy behaviour; none to
pairing or trust; the desktop renderer's settings UI is untouched, only
its catalogue source moved.

## Validation

Everything CI runs, run locally on macOS arm64, Go 1.26.3, Node 22:

- `node scripts/spdx-headers.mjs` — 1039 checked, 0 missing
- `npm --prefix desktop run verify:build-scripts`,
`service-contracts:check`, `typecheck`, `lint`, `dead-code:check`,
`test:unit` (230 passing)
- `npm --prefix desktop run build:modular-binaries -- --force` — 13
binaries
- Go component tests with `-race` across `shared`, `nvpair-tui`,
`nvpair-ui-broker`, `nvpair-engine-manager`
- `cd services/tests && go test ./... -count=1 -timeout=20m` — passes,
no failures
- `services/build.sh`, then the staged-binary comparison this branch
updates

Two failures are pre-existing and reproduce identically on `develop` in
a clean worktree: `TestUninstallTerminatesRunningInstance` in
`nvpair-engine-manager`, and
`services/shared/splitlisten/splitlisten_test.go` is not `gofmt`-clean.

Manual, on a Mac and a DGX Spark over SSH: pairing, engine lifecycle,
model download, the demo, and the settings editor. The
`ui.ReleaseVersion` stamp was read back out of both binary sets, since a
`-X` against a wrong symbol path fails silently. Terminal background
detection was checked in a pty against a terminal that answers the query
and one that ignores it.

## Risk

The engine settings path is the part to look at. The two port fields
previously wrote through `engine:set-port` and the proxy's own
`set-port`, which meant two writers with no shared revision — the last
to finish won and neither could tell it had lost. All three fields are
now one revision-checked write through `engine:preview-settings` and
`engine:apply-settings`. That also makes them editable on a peer, which
the old path refused because it had no remote form; whether a given
engine accepts the write is the snapshot's `Editable` answer rather than
a rule kept in the front end.

`jsonrpc.WorkerFrameBytes` raises the inbound frame cap on three hops
that carry large worker replies, from the 1 MiB default to 8 MiB. The
Ollama catalogue is roughly 1.9 MiB and could not previously cross them;
an over-long frame is a terminal read error rather than a dropped
message, so the hops have to agree.

No wire-format removals, no persisted-data migrations, no change to how
nodes authenticate to each other.

## Release intent

<!-- pair-release-intent:v1 -->
### Changelog title
A rebuilt terminal interface, and one model catalogue for both front
ends

### Changelog body
- The terminal interface is organised around machines: Nodes, Jobs,
Service, Errors, and Logs, with a per-machine drill-down for engines,
models, ports, and hardware.
- Engines can be started with your own arguments and environment
variables, edited from the terminal interface on this machine or on a
paired one. Changes are validated before they are saved.
- Downloadable models are served by the engine manager, so the terminal
interface and the desktop application browse the same catalogue.
- The Inference Demo runs from the terminal interface's Jobs tab, so a
headless machine can demonstrate routing.
- The terminal interface says when a newer release of PAIR is published,
on every tab until dismissed. It installs nothing.
- Pairing keys follow the words on screen: `p` pairs, `a` accepts, `f`
finds a machine by address.

### Bumps
- services: major
- nvpair-cluster-manager: none
- nvpair-engine-manager: minor
- nvpair-errors: none
- nvpair-job-scheduler: none
- nvpair-manual-nodes: none
- nvpair-node-info: none
- nvpair-node-scanner: none
- nvpair-node-settings: none
- nvpair-proxy: none
- nvpair-tui: major
- nvpair-ui-broker: patch
- nvpair-workload-manager: none
<!-- /pair-release-intent:v1 -->

## Checklist

- [x] I have read the [Contributing
Guidelines](https://github.com/NVIDIA/Personal-AI-Router/blob/main/CONTRIBUTING.md).
- [x] Every commit is signed off (`git commit -s`), certifying the
[Developer Certificate of Origin](https://developercertificate.org/).
- [x] New or existing tests cover the change.
- [x] Relevant documentation is updated.
- [x] I checked the diff, changed filenames, and commit messages for
credentials, private data, internal URLs, internal issue identifiers,
and generated artifacts.
- [x] I recorded the validation commands and results above.
- [x] I declared version bumps in the release-intent block above.
`services/versions.json` is written by automation — do not edit it by
hand.

---------

Signed-off-by: Terve <ntervalon@nvidia.com>
kjlubick pushed a commit that referenced this pull request Oct 8, 2026
<!-- pair-release-intent:v1 -->
### Changelog title
Safer engine removal and data resets

### Changelog body
- Removing an engine never follows a symlink or Windows junction into
another folder.
- An LM Studio models folder linked from `~/.lmstudio/models` is kept
when LM Studio is removed.
- Resetting app data, from Settings or the terminal interface, stops
before deleting anything if an engine PAIR installed cannot be removed,
and says which one.

### Bumps
- services: patch
- nvpair-cluster-manager: none
- nvpair-engine-manager: patch
- nvpair-errors: none
- nvpair-job-scheduler: none
- nvpair-manual-nodes: none
- nvpair-node-info: none
- nvpair-node-scanner: none
- nvpair-node-settings: none
- nvpair-proxy: none
- nvpair-tui: patch
- nvpair-ui-broker: none
- nvpair-workload-manager: none
<!-- /pair-release-intent:v1 -->

## Summary

Review fixes for the work below this in the stack. **Stacked on**
`fix/indeterminate-install-progress` (#145), at the top of #147 → #148 →
#117 →
#150 → #145; merge it after them, and before the stack ships.

### Engine removal (#150)

- **Junctions are never followed.** Since Go 1.23 a Windows junction is
reported
as an irregular file, not a symlink, so the symlink check let the
elevated
uninstaller read through one and empty the directory it names. Removal
now
descends only into what `Lstat` reports as a real directory. Two Windows
tests use real junctions; CI runs Go tests on Linux only, so they run on
  Windows machines.
- **The model store is matched by file identity.** A store that is a
symlink to
  another folder under the removal target was deleted, and case handling
followed the OS rather than the filesystem. `os.SameFile` replaces both.
New
tests cover a symlinked store, a store two levels down (the first to
exercise
  the recursion), and a store spelled in another case.
- `spec.md` now states what removing an engine may delete.

### Resets (#150)

The TUI and desktop resets wiped app data even when an engine could not
be
removed, deleting the only record that it was PAIR's. Both now stop
first and
say which engine, or that the service is not running. The orchestrator
gains its
first tests.

### Smaller fixes

- "Personal AI Router" → "NVIDIA PAIR" in the uninstall work's
user-visible
wording. The data path and firewall rule names keep the old name on
purpose.
- Catalog tests assert the error itself (#117).
- Getting started covers the update that renames the macOS app (#147).
- Engine lifecycle documents an LM Studio models folder moved inside
  `~/.lmstudio`.

## Test plan

- `go vet` (darwin, linux, windows) and `go test ./...` in
  `nvpair-engine-manager` and `nvpair-tui`.
- `npm run typecheck`, `npm run lint`, `npm run test:unit` from
`desktop/`.
- `TestUninstallTerminatesRunningInstance` fails locally on macOS
identically on
  `develop`. The Windows junction tests compile but did not run here.

---------

Signed-off-by: Terve <ntervalon@nvidia.com>
ckelseynv added a commit that referenced this pull request Oct 8, 2026
…creenshots

Bring the docs in line with the release stack as if it were merged: llama.cpp
from #148, the rebuilt terminal interface from #117, and kept model stores
from #150.

- Replace the three terminal-interface screenshots, which showed the old
  ten-tab interface, with renders of the current views, and add one of a
  machine's detail screen.
- Add llama.cpp where docs still described two engines: its install path and
  router port, the proxy, scheduler, workload, and error references, the Linux
  install guide, and troubleshooting a self-started llama-server on 8080.
- Drop the README claim that first-run setup does not preselect llama.cpp.
- Correct the terminal interface docs against the code: ten workers, where
  its own logs go, which engine settings work on a peer, and when the update
  check runs.

Signed-off-by: Chris Kelsey <ckelsey@nvidia.com>
Noah-Tervalon-Nvidia pushed a commit that referenced this pull request Oct 9, 2026
…reenshots (#151)

## Description

A documentation pass over the release that landed in `develop` with
#147, #148, #117, #150, #145, and #153. It covers llama.cpp everywhere
the docs list engines, replaces the terminal-interface screenshots of
the old ten-tab interface, and corrects docs that no longer match the
code.

The branch was replayed onto `develop` after the stack was
squash-merged, so it carries only the docs commits.

<!-- pair-release-intent:v1 -->
### Changelog title
n/a

### Changelog body
n/a

### Bumps
- services: none
- nvpair-cluster-manager: none
- nvpair-engine-manager: none
- nvpair-errors: none
- nvpair-job-scheduler: none
- nvpair-manual-nodes: none
- nvpair-node-info: none
- nvpair-node-scanner: none
- nvpair-node-settings: none
- nvpair-proxy: none
- nvpair-tui: none
- nvpair-ui-broker: none
- nvpair-workload-manager: none
<!-- /pair-release-intent:v1 -->

## Scope

**Terminal interface screenshots**
(`docs/assets/onboarding/terminal-interface/`)

- The three existing images showed the ten-tab interface from before
#117. They are replaced with renders of the current Nodes, pairing, and
Jobs views, at the same size as the originals.
- New `04-tui-node-detail.png`: a machine's Engines and Models panes
with Ollama, LM Studio, and llama.cpp. It is placed in "Look at One
Machine".
- The renders drive the real `ui` views and styles with sample data for
two machines, so the layout, columns, and footer are the interface's own
output. The harness that produced them is not part of this pull request.

**Terminal interface capabilities**

- `docs/known-issues.mdx` and `docs/engine-lifecycle.mdx` said the
terminal interface cannot list or delete models, change ports, manage
another node's engines, or show which node ran a job. It does all of
these now. The lists keep only what is still true: updating an engine,
sending your own inference request, installing PAIR updates, and
surviving a dropped connection.
- `docs/engine-lifecycle.mdx`: ports and startup arguments can be edited
from a paired machine. Only uninstall, update, and CORS settings stay
local.
- `README.md`: leaving a cluster is `l` on the **Nodes** tab. There is
no Cluster tab any more.
- `docs/architecture.mdx`: the broker's `ping` is used by the
**Service** tab, not a health view.

**Terminal interface docs** (`docs/terminal-interface.mdx`,
`services/nvpair-tui/README.md`), checked against the code:

- It runs a broker and ten workers, not eleven.
- While the interface runs, its own logs go to the Logs tab, not stderr.
- The engine port, proxy port, and startup arguments are offered on a
peer and written through `engine:apply-settings`. Only restart and
uninstall are local-only. A peer's screen lists only its engine port,
though `p` still changes that peer's proxy port.
- Only a plain `go build` skips the update check. Builds made with
`services/build.sh` or `services/build.bat` check.
- `r` forgets a machine added with `f` without asking, and the "asks
first" rule says so. The catalog filter also matches tags and clears
with `c`.

**llama.cpp coverage**

- `README.md`: removed the claim that first-run setup does not preselect
llama.cpp. `WELCOME_ENGINE_DEFAULT_SELECTED` preselects it. The macOS
install step names **NVIDIA PAIR**.
- `docs/getting-started.mdx`:
- "Using an Engine's Own Command Line" gives llama.cpp's install path on
Windows, Linux, and macOS, its cache, and querying its router on `8081`.
The `OLLAMA_HOST` advice is marked as Ollama-only.
  - Ollama's path table gains a macOS row.
- The endpoint table gains a llama.cpp row: OpenAI-compatible paths on
`8080`, no Ollama `/api/...` and no Anthropic Messages API.
- `docs/troubleshooting.mdx`: a `llama-server` you started yourself on
`8080` hides requests from PAIR. The fix is to stop it and reopen PAIR,
because the proxy keeps its fallback port until the broker restarts.
- `docs/building.mdx` and `docs/engine-lifecycle.mdx`: model folders and
example ports include `~/.llamacpp`, `~/.lmstudio/models`, and `8080`.
- `services/installer/linux/INSTALL.md` and
`services/installer/macos/INSTALL.md`:
- Models live in each engine's own store, so deleting the bundle or
configuration leaves them.
- The bind-conflict note named `11435`, which is Ollama's relocated
engine port, not a proxy port. It now lists the three proxy ports.
- Linux: the bundle layout lists `nvpair-tui`, and the requirements say
`x86_64` or `arm64`, which `installer_build.sh` produces.
- Service and developer references that still assumed two engines:
  - the `nvpair-proxy` README and spec;
- the `nvpair-job-scheduler` spec, including the `scheduler:get-status`
shape;
  - the `nvpair-workload-manager` README and spec;
  - the `llamacpp-proxy` error IDs in `nvpair-errors`;
  - the `nvpair-ui-broker` README;
  - `.cursor/rules/proxy-inference-routing.mdc`;
  - `desktop/docs/services-backend.md`.
- `services/nvpair-engine-manager/README.md` lists
`engine:uninstall-managed`. `desktop/docs/services-parity.md` marks the
pull `percent` as optional, after #145.

**Not included**

- Desktop screenshots that still show only Ollama, such as
`getting-started/01-first-run-setup.png`. They need a capture from the
Electron app.
- `CHANGELOG.md`, which release-intent automation writes.

## Validation

- Swept every `*.md`, `*.mdx`, and `.cursor/rules/*.mdc` file for engine
lists that name Ollama or LM Studio without llama.cpp, and for
references to the old terminal interface. Each hit was checked against
the code. Statements that PAIR adopts an existing Ollama or LM Studio
are left alone, because llama.cpp is never adopted.
- Each factual change was checked against its source, including the
engine manifests, `engines.All()`, `schedule.go`, `engineproxy.go`,
`nodedetail.go`, `nodes.go`, `logoutput.go`, `updatecheck.go`,
`services/build.sh`, `installer_build.sh`, and `appdir.go`.
- A three-model review pass over the branch. Every finding was verified
against the code before it was applied.
- Reviewed each rendered screenshot against the doc text and alt text it
accompanies.
- `go vet ./ui/` in `services/nvpair-tui` after the harness was removed.
- `python3 scripts/release-intent/validate_pr.py --description-file
<this description> --skip-owned-files-check`

## Risk

Documentation and images only. No code, build, or versioned file
changes.

Found while checking and left for a separate change:

- In a colour terminal, an unplaced job's RAN ON cell renders as `c…`
instead of `choosing`. The styled text is truncated by its byte width,
and the tests render without colour, so they do not see it.
- A comment in `ui/updatecheck.go` still says "eleven workers".
- `nvpair-manual-nodes` probes hand-added nodes for Ollama and LM Studio
only, not llama.cpp.
- `installer_build.sh` does not copy `inference-dispatcher` into the
services tarball, so the Jobs tab's test traffic cannot start from it.

## Checklist

- [x] I have read the [Contributing
Guidelines](https://github.com/NVIDIA/Personal-AI-Router/blob/main/CONTRIBUTING.md).
- [x] Every commit is signed off (`git commit -s`), certifying the
[Developer Certificate of Origin](https://developercertificate.org/).
- [ ] New or existing tests cover the change. (Documentation only.)
- [x] Relevant documentation is updated.
- [x] I checked the diff, changed filenames, and commit messages for
credentials, private data, internal URLs, internal issue identifiers,
and generated artifacts.
- [x] I recorded the validation commands and results above.
- [x] I declared version bumps in the release-intent block above.
`services/versions.json` is written by automation — do not edit it by
hand.

---------

Signed-off-by: Chris Kelsey <ckelsey@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants