Repository navigation
Add Anthropic Messages API Routing - #27
Noah-Tervalon-Nvidia merged 5 commits into
Conversation
7974686 to
5840b2b
Compare
|
Glad to have you here contributing to the project! This lgtm - I am currently working on some major redesigns for the proxies to unify them. I want to wait on merging in changes for the proxies until we get that out to avoid extra work in merging into the new version so I expect this will need a minor rework once we get that out and we'll wait on merging it until then. |
|
Thanks for taking the to review it and heads-up on the proxy rework. Sounds good - I’ll hold off until the proxies are redesigned and then make the necessary adjustments once it’s ready. I’m looking forward to being able to contribute more to the project! |
|
Notes: https://lmstudio.ai/docs/developer/anthropic-compat and https://docs.ollama.com/integrations/claude-code are where the currently supported engines document their Anthropic pass-through support (which this PR adds). |
| // TestHandleHTTP_AnthropicMessages_InferenceRouting proves /v1/messages is | ||
| // treated as an inference endpoint: model-based candidate filtering is applied | ||
| // and the request body is forwarded unchanged to the matching node. | ||
| func TestHandleHTTP_AnthropicMessages_InferenceRouting(t *testing.T) { |
There was a problem hiding this comment.
Thanks for adding this test. I'm going to upload a follow-on commit that reduces the duplication here with a table-based test (that seemed easier than me describing what I had in mind).
|
One more review comment, could you add something in the description about improving support for Claude Code (I'm guessing that's the motivation) or whatever problem you bumped into that inspired this change? It's nice to include the "how does this benefit the user" a bit more explicitly in changes. Thanks again for your contribution! |
|
Thanks Kaylee! I’ve improved the PR description to better describe the user benefit and Claude Code compatibility. Also really appreciate you adding the test refactor based on tables, it's a much cleaner way to express the test coverage. |
|
I really enjoy working with PAIR and look forward to doing more for the project. Fingers crossed it gets merged! |
|
Just to touch base here, we are working on getting our final batch of internal changes out into the public repo. Once those are in, we'll get this rebased (if necessary) and landed. Thanks again for your contribution! |
|
Heads up — This PR edits files in those directories, so it will need updating before it can Apologies for the churn, and thanks for the contribution. Happy to help work out |
Signed-off-by: Rayees <rayiesamin@gmail.com>
Signed-off-by: Rayees <rayiesamin@gmail.com>
Signed-off-by: Rayees <rayiesamin@gmail.com>
40b1862 to
7749b54
Compare
Signed-off-by: Kaylee Lubick <klubick@nvidia.com>
|
Hi Noah, thanks again for the kind words and for the heads-up about the proxy redesign. I’ve now rebased the PR onto the current The implementation now:
For validation:
The updated branch is pushed at |
|
You and I were apparently trying to rebase at the same time. You got in first, I pushed a few other changes on top of yours (e.g. documentation and cleanup tests). I think we should be good to go |
|
Hehe, looks like we were racing each other there 😄 Glad we got it sorted. Thanks for jumping in and cleaning things up! |
Signed-off-by: Kaylee Lubick <klubick@nvidia.com>
Two hunks conflicted with develop and are resolved here: - services/nvpair-engine-manager/lifecycle.go: the HTTP probe now defers httpcon.DrainAndClose (develop's connection-reuse fix from NVIDIA#37) after the llama.cpp identity check has read the body, instead of closing it unread. - services/nvpair-proxy/engines.go: develop's per-engine base route sets (NVIDIA#27) gain llamaCppBaseRoutes, and the llama.cpp facade profile is built with slices.Concat like Ollama and LM Studio, so it also classifies the Anthropic /v1/messages route. The llama.cpp router at b10826 serves /v1/messages itself, so the facade routes it like the other engines. The route-role test gains llama.cpp cases for the shared inference routes and for the router's own /models passthrough. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Signed-off-by: pgoode41 <pgoode41@gmail.com>
## Description
Adds two pages to the public docs: **Release Notes**, starting with
1.0.0, and an **Upgrade Guide** for updating from a version before
1.0.0. People who read the docs rather than GitHub had no release notes
to look at.
The 1.0.0 notes are adapted from the release notes master doc. The
Upgrade Guide starts from the fact that most machines need nothing
extra, then covers the few setups that need a small step:
- tidying up the old `PAIR.app` on macOS;
- an LM Studio that an earlier PAIR installed;
- browser clients affected by the CORS change.
<!-- pair-release-intent:v1 -->
### Changelog title
n/a
### Changelog body
n/a
### Bumps
- services: none
- nvpair-cluster-manager: none
- nvpair-engine-manager: none
- nvpair-errors: none
- nvpair-job-scheduler: none
- nvpair-manual-nodes: none
- nvpair-node-info: none
- nvpair-node-scanner: none
- nvpair-node-settings: none
- nvpair-proxy: none
- nvpair-tui: none
- nvpair-ui-broker: none
- nvpair-workload-manager: none
<!-- /pair-release-intent:v1 -->
## Scope
- `docs/release-notes.mdx` (new): the 1.0.0 notes, covering what's new,
compatibility notes, targeted bug fixes, community contributions with PR
links, and special thanks. Notes before 1.0.0 are linked on GitHub
Releases.
- `docs/upgrade-guide.mdx` (new). It opens by saying an update keeps
settings, cluster membership, and models, and that every machine in a
cluster should run the same version. Each section then says who it
applies to:
- **On macOS, tidy up the old app:** run the old app's
`uninstall-macos.sh` without `--purge`, then install `NVIDIA PAIR.app`.
0.1.0 and 0.1.1 both ship that script, and both keep data without
`--purge`.
- **If an earlier PAIR installed LM Studio for you:** it keeps working,
but it has no install marker (`installed-by-pair.json`), so Uninstall,
Reset app data, and uninstalling PAIR leave it. Steps to remove it while
keeping `models`, or to reinstall it so PAIR manages it.
- **If you use PAIR from a web page or browser extension:** which launch
option allows a site for each engine.
- **Checking everything is working.**
- `fern/docs.yml`: both pages under Guides, after Getting Started.
- `README.md`, `docs/getting-started.mdx` ("Keeping PAIR Up to Date" and
"Learn More"), and `docs/known-issues.mdx`: link to the new pages.
`known-issues` gains an entry for the LM Studio limitation.
## Validation
- Each upgrade step was checked against the code:
`desktop/scripts/build/macos/uninstall.sh` at HEAD, `v0.1.0`, and
`v0.1.1`; `services/nvpair-engine-manager/install.go` and
`provenance.go`; `manifests/lmstudio.json`;
`desktop/docs/macos-privileged-helper.md`; and
`docs/engine-settings.mdx` for the CORS controls.
- Contributor PRs #27, #34, #37, #62, #80, #95, #120, #122, and #126
were checked for merge state and author.
- `node scripts/spdx-headers.mjs`
- `python3 scripts/release-intent/validate_pr.py --description-file
<this description> --skip-owned-files-check`
## Risk
Documentation only. No code, build, or versioned file changes.
Left for the product owner to confirm against the master doc:
- The app is named **NVIDIA PAIR** (`NVIDIA PAIR.app`).
- The macOS note does not say models could be lost, because models never
live inside the app.
- The bug-fix line names LM Studio only, because llama.cpp never shipped
publicly.
- `README.md` and `known-issues.mdx` still call Windows on ARM
experimental, beside the RTX Spark validation note.
## Checklist
- [x] I have read the [Contributing
Guidelines](https://github.com/NVIDIA/Personal-AI-Router/blob/main/CONTRIBUTING.md).
- [x] Every commit is signed off (`git commit -s`), certifying the
[Developer Certificate of Origin](https://developercertificate.org/).
- [ ] New or existing tests cover the change. (Documentation only.)
- [x] Relevant documentation is updated.
- [x] I checked the diff, changed filenames, and commit messages for
credentials, private data, internal URLs, internal issue identifiers,
and generated artifacts.
- [x] I recorded the validation commands and results above.
- [x] I declared version bumps in the release-intent block above.
`services/versions.json` is written by automation — do not edit it by
hand.
---------
Signed-off-by: Chris Kelsey <ckelsey@nvidia.com>
Summary
Added support for routing requests to the Anthropic Messages API via PAIR's unified nvpair-proxy service whether using Ollama or LM Studio.
Since the POST /v1/messages endpoint is an inference route, it makes use of PAIR's existing model selection, routing, and failover path without protocol translation.
It enables compatible clients such as Claude Code to make Messages API requests via PAIR and at the same time keep the upstream request/response format.
Fixes #16
Changes
Include the POST request to the /v1/messages endpoint in the shared inference route set that is used by both the Ollama and LM Studio facades.
Extend the route-role tests for both engine profiles.
Extend the failover coverage of unified-proxy for requests to Anthropic Messages.
Extend the test of the cross-process strict model-routing to include requests to Ollama and LM Studio via the /v1/messages endpoint.
Update the README file for the unified nvpair-proxy to include information about the Anthropic Messages route.
ollama-proxy:andlmstudio-proxy:namespaces unchanged.We do not make any changes to services/versions.json or to CHANGELOG.md; the release metadata is provided via the required PR release-intent block listed below.
Design
The implementation employs passthrough rather than protocol translation; PAIR simply adds the
/v1/messagesendpoint to the current inference-routing interface, after which the chosen engine receives the original request and returns a response in the original format.The same system components that are used by the other inference paths—that is, the model eligibility mechanism, the candidate selection process, the scheduler reservation, the first-byte commit, the retry function, and the failover mechanism—are also responsible for handling this route.
There is no Anthropic-specific routing subsystem or separate proxy process.
Validation
Automated
cd services/nvpair-proxy && go test -count=1 ./.... PASS.cd services/tests && go test -count=1 -run 'TestStrictModelRoutingAcrossProcesses' -v .yielded a pass in the cases involving Ollama, Ollama Anthropic, LM Studio, and LM Studio Anthropic.cd services/tests && go test ./...resulted in one failure in theTestBrokerProxySetPortRebindstest (in the file broker_supervision_test.go at line 551). This test also fails when run against the currentorigin/developbaseline, showing the same port/settings resolution error.The script node scripts/spdx-headers.mjs has passed; out of 1014 files checked, 0 were missing and 97 were skipped.
git diff --checkpasses.End-to-end
I verified that a genuine Claude Code session had been established via PAIR using Ollama. Through the routed session, Claude Code produced the output with the required content, thus confirming that there had actually been an interaction between an Anthropic client and the backend via PAIR.
Notes
The implementation is deliberately restricted so that the Anthropic Messages endpoint will make use of PAIR's current routing and failover facilities and will not carry out Anthropic-specific request/response translation.
Changelog title
Route Anthropic Messages API requests
Changelog body
The Ollama and LM Studio endpoints offered by PAIR now accept requests for the Anthropic Messages API via POST to the /v1/messages endpoint and direct them to the owner of the requested model, just as happens with the other inference methods.
Bumps