Repository navigation
[Feature]: Expose the Anthropic Messages API (POST /v1/messages) on the PAIR proxy #16
Description
Activity
Additional validation
I carried out an end-to-end test of the implementation using Claude Code 2.1.261 and a local Ollama/Qwen3 setup:
The request to the POST /v1/messages endpoint was successfully forwarded via PAIR to Ollama.
The response was successfully obtained by Claude Code via the new route.
Streaming now features native Anthropic SSE events (message_start, content_block_*, message_delta, message_stop).
The Claude Code managed to use the file-writing tool via the Messages API.
The current routing and failover tests work.
The endpoint for /v1/messages/count_tokens is currently returning a 404 from PAIR since it is not exposed by the proxy.The implementation is still based on passthrough, so PAIR does not translate Anthropic's request and response fields.
I couldn't check the routing for the remote node since at the moment I have only one PAIR node available.
I offered to test on Apple Silicon, so here is that, run against #27 at 5840b2b on darwin/arm64 with Go 1.26.5. Short version: #27 passes and introduces no failures, and the
count_tokens404 looks like it is coming from the engine rather than from PAIR.Tests.
services/lmstudio-proxyandservices/ollama-proxyboth pass at #27, including its newfailover_test.goin each. The one red test inollama-proxy,TestAliasSelfTargetMatchesBoundLoopbackAddressNotPortAlone, also fails at #27's merge base and is unrelated: it binds127.0.0.2, and macOS configures only127.0.0.1onlo0where Linux treats all of127.0.0.0/8as local.On
/v1/messages/count_tokens. Checked against Ollama 0.33.3 directly, with no PAIR in the path:POST http://127.0.0.1:11434/v1/messages/count_tokens -> 404 "404 page not found" POST http://127.0.0.1:11434/v1/messages -> a well-formed Anthropic error envelopeThe bare
404 page not foundis the engine's own mux answering, so the route does not exist upstream.inferenceEndpointsis an exact-match map and the proxy handler is a catch-all rather than a registered route set, which reads to me as PAIR forwardingcount_tokensand the engine 404ing it, rather than PAIR declining to expose it. If that is right, this is already the "pass it through where an engine implements it, clear error where none does" outcome, and nothing further is needed for it. I did not stand up the full proxy to confirm the hop, so treat that half as a source reading.What #27 changes, and what neither of us has tested. The diff adds
/v1/messagestoinferenceEndpoints, which is the classification that decides whether a request is a cluster workload: model-eligible candidate selection, and failover to the next candidate. Since the handler already forwards any path, a single-node setup exercises the passthrough but not the part this PR actually changes. Your end-to-end run and this one are complementary, and neither covers remote-node routing. I have one node too, so I could not close that gap.One thing I checked and found fine, in case it comes up in review: Ollama answers an unknown model with HTTP 404 on
/v1/messages, and failover treats a 404 on an inference route as a reason to try the next candidate. It answers/v1/chat/completionsand/api/chatthe same way, so/v1/messagesnow behaves exactly like the routes already in that map. Consistent, not a new edge.Reacted by Rayees AminThank you for taking the time to test and review this, and for confirming that the original changes pass on Apple Silicon.
I have included the extra cross-process integration testing which you recommended for /v1/messages. The test now runs Anthropic Messages through the current multi-node routing system which includes the routing/failover path.
I also ran the relevant validation locally:
node scripts/spdx-headers.mjs — 871 checked, 0 missing
go test -run TestStrictModelRoutingAcrossProcesses -count=1 — PASS
cd services/tests && go test ./... — PASS (334.186s)
The integration-test change is now pushed to the PR as commit a5500b1.
Thank you for providing complete testing results together with your valuable comments. The testing process revealed an essential gap which needed to be addressed because it affected the remote-node routing system.
Reacted by Michael Pursifull- added a commit that references this issue
on Sep 23, 2026
PAIR pools capacity so that a machine too small to hold a useful model can still use one. Coding agents that speak the Anthropic Messages API are a large and growing class of client that cannot reach PAIR at all today, and the reason is protocol rather than capability: both engines PAIR already routes to serve
POST /v1/messagesnatively, and PAIR re-exposes them as OpenAI only. Adding the endpoint alongside the existing ones would open pooled local capacity to those clients without changing anything for current ones.Area
API compatibility
User problem
Who is affected: people running an agentic coding client that speaks the Anthropic Messages API against local models, across a mix of machines where no single one is large enough to hold a useful model roster. That is the case PAIR is built for, and it is ours: a team where a few machines have 64 GB or more and most have far less.
What cannot be accomplished today: such a client sends every request to one base URL and speaks only
/v1/messages. PAIR presents Ollama-compatible and OpenAI-compatible proxy endpoints, so the client cannot reach PAIR at all.The gap is narrow rather than architectural, because the destination engines already understand the dialect:
/v1/messagesAPI".Checked directly rather than taken from the notes: on LM Studio 0.4.20 and Ollama 0.33.2, both return well-formed Anthropic Messages responses, content blocks and
stop_reasonandusageincluded.So a request arrives in a dialect the destination engine already speaks, and the only component in the path that does not is the proxy in the middle.
Desired outcome
POST /v1/messageson the PAIR proxy, alongside the existing endpoints, routed by the same scheduler and answered by the selected node's engine. Observable behavior:tool_useandtool_resultcontent blocks survive intact in both directions. This is the part that matters most for agent clients, because a dropped or reshaped tool block fails silently and looks like a confused model rather than a transport bug./v1/messages/count_tokensbehaves predictably. Agent clients call it, and neither engine implements it: at LM Studio 0.4.20 it returns HTTP 200 carrying{"error":"Unexpected endpoint or method."}, and at Ollama 0.33.2 it returns 404. Passing it through where an engine implements it and returning a clear error where none does would improve on both, and PAIR is not the right place to invent a tokenizer.Passthrough rather than translation looks like the smaller and safer change: when the selected node's engine speaks Anthropic Messages, forwarding the body preserves fidelity by construction and needs no per-field mapping to maintain as the API evolves.
Alternatives considered
Each alternative pays a translation cost to reach engines that already speak the dialect.
Compatibility and security implications
Additive. Existing Ollama-compatible and OpenAI-compatible endpoints are unchanged, and clients that do not use
/v1/messagessee no difference. No change to discovery, pairing, or the trust model is implied.Routing an unfamiliar path to a node whose engine does not implement it should fail with a clear error rather than a partial or reshaped response, so that a capability gap stays visible instead of silent.
Validation approach
tool_useblock with valid arguments, and a streaming request whose SSE event sequence is compared against the Anthropic grammar.count_tokenscall that returns either a clear answer or a clear error, never a hang and never an HTTP 200 carrying an error body.Scope of what I checked
I read the proxy service only, not every service, and I did not locate in source the Ollama-native surface the README describes. The engine behavior above was executed at the pinned versions named; the claim that PAIR has no Anthropic surface today comes from finding no
/v1/messagesor Anthropic reference in the repository.This sits next to two open threads rather than pointing in a fresh direction: #9 adds vLLM as a third engine, which serves the same endpoint, and #7 asks for richer placement signals, which is the same question of whether a node can serve a given request rather than merely being idle.
Happy to test on Apple Silicon and report back, and happy to follow up with a pull request if this is a direction you would take.
Confirmations