Skip to content

feat(proxy): expose committed serving node response header - #2

Draft
kvnloo wants to merge 7 commits into
developfrom
feat/proxy-served-on-header
Draft

kvnloo wants to merge 7 commits into
developfrom
feat/proxy-served-on-header

Conversation

@kvnloo

@kvnloo kvnloo commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

Description

Downstream implementation for NVIDIA#42.

PAIR's published OpenAPI already documents the response header:

X-PAIR-Served-By: <stable host UUID>

for proxied inference responses, but the proxy did not emit it.

This branch implements that existing documented contract at the authoritative commit point in ReverseProxy.ModifyResponse, after retries/failover have settled and before response headers are copied to the client.

Contract

  • names the candidate that actually committed the response;
  • after failover, reports the successful final node rather than the first attempted node;
  • proxy-generated rejection/errors with no committed upstream do not invent the header;
  • an upstream-supplied spoofed value is overwritten with PAIR's authoritative stable node ID;
  • uses the existing documented header name X-PAIR-Served-By;
  • no routing or CORS policy changes.

Release intent

Changelog title

Expose the serving node on proxied inference responses

Changelog body

PAIR now returns the documented X-PAIR-Served-By response header with the stable host UUID of the node that actually served a proxied inference request, including after failover.

Bumps

  • services: minor
  • nvpair-cluster-manager: none
  • nvpair-engine-manager: none
  • nvpair-errors: none
  • nvpair-job-scheduler: none
  • nvpair-manual-nodes: none
  • nvpair-node-info: none
  • nvpair-node-scanner: none
  • nvpair-node-settings: none
  • nvpair-proxy: minor
  • nvpair-tui: none
  • nvpair-ui-broker: none
  • nvpair-workload-manager: none

Scope

In scope:

  • implement the already-documented serving-node response header;
  • preserve authoritative node identity after failover;
  • prevent upstream spoofing of the PAIR-owned header.

Out of scope:

  • adding prompt/response logging;
  • changing routing policy;
  • changing CORS policy;
  • exposing private node addresses.

Validation

Focused proxy regressions cover:

  • single healthy target -> header names that node;
  • retry/failover -> header names the successful final node;
  • local/no-target error -> header absent;
  • non-inference control request -> header absent;
  • upstream value cannot override PAIR's authoritative value.

Fork Build passed on the initial branch. Fresh CI is running after aligning the implementation with the documented header, narrowing it to inference responses, and adding the required release-intent contract.

Risk

The value is the existing stable host UUID already used by PAIR for workload scheduledOn / proxy node attribution. No prompt, response body, credentials, hostname, or network address is exposed.

All commits are DCO-signed and the branch is based on upstream develop @ 937cf6c.

kvnloo added 7 commits October 1, 2026 01:05
Signed-off-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
Signed-off-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
Signed-off-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
Signed-off-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
Signed-off-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
Signed-off-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
Signed-off-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>

kvnloo commented Oct 1, 2026

Copy link
Copy Markdown
Owner Author

[READY TO POST] blocking review — upstream NVIDIA#98

Upstream review submission returned 403, so this is staged here for promotion.

NVIDIA#98 has a stale-registry recovery bug:

  • when the registry matches none of this boot's LUIDs, selectPhysicalAdapters keeps all DXGI candidates, potentially including the phantom duplicate;
  • those become the startup/static GPU inventory;
  • after the registry becomes correct, the refresh can return only the real adapter;
  • but mergeGPUInventory(static, recovered) is additive/enrichment-only and never removes a static statsKey missing from the recovered set.

So:

  1. stale boot gate -> static = [real, phantom]
  2. fresh gate -> recovered = [real]
  3. merge -> [real, phantom]

The duplicate persists until restart.

Suggested fix: Windows refresh needs an authoritative replacement path once the LUID gate is valid, rather than always using the Darwin-oriented additive merge.

Regression: start with two same-model candidates under a stale gate; refresh to a valid registry containing one LUID and assert the phantom disappears. Also verify two genuine same-model GPUs remain when both LUIDs are valid.

Status: READY TO POST to upstream PR NVIDIA#98 when org-side writes are available.

kvnloo commented Oct 1, 2026

Copy link
Copy Markdown
Owner Author

[READY TO POST] blocking review — upstream NVIDIA#82

Upstream review submission returned 403; staging here for promotion.

The PR serializes the zone returned by outboundIP(target), but that function runs on the inviter, so the zone is the inviter's local interface name. IPv6 link-local scopes are local to the dialing host. The original repro uses macOS en0 on both nodes, but an inviter on en0 and joiner on eth0/wlan0 would receive an unusable remote scope.

The repo already has the correct pattern in the opposite direction: onInviterPaired preserves the advertised listening port but replaces the host with sess.peerSrcIP, the address actually observed on this machine.

Suggested fix for Completion:

  • capture hostOnly(r.RemoteAddr) from the Initial Exchange on the joiner;
  • after parsing inviter PairingInfo, keep its advertised listening port but prefer that observed source host for sess.addr;
  • keep peerURL solely for RFC 6874 URL encoding.

Regression: inviter-local zone en0, joiner-observed source zone eth0; Completion must dial %25eth0, not %25en0. Preserve the same-zone macOS case too.

The doc's current "both hosts must use the same interface name" limitation should then be unnecessary.

Status: READY TO POST to upstream PR NVIDIA#82 when org-side writes are available.

kvnloo commented Oct 1, 2026

Copy link
Copy Markdown
Owner Author

[READY TO POST] blocking review — upstream NVIDIA#38

Upstream review submission returned 403; staging here for promotion.

NVIDIA#38 currently authenticates a LAN request using either Authorization: Bearer <PAIR key> or X-Api-Key: <PAIR key>, then StripCredential unconditionally deletes both headers.

That makes the new LAN ingress incompatible with an upstream engine that has its own credential. Example:

  • X-Api-Key: <PAIR key>
  • Authorization: Bearer <vLLM key>

PAIR admits the request using X-Api-Key, then deletes the unrelated Authorization header before forwarding, so vLLM receives no engine credential. The symmetric case has the same problem.

Suggested invariant: never forward the credential that authenticated PAIR, but preserve unrelated upstream credentials. The auth decision should retain which header/value matched and strip only PAIR-owned credential material. A dedicated proxy-auth header would fully avoid same-header ambiguity.

Regression:

  1. PAIR key A configured;
  2. LAN request carries X-Api-Key A + distinct engine Authorization;
  3. request is admitted;
  4. upstream sees engine Authorization and never sees A.
    Also test the symmetric arrangement.

Status: READY TO POST to upstream PR NVIDIA#38 when org-side writes are available.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant