Repository navigation
[Feature]: Allow authenticated opt-in LAN access to local compatibility endpoints #28
Description
Activity
Same need from the Kubernetes side, with a working pattern you may find useful in the meantime.
I run a services-only node as a pod with host networking on a single-node OpenShift cluster, paired to a Linux GPU node and a macOS node. In-cluster workloads hit the same
403 loopback-onlyyou quote. What made it work without touching the policy: a small relay in the same pod that listens on a second port and re-originates each request from127.0.0.1, so the node's proxy sees a loopback client. AServicefronts the relay, so workloads callhttp://pair.pair.svc.cluster.local:1234/v1and never see the node's LAN address. ANetworkPolicykeeps the relay inside the cluster, which is the trust boundary here since the relay itself does no authentication.Evidence from that setup, PAIR v0.1.1: a five-minute soak from an in-cluster Job, 71 of 71 requests OK, routed to the GPU node over the cluster's mTLS; pod restart kept the node's identity and membership.
The relay, Dockerfile, and manifests are in #31, and #32 asks the maintainers whether a container is a distribution channel they want. A native, authenticated opt-in like you describe would be the better answer for the Windows workstation case, where there is no pod to put a relay in; the relay is what you can do today when the client can run in the same host boundary as the node.
PR #38 implements this as an opt-in: #38
With nothing configured, behavior is unchanged — a non-loopback plaintext request still gets
403 loopback-only. An operator enables LAN access by configuring at least one API key:NVPAIR_PROXY_API_KEYS_FILE— key file, one key per line (defaultproxy-api-keysin the PAIR data directory, consulted only if it exists; must be0600and owned by the proxy's user on Unix)NVPAIR_PROXY_API_KEYS— inline, comma-separated, for containers and headless nodesNVPAIR_PROXY_ALLOWED_CIDRS— optional source allowlist, checked before the key
A LAN client then sends the key as
Authorization: Bearer <key>(orX-Api-Key: <key>) and is routed exactly like a loopback client; the key is stripped before the request is forwarded. Loopback callers are never asked for a key, so the desktop app and TUI are unaffected. Keys hot-reload from the file, so rotation and revocation need no restart. Full details, validation on macOS and Linux, and a live two-node LAN check are in the PR.@DustinTrap — this is the native path for the case you described in #31/#32: a workload sends a bearer key to the node's own proxy instead of relaying through a loopback sidecar, and the CIDR allowlist can pin it to the cluster's pod network. If you get a chance to try the branch against your OpenShift setup, a note on the PR either way would be useful.
Proudly Made in Nebraska. Go Big Red! 🌽 https://xkcd.com/2347/
Tried #38 on the OpenShift setup and left the results there. It works for a Kubernetes workload with the key passed inline from a Secret; the key-file path fails closed under a pod fsGroup, details on the PR. It is the right native answer to this issue, and once it lands the relay I described above becomes unnecessary.
Reacted by Aaron K. ClarkTried #38 on the OpenShift setup and left the results there. It works for a Kubernetes workload with the key passed inline from a Secret; the key-file path fails closed under a pod fsGroup, details on the PR. It is the right native answer to this issue, and once it lands the relay I described above becomes unnecessary.
Thank you for confirming!! Fingers crossed that the PR gets merged!!
Thanks for trying PAIR out! I'm the tech lead for PAIR and we focused on the prosumer use case for our initial release, but I will work with the team to assess the security requirements for your deployment and evaluate the pending PRs.
Reacted by Aaron K. ClarkThank you so much for responding. If this decision is to keep security work in-house I completely understand! I am just excited at the opportunity to give back to NVIDIA and the Open-Source Community!\ Thank you to you and your team for everything you do!
We would like to be open and I will share any security concerns about this feature that come up here for discussion - in this case what I did was ask the relevant security people to consider what changes if any would be needed to the existing posture to support this feature. I'm excited to see use cases beyond the home prosumer landscape and please, keep them coming!
Reacted by Aaron K. ClarkFollowing up after @sherief-nv's 09-09/09-10 notes — thanks for taking the security review seriously and keeping the discussion in the open.
Supporting @CryptoJones's point about #38: the PR's default-deny posture (non-loopback plaintext still rejected unless a token is explicitly configured) is exactly the right shape for this — it's an opt-in LAN exposure for trusted clients, not a loosening of the loopback policy. The OpenShift-side testing @DustinTrap did (inline Secret key works; key-file path fails closed under a pod user) also already maps the failure modes: if the team adopts the key-file variant, it needs a permission/pre-flight check with a clear error rather than a silent closed door.
My original use case, for the security review's scope: a dedicated Windows GPU workstation serving automation clients on other LAN workstations through an established firewall/NAT path. Today the only supported alternative is mTLS ingress, which requires cluster membership for every consumer — disproportionate for a few trusted static workstations on an already-firewalled LAN segment. An authenticated token gate on the existing compatibility endpoint (what #38 implements) covers it with no change to default behavior.
No rush on the security side — just noting demand is live on three independent setups (Windows workstation, OpenShift pod, and whatever else joins the thread), so this stays on the radar.
Reacted by Aaron K. Clark
Feature request
Please provide a supported, explicit opt-in way to expose PAIR's OpenAI-compatible local endpoint to trusted LAN clients.
Current behavior
PAIR 0.91.7 on Windows 11 binds the LM Studio-compatible proxy on port 1234, but rejects a request arriving through an established firewall/NAT path:
{"code":"loopback-only","error":"plaintext requests are accepted only from loopback; cluster peers must use the mTLS ingress"}Changing the proxy port preserves the port number but does not change this policy.
Use case
A dedicated Windows GPU workstation already serves:
Installing and pairing PAIR inside every client, pod, or workload is substantially more operationally complex than exposing one authenticated endpoint.
Security expectations
I understand why unrestricted plaintext LAN exposure is disabled and am not asking for an insecure default. A secure opt-in mode could require some combination of:
0.0.0.0;Requested outcome
Provide a documented, supported mechanism for ordinary OpenAI-compatible clients on trusted networks to access a PAIR endpoint without running a PAIR node locally. The existing loopback-only behavior should remain the default.