NVIDIA Switchyard: Route Claude Code to Your Own vLLM

NVIDIA Switchyard LLM proxy translates OpenAI and Anthropic APIs in Rust, so Claude Code can hit your own vLLM box. Pre-alpha, so keep it off prod.

Your agent speaks the wrong API

Claude Code talks Anthropic Messages. Codex CLI talks OpenAI Chat Completions. The model you actually pay to run, whether that is your own vLLM box, an NVIDIA NIM deployment, or an Ollama instance, talks something else entirely. Rewriting the agent to match the backend is not realistic, so something in the middle has to translate.

On September 2, 2026, NVIDIA published Switchyard (opens in new tab), a Rust proxy and library under Apache 2.0 that does exactly that job. Documentation lives at docs.nvidia.com/nemo/switchyard. NVIDIA labels the project pre-alpha and experimental, with an explicit warning against production use. API and algorithms are expected to change significantly before v1.0.

My position up front, so nobody wastes an afternoon: this belongs on a dev box or a single team’s sandbox, not anywhere near traffic you invoice for. The process sits on the path of every agent turn, which means a crash takes out every developer pointed at it, and the config schema is not stable yet.

Background: why the translation layer has to live outside the agent

If you cannot change the agent, you change what the agent points at. Switchyard decodes the inbound request into provider-neutral Rust types, runs a routing algorithm to pick a backend, re-encodes the request in that backend’s wire format, calls it, and translates the response (streaming events included) back into the shape the client expects. That is the whole job.

The decoupling that matters is on the server side: it accepts OpenAI Chat Completions, OpenAI Responses, and Anthropic Messages. Any of those three inbound formats can address any route, and each configured LLM client picks one upstream format of its own. The agent’s API and the backend’s API no longer have to match.

Homemade shims most people run today fall apart in two specific places: streaming event shapes and tool-call blocks. Both have to be reshaped token-by-token or block-by-block, and that is the work you are actually outsourcing when you hand this to Switchyard. If you have ever watched a tool_use block get mangled mid-stream and had to bisect a FastAPI middleware to find it, you already know why a typed Rust implementation is appealing.

What’s new: three ways to run the NVIDIA Switchyard LLM proxy

There is a launcher path aimed squarely at coding agents. You install it with uv tool install --python 3.12 "nemo-switchyard[cli]", then run switchyard launch claude, switchyard launch codex, or switchyard launch openclaw against a packaged deployment or your own TOML. The launcher handles the client wiring so the agent sees a local endpoint.

There is a server path for teams that want to host the proxy themselves. cargo install --locked switchyard-server gives you the standalone binary, --dry-run validates the config, and you bind the host and port yourself. This is the path I would pick if I wanted to see what Switchyard is actually doing.

There is a library path. switchyard-libsy embeds the routing algorithms without owning an HTTP stack. It never calls a model itself. The algorithm decides which target to use and hands every model call back to your code. If you already have a request pipeline in Rust and want to add routing without adding a network hop, this is the one.

The algorithms are the part that will tempt people into overreach. A route is one client-visible model ID plus the algorithm behind it, and roles inside a route (strong, weak, capable, efficient) are not fixed properties of a model. passthrough sends every request to one target. random splits traffic across targets with optional relative weights and an optional seed that reproduces the selection sequence, which is your A/B and cost-experiment path. llm_classifier calls a classifier target for a capability verdict and routes to weak or strong; base_threshold is required, and mode = "escalation" runs every turn on the weak tier first and lets a judge decide whether to rerun on the strong tier. stage_router scores tool-result and agent-progress signals from recent turns to skip the extra classifier call on most turns.

Is the NVIDIA Switchyard LLM proxy ready for production traffic?

No. NVIDIA says so directly, and I agree.

What does the churn cost me? The TOML you write today across llm_clients, targets, and routes may not parse after an upgrade. Pin the version in your lockfile and expect config rewrites rather than clean bumps. If you treat this like a normal semver-stable crate, you will be debugging TOML at midnight.

Can I trust the numbers? switchyard_routing_overhead_ms reports the algorithm’s run time minus the model call that served the request. Classifier calls are not subtracted, so an llm_classifier route reports its classification time as routing overhead while passthrough and random report the sub-millisecond cost of picking a target. Buckets start at 0.1 ms, so read the histogram knowing which route produced it. The metric is genuinely useful, but only if you remember that “overhead” means different things on different routes.

Where do my credentials sit? api_key_env only names an environment variable, which is the right decision. It is still a plaintext key in the process environment of whatever box you put this on, so the usual rules apply: dedicated runtime user, no shared shell history, rotate on box rebuild. Also note max_retries defaults to 2 and applies to transport failures, timeouts, HTTP 408/429, and 5xx. That is worth understanding before you aim it at a rate-limited upstream, because two retries on a 429 can turn a soft limit into a hard ban.

What to expect before v1.0

Pre-alpha projects usually ship the same kind of breakage: renamed algorithms, moved config keys, changed metric family names. All of that lands on your dashboards.

I cannot answer from the docs how stage_router behaves over long agent sessions, or whether escalation mode’s weak-tier-first pass genuinely saves money once you count the reruns the judge orders. That is an empirical question, and I would want to see real session traces before I trusted either claim.

I also cannot tell you whether the project stays honestly provider-neutral or gradually bends toward NIM. The source release notes do not commit either way, and the answer decides how much of your routing config is portable later.

What would move this from evaluation to something I would run for paying customers? A stable config schema first, then a documented upgrade path with a deprecation table, then a real concurrency test suite. Without those, the proxy is a clever piece of engineering you cannot yet depend on.

I expect, based on the cadence of NVIDIA’s other NeMo releases, that pre-alpha churn will continue for at least two minor versions before anything resembling a stable contract lands. Treat that as my read, not a roadmap.

Do this on one machine, not on your team’s default config

Write a small TOML with two targets, one vLLM client of your own and one hosted model, and validate it with switchyard-server --dry-run before anything else. Point a single switchyard launch claude session at it, scrape GET /metrics for a day, and read the token families and the overhead histogram before you touch a routing algorithm. Turn on --routing-log-file and check GET /v1/routing/session-stats while you are there.

Keep your unproxied configuration intact so reverting is one environment variable and not an evening. That is the rule.