linkgo
Switchyard

Switchyard

AI

An open-source Rust proxy and library that routes LLM traffic across models and providers while preserving native OpenAI and Anthropic API compatibility.

-(0 Reviews)
Free Available
Starting from Free

About Switchyard

Switchyard is a Rust proxy and embeddable library for LLM traffic, built by NVIDIA and released under Apache 2.0. It sits between an application and its model backends, translating between the OpenAI Chat Completions, OpenAI Responses and Anthropic Messages formats so a coding agent such as Claude Code or Codex keeps speaking its native API while the request is actually served by vLLM, NVIDIA NIM, Ollama or any OpenAI-compatible endpoint. Routing is handled by typed, composable algorithms — an LLM-as-classifier router that decides whether a turn needs the weak or strong tier, a stage router that uses signals already present in the conversation such as tool results and errors, an escalation router that runs every turn on the weak tier first and lets a judge decide whether to retry on the strong tier, and random routing for fixed A/B traffic splits — or a custom algorithm you write yourself. Prometheus metrics cover requests, errors, latency, tokens and routing overhead. The project ships two paths: a standalone server binary configured with a routes.toml, and switchyard-libsy, which embeds the routing algorithms in your own Rust application and hands every model call back to you so it drops into an existing proxy, gateway or agent runtime without owning an HTTP stack. Switchyard is explicitly pre-alpha and its API and algorithms are expected to change significantly before v1.0.

Key Features

Protocol Translation: Converts between OpenAI Chat Completions, OpenAI Responses and Anthropic Messages formats so clients keep their native API while any backend serves the request.
Multi-Backend Routing: Spreads traffic across vLLM, NVIDIA NIM, Ollama and any OpenAI-compatible endpoint, letting you point an existing coding agent at an open-source model without changing the agent.
LLM Classifier Router: Uses request content to decide whether a given turn needs the weak or the strong model tier, cutting spend on turns that do not need frontier capability.
Stage Router: Routes most turns from signals already in the conversation — tool results, errors, conversation stage — so no extra model call is needed to make the decision.
Escalation Router: Runs every turn on the weak tier first, then has a judge read that answer and decide whether the same request should be re-sent to the strong tier.
Random Routing for A/B Tests: Applies a fixed traffic split across targets for benchmarking, baselines and cost experiments.
Operational Metrics: Exposes Prometheus metrics for requests, errors, latency, token counts and the overhead added by routing itself.
Server or Library Deployment: Run it as a standalone Rust proxy configured by routes.toml, or embed switchyard-libsy in your own application so it decides the target and hands the model call back to you.

Use Cases

Pointing Coding Agents at Open Models: Serve Claude Code or Codex from vLLM, NIM or Ollama without the agent knowing the API changed.
Cost/Performance Optimization: Send routine turns to a cheap weak-tier model and reserve the strong tier for turns a classifier or judge says need it.
Model A/B Benchmarking: Split traffic on a fixed ratio across two models to compare quality, latency and cost on real production requests.
Provider Migration and Failover: Keep application code on one API shape while swapping or mixing the providers behind it.
Embedding Routing in an Agent Runtime: Drop the routing algorithms into an existing gateway or agent framework via the library path without adopting a new HTTP stack.
Operational Visibility: Track per-route latency, error rates and token spend through Prometheus to find which routes are actually costing money.

Frequently asked questions about Switchyard

What is Switchyard?

An open-source Rust proxy and library that routes LLM traffic across models and providers while preserving native OpenAI and Anthropic API compatibility.

How does Switchyard work?

Switchyard works by combining Protocol Translation: Converts between OpenAI Chat Completions, OpenAI Responses and Anthropic Messages formats so clients keep their native API while any backend serves the request., Multi-Backend Routing: Spreads traffic across vLLM, NVIDIA NIM, Ollama and any OpenAI-compatible endpoint, letting you point an existing coding agent at an open-source model without changing the agent., LLM Classifier Router: Uses request content to decide whether a given turn needs the weak or the strong model tier, cutting spend on turns that do not need frontier capability., Stage Router: Routes most turns from signals already in the conversation — tool results, errors, conversation stage — so no extra model call is needed to make the decision., Escalation Router: Runs every turn on the weak tier first, then has a judge read that answer and decide whether the same request should be re-sent to the strong tier. to help users with Pointing Coding Agents at Open Models: Serve Claude Code or Codex from vLLM, NIM or Ollama without the agent knowing the API changed., Cost/Performance Optimization: Send routine turns to a cheap weak-tier model and reserve the strong tier for turns a classifier or judge says need it., Model A/B Benchmarking: Split traffic on a fixed ratio across two models to compare quality, latency and cost on real production requests., Provider Migration and Failover: Keep application code on one API shape while swapping or mixing the providers behind it., Embedding Routing in an Agent Runtime: Drop the routing algorithms into an existing gateway or agent framework via the library path without adopting a new HTTP stack..

What are the main features of Switchyard?

Key features include Protocol Translation: Converts between OpenAI Chat Completions, OpenAI Responses and Anthropic Messages formats so clients keep their native API while any backend serves the request., Multi-Backend Routing: Spreads traffic across vLLM, NVIDIA NIM, Ollama and any OpenAI-compatible endpoint, letting you point an existing coding agent at an open-source model without changing the agent., LLM Classifier Router: Uses request content to decide whether a given turn needs the weak or the strong model tier, cutting spend on turns that do not need frontier capability., Stage Router: Routes most turns from signals already in the conversation — tool results, errors, conversation stage — so no extra model call is needed to make the decision., Escalation Router: Runs every turn on the weak tier first, then has a judge read that answer and decide whether the same request should be re-sent to the strong tier..

Who is Switchyard for?

Switchyard is useful for anyone interested in Pointing Coding Agents at Open Models: Serve Claude Code or Codex from vLLM, NIM or Ollama without the agent knowing the API changed., Cost/Performance Optimization: Send routine turns to a cheap weak-tier model and reserve the strong tier for turns a classifier or judge says need it., Model A/B Benchmarking: Split traffic on a fixed ratio across two models to compare quality, latency and cost on real production requests., Provider Migration and Failover: Keep application code on one API shape while swapping or mixing the providers behind it., Embedding Routing in an Agent Runtime: Drop the routing algorithms into an existing gateway or agent framework via the library path without adopting a new HTTP stack..

How much does Switchyard cost?

Switchyard is free to use.

How do I get started with Switchyard?

Visit https://github.com/NVIDIA-NeMo/Switchyard to sign up and explore Switchyard.

Explore more AI Ai Services tools

Browse all Ai Services tools →

Compare Switchyard: vs Claude Academy · vs Router by Ramp · vs Supernova · vs bitdrift