

Open-source AI gateway that routes every model through one endpoint at provider cost, then improves that traffic with caching, routing and fine-tuning.

Open-source AI gateway that routes every model through one endpoint at provider cost, then improves that traffic with caching, routing and fine-tuning.
Experiential Labs is an open-source AI gateway that puts hosted providers, your own provider keys and your own GPUs behind a single OpenAI-style endpoint, charging zero markup on routed tokens so you pay each provider's list price. On top of the gateway sits an intelligence layer that watches real production traffic and turns it into improvements: it flags when switching to a different or newly released model would win, surfaces caching opportunities that return repeated tokens at up to 90% off, and can fine-tune a model on your own traffic that is proved in closed-loop simulation before it serves and then reached through the same endpoint. A console reports every request and every dollar across the organization, broken down by agent, person, model or day, with catalog, usage limits and live request logs in one place. Published case studies describe a 9B model distilled for computer use that cut cost 97% while completing more tasks, and a 4B model trained with GRPO that improved claims-research verdict accuracy 10.9% at 90% lower cost. The company is backed by Y Combinator, and the gateway itself is free and open source while revenue comes from hosted inference and the Pro plan.
Open-source AI gateway that routes every model through one endpoint at provider cost, then improves that traffic with caching, routing and fine-tuning.
Experiential Labs works by combining Unified Model Endpoint: One OpenAI-compatible POST endpoint fronts every hosted provider, your own bring-your-own keys and your own GPUs, so switching models is a parameter change rather than an integration., Zero-Markup Routed Tokens: Routed traffic bills at the provider's list price with 0% added on top, with the company earning on hosted inference and the Pro plan instead of on your tokens., Model Recommendation from Traffic: The intelligence layer watches real request patterns and tells you when switching models would win, including newly released models on the day they ship, with optional per-prompt optimization., Caching Opportunity Detection: Identifies where cache hit rate could improve and shows the projected savings, with repeated tokens returning at 90% off once enabled., Traffic-Trained Custom Models: Fine-tunes a model on your own traffic and proves it in closed-loop simulation before it ever serves, then exposes it through the same endpoint you already call. to help users with Consolidating Multi-Provider Access: Replace separate SDKs and keys for OpenAI, Anthropic, Google and others with a single endpoint and key across every application., Cutting Inference Spend: Use caching recommendations and model-switch suggestions to lower the cost of an existing production workload without changing application code., Replacing a Frontier Model with a Small One: Distill or fine-tune a small model on your own traffic for a narrow repetitive task and serve it at a fraction of frontier-model cost and latency., Chargeback and Budgeting: Attribute AI spend to individual agents, teams or people for internal cost allocation and to enforce per-key budget caps., Evaluating New Model Releases: Compare a newly shipped model against your current one on your own traffic before committing to a migration..
Key features include Unified Model Endpoint: One OpenAI-compatible POST endpoint fronts every hosted provider, your own bring-your-own keys and your own GPUs, so switching models is a parameter change rather than an integration., Zero-Markup Routed Tokens: Routed traffic bills at the provider's list price with 0% added on top, with the company earning on hosted inference and the Pro plan instead of on your tokens., Model Recommendation from Traffic: The intelligence layer watches real request patterns and tells you when switching models would win, including newly released models on the day they ship, with optional per-prompt optimization., Caching Opportunity Detection: Identifies where cache hit rate could improve and shows the projected savings, with repeated tokens returning at 90% off once enabled., Traffic-Trained Custom Models: Fine-tunes a model on your own traffic and proves it in closed-loop simulation before it ever serves, then exposes it through the same endpoint you already call..
Experiential Labs is useful for anyone interested in Consolidating Multi-Provider Access: Replace separate SDKs and keys for OpenAI, Anthropic, Google and others with a single endpoint and key across every application., Cutting Inference Spend: Use caching recommendations and model-switch suggestions to lower the cost of an existing production workload without changing application code., Replacing a Frontier Model with a Small One: Distill or fine-tune a small model on your own traffic for a narrow repetitive task and serve it at a fraction of frontier-model cost and latency., Chargeback and Budgeting: Attribute AI spend to individual agents, teams or people for internal cost allocation and to enforce per-key budget caps., Evaluating New Model Releases: Compare a newly shipped model against your current one on your own traffic before committing to a migration..
Experiential Labs offers a free tier with paid plans for advanced features.
Visit https://www.experientiallabs.ai/ to sign up and explore Experiential Labs.
Compare Experiential Labs: vs Cadenya · vs OpenObserve · vs Dial · vs Articos