Ramp, the corporate card and spend management company, opened its internal model routing engine to outside developers this week under the name Router. The tool sits between an application and its AI providers, sending every inference call to whichever supported model clears a set performance bar at the lowest price, adjusting in real time to the latency and failure rates it observes on each provider.

Ramp says Router cuts AI inference costs by an average of 40 percent, a figure the company presents without a named methodology or an outside audit. That claim sits on Router’s own marketing page next to a chart built from 18 of Ramp’s own sample sets, showing cost falling as a larger share of requests gets shifted onto cheaper models. It is marketing math, not a benchmark another lab has replicated.

The more concrete number in Ramp’s favor comes from its own operations. According to Ramp, running Router across its production workloads for three years cut its internal AI costs by 30 percent without hurting performance. Rahul Sengottuvelu, Ramp’s chief technology officer, is quoted on the same page saying, “At Ramp, Router cut our overall LLM cost by 30 percent while making our features smarter and faster.” That is a company executive crediting a company product, not an independent evaluation.

Router connects to 27 models spanning OpenAI, Anthropic, and open-weight labs including DeepSeek, Kimi, GLM, and Qwen, plus xAI’s Grok. Because its interface mirrors the OpenAI and Anthropic SDKs developers already use, Ramp says switching from a direct provider integration takes a one-line change to an application’s base URL. If a provider degrades or gets rate-limited, Router says it can shift eligible traffic to another model automatically, a fallback pitch aimed at teams tired of building that logic themselves.

Pricing is built to pull in individual developers first. Routing itself is free through the end of 2026, new accounts get $26 in credits, and no Ramp card, corporate account, or business entity is required to sign up. Enterprise features and support for countries beyond the United States are both listed as coming later, a sequencing choice that reads like Ramp is running this launch as a funnel rather than shipping a finished enterprise product.

Model routing is not a new idea. OpenRouter, Martian, and Not Diamond already sell versions of the same arbitrage, and both AWS Bedrock and Azure AI Foundry now ship native prompt routing. What Ramp is selling instead of the routing concept itself is a track record: three years of production traffic pulled from its own expense platform, plus a proprietary benchmark, Ramp SWE-Bench, built from real engineering tickets rather than a public leaderboard. Whether that history persuades developers to route frontier model traffic through a fintech company’s infrastructure, instead of a neutral routing specialist or the cloud provider they already trust, is the question a 40 percent headline number does not answer on its own.

Teams evaluating inference spend over the next quarter should treat Router’s savings figures as a hypothesis to test against their own traffic patterns, not a number to budget around before running the comparison themselves.

This account is based on Router’s product page at router.com, published by Ramp, which carries no dateline or byline date.