Cut your token costs by up to 40%
Eden AI routes every request to the cheapest provider that meets your bar, serves repeated requests from cache, falls back without paying twice and returns the cost of every call. Same OpenAI-compatible API, 500+ models, provider prices with no markup.
Built to shrink the bill, not the quality.
Six levers, all at the gateway
Most AI spend goes to habits: one frontier model for everything, retries that bill twice, prompts repeated word for word, and nobody watching the meter. Each lever below works without changing your application code.
Cost-based smart routing
Leave the provider off the model name and Eden AI sorts providers by cost, sending more traffic to the cheaper ones. Sort by speed or latency when that matters more.
The right model for each step
Agents spend most turns reading, searching and verifying. Send those to small models and keep frontier models for the hard steps.
Fallbacks billed once
Declare up to three fallback models or providers. A failed attempt is never billed; you pay only for the attempt that produced a response.
Response caching
Identical requests are served from cache at no additional cost. On by default, toggled per project.
Prompt caching passthrough
Reuse stable prompt prefixes with OpenAI, Anthropic, Bedrock, Gemini and DeepSeek at the provider's discounted cache-read rate.
Caps per key, cost per call
Give each API key a budget that resets daily, weekly or monthly, and read the cost of every request in its response.
Same quality, 82% less per request
We ran the same five task types through a single frontier model and through Smart Routing, graded both, and compared the bill. The 82% below is the result measured in that benchmark, on a controlled workload. The "up to 40%" in the page title is the conservative estimate we give for real production traffic, which mixes simple and complex requests; how close you land to the benchmark depends on how much of your traffic a cheaper model can serve.
Teams that compare before they commit

We use Eden AI because of its standard interface that connects into various AI providers, so we can test & compare the accuracy and manage vendor risks down the road.
Jan Vosecky
Head of Product @ Re-Hub

I'm using Eden AI for parsing identity documents, it works very well and is much more cost-effective than the previous solution I was using. With Eden AI, I'm billed per request, and not only is it much more affordable, but also I can have backup OCR providers for improved reliability.
Jose Vilches
CEO @ SIM Store
What would it mean for your bill?
Move the sliders to your own numbers. The estimate uses only two levers, cheaper models on routable traffic and the response cache, so fallbacks and prompt caching are upside.
Three steps to a smaller invoice
Point your SDK at Eden AI
Change the base URL of your OpenAI-compatible client and keep your code. Prompts, tools, streaming and structured outputs keep working as before.
Pick a strategy, add fallbacks
Sort providers by cost, speed or latency, restrict the pool to approved providers or the EU region, and list fallback models. It is a few fields in the request, not a rewrite.
Set your quality bar, then cut the cost
Compare models on accuracy, latency and price before you move traffic, then read the cost field on every response and cap each key.
Spend less without giving up quality
Three places where the gateway does the work: how each key routes, what the catalog lets you compare before you move traffic, and how little has to change in your code.
A routing strategy per key, set once
Sort providers by cost, speed or latency, narrow the pool to approved providers, pin regulated traffic to Europe with the @eu suffix, keep a conversation on one provider with sticky routing, and list fallbacks. Configure it in the dashboard or per request.
Compare price, latency and region before you commit
The model catalog lists every model with its context window, input and output price per million tokens, capabilities and hosting region. The same task can cost 2 to 3x more on one provider than on another, so choose with the numbers in front of you.

Keep your SDK, and your own keys if you want
Eden AI is a drop-in for the OpenAI SDK. With Bring Your Own Keys you keep your negotiated provider prices while Eden AI still handles routing, fallbacks, caching and cost monitoring.

Calling providers directly vs. Eden AI
Direct provider integrations
Eden AI
One frontier model for every request, including the simple ones.
Cost-based routing sends each request to the cheapest capable provider.
Prices checked by hand, once, when you integrate.
Live price, latency and region comparison across 50+ providers.
Client-side retries that bill for every attempt.
Fallbacks across models and providers, billed only for the attempt that answers.
Repeated prompts paid in full every time.
Response caching at no extra cost and provider prompt-cache passthrough.
Spend discovered on the invoice at month end.
Cost on every response, budgets per key, one consolidated invoice.
Switch the base URL, keep your code
Eden AI is OpenAI-compatible. Routing and fallbacks are fields on the request, and the response tells you which provider answered and what it cost.
Questions we're usually asked
How does Eden AI reduce token costs?
Six levers, all at the gateway, none of them touching your application code.
Cost-based routing sends each request to the cheapest provider that can serve the model you asked for.
The right model for each step keeps frontier models for the hard turns.
Fallbacks bill only the attempt that produced a response.
Response caching serves repeated requests at no cost.
Prompt caching passthrough bills cached input at the provider's discounted rate.
And a budget cap on every key, with the cost returned on each call, stops runaway spend.
Where does the “up to 40%” figure come from?
It is a conservative estimate for a production workload that mixes simple and complex requests. In our published benchmark, Smart Routing cut cost per request by 82% versus a single frontier model, with a quality score of 9.40 versus 9.48. Your savings depend on how much of your traffic cheaper models can serve and how often your prompts repeat.
Does Eden AI mark up provider prices?
No. The prices in the model catalog are the providers' exact prices. The only additional cost is a 5.5% platform fee applied when you buy credits. There is no subscription and no API call limit.
Will routing to cheaper models hurt quality?
You stay in control. Choose the model yourself and let Eden AI only pick the cheapest provider for it, or route among models you have approved. Restrict providers with allowed_providers, pin a conversation to one provider with sticky routing, and compare models on accuracy, latency and price in the catalog before you move traffic.
Can I cap spend?
Yes. Each API key can carry its own budget that resets daily, weekly or monthly. When it reaches $0 the key stops working until the next period. Guardrails add rate limits, spend caps and model allow-lists per role, member or key.
Can I use my own provider API keys?
Yes. With Bring Your Own Keys you plug in your OpenAI, Google, Anthropic or other keys and keep your negotiated prices, while still using Eden AI routing, fallbacks, caching and monitoring.
Do I need to change my code?
No. Eden AI exposes an OpenAI-compatible chat completions endpoint. Point your existing SDK at https://api.edenai.run/v3 and add routing or fallback fields to the request body when you want them.
Does the gateway add latency?
Eden AI adds a thin layer in front of the provider. Sort by latency or speed when response time matters more than price, and use response caching to answer repeated requests instantly.
How do I see what I spend?
Every response includes a cost field in USD. The monitoring dashboard filters and groups usage by API key, member, model, provider, feature and tag, and the management API returns the same data for your own reporting.
Is there a subscription or a minimum spend?
No. There is no subscription and no API call limit. You buy credits and pay the providers' own prices, plus a 5.5% platform fee applied at checkout.
Can I keep regulated traffic in Europe?
Yes. Point those calls at the EU endpoint and only models eligible for European processing will serve them. A model that is not eligible is refused with HTTP 451 before the provider is contacted, so the request never spends credits and never leaves the European boundary.
What is the rate limit?
A per-second rate limit applies by default and can be raised on request. Advanced plans get a higher custom limit, and guardrails can set a lower one per role, member or key.
Pay provider prices, not provider habits.
One API, 500+ models, cost-based routing, fallbacks and caching built in. Get your key and see the cost of your first request in the response.
NO MARKUP ON PROVIDER PRICES • 5.5% PLATFORM FEE ON CREDITS • NO SUBSCRIPTION