Cost optimization

Cut your token costs by up to 40%

Eden AI routes every request to the cheapest provider that meets your bar, serves repeated requests from cache, falls back without paying twice and returns the cost of every call. Same OpenAI-compatible API, 500+ models, provider prices with no markup.

NO MARKUP ON PROVIDER PRICES • PAY ONLY FOR WHAT YOU USE • SOC 2 & ISO 27001 CERTIFIED
Monitoring · Cost overview
Platform
Models
Monitoring
API keys
Routing
Organization
Members
Guardrails
Billing
Sep 1 – Sep 30, 2026All API keysCompared with direct provider pricesExport CSV
Spend
$4,500
−39% vs. direct prices
Saved this month
$2,870
routing · cache · fallbacks
Cache hits
23%
441k requests served at $0.00
Fallbacks
142
every one billed once
Daily spend · direct provider prices vs. with Eden AI
Sep 1Sep 10Sep 20Sep 30
Direct provider pricesWith Eden AI
Where the savings came from
Cheaper providers
$1,842
Response cache
$612
Prompt caching
$318
Fallbacks billed once
$98
Recent requests
TimeModel askedServed byLatencyCost
14:02:11gpt-5.4-nanoopenai · eu-west143 ms$0.000202
14:02:09claude-sonnet-5-5anthropic1.2 s$0.001784
14:02:08gpt-5.4-nanocache hit2 ms$0.000000
14:02:05mistral-large-latestmistral fallback890 ms$0.001400
14:02:01gemini-3.8-flashgoogle310 ms$0.000669
Spending cap · prod-support
$640 / $1,000
64%
Resets monthly · key stops at $0

Built to shrink the bill, not the quality.

Lower token spend
Up to 40%
AI models
500+
Providers
50+
Markup on provider prices
0%
No markup on provider prices
Cost on every response
Budget caps per API key
Response caching included
Where the savings come from

Six levers, all at the gateway

Most AI spend goes to habits: one frontier model for everything, retries that bill twice, prompts repeated word for word, and nobody watching the meter. Each lever below works without changing your application code.

Cost-based smart routing

Leave the provider off the model name and Eden AI sorts providers by cost, sending more traffic to the cheaper ones. Sort by speed or latency when that matters more.

ProviderLatency$ / M inTraffic
deepinfra410 ms$0.0952%
azure · eu180 ms$0.1531%
openai160 ms$0.1517%
routing.sort = cost · cheaper providers receive more traffic

The right model for each step

Agents spend most turns reading, searching and verifying. Send those to small models and keep frontier models for the hard steps.

Fallbacks billed once

Declare up to three fallback models or providers. A failed attempt is never billed; you pay only for the attempt that produced a response.

Response caching

Identical requests are served from cache at no additional cost. On by default, toggled per project.

Prompt caching passthrough

Reuse stable prompt prefixes with OpenAI, Anthropic, Bedrock, Gemini and DeepSeek at the provider's discounted cache-read rate.

Input tokensRateCost
Cached prefix · 1.2M0.1×$0.36
Fresh · 0.4M1×$1.20
The cost field already includes the provider's cache-read price

Caps per key, cost per call

Give each API key a budget that resets daily, weekly or monthly, and read the cost of every request in its response.

Benchmark

Same quality, 82% less per request

We ran the same five task types through a single frontier model and through Smart Routing, graded both, and compared the bill. The 82% below is the result measured in that benchmark, on a controlled workload. The "up to 40%" in the page title is the conservative estimate we give for real production traffic, which mixes simple and complex requests; how close you land to the benchmark depends on how much of your traffic a cheaper model can serve.

Customer stories

Teams that compare before they commit

Jan Vosecky, Head of Product at Re-Hub

We use Eden AI because of its standard interface that connects into various AI providers, so we can test & compare the accuracy and manage vendor risks down the road.

Jan Vosecky

Head of Product @ Re-Hub

Jose Vilches, CEO at SIM Store

I'm using Eden AI for parsing identity documents, it works very well and is much more cost-effective than the previous solution I was using. With Eden AI, I'm billed per request, and not only is it much more affordable, but also I can have backup OCR providers for improved reliability.

Jose Vilches

CEO @ SIM Store

Savings estimate

What would it mean for your bill?

Move the sliders to your own numbers. The estimate uses only two levers, cheaper models on routable traffic and the response cache, so fallbacks and prompt caching are upside.

Your workload
Monthly AI spend today
Requests a smaller model could handle just as well
The easy ones: classification, extraction, short summaries, routine replies. Frontier models keep the hard turns.
Requests that come back word for word (cache hits)
Identical prompts served from cache at no cost. Common with FAQs, templates and fixed instructions.
Estimated monthly savings
Estimate only. The low end assumes cheaper models cost 35% less on routable traffic and 80% of repeated requests hit the cache; the high end assumes 60% and 100%. Your actual figure depends on your prompts and traffic mix.
How it works

Three steps to a smaller invoice

Connect

Point your SDK at Eden AI

Change the base URL of your OpenAI-compatible client and keep your code. Prompts, tools, streaming and structured outputs keep working as before.

Route

Pick a strategy, add fallbacks

Sort providers by cost, speed or latency, restrict the pool to approved providers or the EU region, and list fallback models. It is a few fields in the request, not a rewrite.

Measure

Set your quality bar, then cut the cost

Compare models on accuracy, latency and price before you move traffic, then read the cost field on every response and cap each key.

Cost control in practice

Spend less without giving up quality

Three places where the gateway does the work: how each key routes, what the catalog lets you compare before you move traffic, and how little has to change in your code.

01 · Routing

A routing strategy per key, set once

Sort providers by cost, speed or latency, narrow the pool to approved providers, pin regulated traffic to Europe with the @eu suffix, keep a conversation on one provider with sticky routing, and list fallbacks. Configure it in the dashboard or per request.

02 · Comparison

Compare price, latency and region before you commit

The model catalog lists every model with its context window, input and output price per million tokens, capabilities and hosting region. The same task can cost 2 to 3x more on one provider than on another, so choose with the numbers in front of you.

Eden AI model catalog showing price, latency and region for each model
03 · Integration

Keep your SDK, and your own keys if you want

Eden AI is a drop-in for the OpenAI SDK. With Bring Your Own Keys you keep your negotiated provider prices while Eden AI still handles routing, fallbacks, caching and cost monitoring.

Code console calling Eden AI with an OpenAI-compatible client
Comparative table

Calling providers directly vs. Eden AI

Direct provider integrations

Eden AI

One frontier model for every request, including the simple ones.

Cost-based routing sends each request to the cheapest capable provider.

Prices checked by hand, once, when you integrate.

Live price, latency and region comparison across 50+ providers.

Client-side retries that bill for every attempt.

Fallbacks across models and providers, billed only for the attempt that answers.

Repeated prompts paid in full every time.

Response caching at no extra cost and provider prompt-cache passthrough.

Spend discovered on the invoice at month end.

Cost on every response, budgets per key, one consolidated invoice.

Developer experience

Switch the base URL, keep your code

Eden AI is OpenAI-compatible. Routing and fallbacks are fields on the request, and the response tells you which provider answered and what it cost.

PYTHON · OPENAI SDK
# pip install openai
from openai import OpenAI

client = OpenAI(
    base_url="https://api.edenai.run/v3",
    api_key=EDENAI_API_KEY,
)

response = client.chat.completions.create(
    model="gpt-5.4-nano",        # no provider prefix: Eden AI picks the provider
    messages=[{"role": "user", "content": "Summarize this ticket: ..."}],
    extra_body={
        "routing": {"sort": "cost", "allow_fallbacks": True},
        "fallbacks": ["mistral/mistral-small-latest"],
    },
)
print(response.choices[0].message.content)
RESPONSE · COST INCLUDED
{
  "id": "chatcmpl-8f1c…",
  "model": "openai/gpt-5.4-nano",   // the provider that served the request
  "choices": [
    { "message": { "role": "assistant", "content": "The customer reports…" } }
  ],
  "usage": {
    "prompt_tokens": 412,
    "completion_tokens": 96,
    "total_tokens": 508
  },
  "cost": 0.000202                   // USD, on every response
}

Questions we're usually asked

How does Eden AI reduce token costs?

Six levers, all at the gateway, none of them touching your application code.
Cost-based routing sends each request to the cheapest provider that can serve the model you asked for.
The right model for each step keeps frontier models for the hard turns.
Fallbacks bill only the attempt that produced a response.
Response caching serves repeated requests at no cost.
Prompt caching passthrough bills cached input at the provider's discounted rate.
And a budget cap on every key, with the cost returned on each call, stops runaway spend.

Where does the “up to 40%” figure come from?

It is a conservative estimate for a production workload that mixes simple and complex requests. In our published benchmark, Smart Routing cut cost per request by 82% versus a single frontier model, with a quality score of 9.40 versus 9.48. Your savings depend on how much of your traffic cheaper models can serve and how often your prompts repeat.

Does Eden AI mark up provider prices?

No. The prices in the model catalog are the providers' exact prices. The only additional cost is a 5.5% platform fee applied when you buy credits. There is no subscription and no API call limit.

Will routing to cheaper models hurt quality?

You stay in control. Choose the model yourself and let Eden AI only pick the cheapest provider for it, or route among models you have approved. Restrict providers with allowed_providers, pin a conversation to one provider with sticky routing, and compare models on accuracy, latency and price in the catalog before you move traffic.

Can I cap spend?

Yes. Each API key can carry its own budget that resets daily, weekly or monthly. When it reaches $0 the key stops working until the next period. Guardrails add rate limits, spend caps and model allow-lists per role, member or key.

Can I use my own provider API keys?

Yes. With Bring Your Own Keys you plug in your OpenAI, Google, Anthropic or other keys and keep your negotiated prices, while still using Eden AI routing, fallbacks, caching and monitoring.

Do I need to change my code?

No. Eden AI exposes an OpenAI-compatible chat completions endpoint. Point your existing SDK at https://api.edenai.run/v3 and add routing or fallback fields to the request body when you want them.

Does the gateway add latency?

Eden AI adds a thin layer in front of the provider. Sort by latency or speed when response time matters more than price, and use response caching to answer repeated requests instantly.

How do I see what I spend?

Every response includes a cost field in USD. The monitoring dashboard filters and groups usage by API key, member, model, provider, feature and tag, and the management API returns the same data for your own reporting.

Is there a subscription or a minimum spend?

No. There is no subscription and no API call limit. You buy credits and pay the providers' own prices, plus a 5.5% platform fee applied at checkout.

Can I keep regulated traffic in Europe?

Yes. Point those calls at the EU endpoint and only models eligible for European processing will serve them. A model that is not eligible is refused with HTTP 451 before the provider is contacted, so the request never spends credits and never leaves the European boundary.

What is the rate limit?

A per-second rate limit applies by default and can be raised on request. Advanced plans get a higher custom limit, and guardrails can set a lower one per role, member or key.

let’s start

Pay provider prices, not provider habits.

One API, 500+ models, cost-based routing, fallbacks and caching built in. Get your key and see the cost of your first request in the response.

NO MARKUP ON PROVIDER PRICES • 5.5% PLATFORM FEE ON CREDITS • NO SUBSCRIPTION