Gemma 3 12B IT

gemma-3-12b-it by Google Cloud Released March 14, 2025
ReasoningFunction callingStructured outputPrompt cachingWeb searchComputer useImage inputFile inputAudio inputVideo inputEU hostingFree endpointUp to % off
Input from
$0.05 / 1M
Output from
$0.15 / 1M
Cached input from
$0.015 / 1M
Context window
131K
Endpoints
5
from 3 providers
Regions
Global, Europe, United States

About Gemma 3 12B IT

Gemma 3 12B is a state-of-the-art multimodal language model built and trained by Google. The model supports a context length of 128K tokens and can analyze images and text. With support for over 140 languages and optimized for dialogue use cases, Gemma 3 12B is aligned with human preferences for helpfulness and safety. This endpoint is hosted by Databricks.

Gemma 3 12B IT is a chat model by Google Cloud, available on Eden AI through 5 endpoints from 3 providers, from $0.05 per million input tokens and $0.15 per million output tokens, with a context window of up to 131K tokens.

Providers and prices

5 endpoints from 3 providers, cheapest first, in USD per million tokens. Every one is called with the same Eden AI key: pin one with its endpoint ID, or let Eden AI route for you. .

Endpoint
Region
Context
Input / 1M
Output / 1M
Cached / 1M
Supports
DeepInfra (United States)Official−%deepinfra/google/gemma-3-12b-it
United States
131K
$0.05$0.05
$0.15
$
ReasoningToolsJSONCacheVisionWeb
Amazon Web Services (Europe)Official−%amazon/google.gemma-3-12b-it
Europe
128K
$0.09$0.09
$0.29
$
ReasoningToolsJSONCacheVisionWeb
Amazon Web Services (United States)Official−%amazon/google.gemma-3-12b-it@us
United States
128K
$0.09$0.09
$0.29
$
ReasoningToolsJSONCacheVisionWeb
Databricks (Europe)Official−%databricks/databricks-gemma-3-12b@eu
Europe
128K
$0.15$0.15
$0.5
$0.015
ReasoningToolsJSONCacheVisionWeb
DatabricksOfficial−%databricks/databricks-gemma-3-12b
Global
128K
$0.15$0.15
$0.5
$0.015
ReasoningToolsJSONCacheVisionWeb

Performance

Measured on real Eden AI traffic over the last 30 days. Endpoints with too little traffic are not shown.

Fastest first token
1.53 s
Databricks (Europe)
Lowest latency
0.81 s
Databricks (Europe)
Highest throughput
59.08 tok/s
DeepInfra (United States)
Best availability
100%
Databricks (Europe)
Endpoint
First token
Latency p50
Latency p95
Throughput
Availability
Cache hits
Tool errors
JSON errors
DeepInfra (United States)
s
9.28 s
12.41 s
59.08 tok/s
100%
0.73%
0%
0%
Amazon Web Services (Europe)
s
3.1 s
12.53 s
57.31 tok/s
89.55%
1.52%
0%
0%
Amazon Web Services (United States)
s
s
s
tok/s
%
%
%
%
Databricks (Europe)
1.53 s
0.81 s
4.06 s
36.06 tok/s
100%
0.11%
1.15%
0%
Databricks
1.53 s
1.04 s
7.67 s
41.59 tok/s
100%
0.21%
1.15%
0%

Estimate your monthly cost

Enter your expected volume to compare every endpoint at once.

Call Gemma 3 12B IT with Eden AI

Use the model ID to let Eden AI pick an endpoint, or an endpoint ID from the table above to pin one.

import requests

response = requests.post(
    "https://api.edenai.run/v3/chat/completions",
    headers={"Authorization": "Bearer YOUR_EDENAI_API_KEY"},
    json={
        "model": "",
        "messages": [{"role": "user", "content": "Hello!"}],
    },
)
print(response.json())

Questions about Gemma 3 12B IT

On Eden AI, prices for this model start at $0.05 per million input tokens and $0.15 per million output tokens. Prices differ by endpoint; the providers table lists each one. .

Up to 131K tokens, depending on the endpoint.

5 endpoints from 3 providers, available in: Global, Europe, United States. All of them are called with the same Eden AI API key.

Yes. At least one endpoint is hosted in Europe; pin it with its endpoint ID from the providers table to keep requests in the EU.

Over the last 30 days, the fastest first token was 1.53 s (Databricks (Europe)) and the highest throughput was 59.08 tokens per second (DeepInfra (United States)).