Gemma 4 31B IT

gemma-4-31b-it by Google Cloud Released April 3, 2026
ReasoningFunction callingStructured outputPrompt cachingWeb searchComputer useImage inputFile inputAudio inputVideo inputEU hostingFree endpointUp to % off
Input from
$0 / 1M
Output from
$0 / 1M
Cached input from
$0.015 / 1M
Context window
262K
Endpoints
4
from 4 providers
Regions
Global, Europe, United States

About Gemma 4 31B IT

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output.

Gemma 4 31B IT is a chat model by Google Cloud, available on Eden AI through 4 endpoints from 4 providers, including a free endpoint (paid endpoints from $0.1 per million input tokens), with a context window of up to 262K tokens.

Providers and prices

4 endpoints from 4 providers, cheapest first, in USD per million tokens. Every one is called with the same Eden AI key: pin one with its endpoint ID, or let Eden AI route for you. .

Endpoint
Region
Context
Input / 1M
Output / 1M
Cached / 1M
Supports
Google CloudOfficial−%google/gemma-4-31b-it
Global
262K
$0$0
$0
$
ReasoningToolsJSONCacheVisionWeb
FlexAIOfficial−%flexai/gemma-4-31b-it
Global
262K
$0.1$0.1
$0.34
$0.015
ReasoningToolsJSONCacheVisionWeb
Greenference (Europe)Official−%greenference/gemma-4-31b-it
Europe
262K
$0.1485$0.1485
$0.42
$0.0578
ReasoningToolsJSONCacheVisionWeb
DeepInfra (United States)Official−%deepinfra/google/gemma-4-31B-it
United States
262K
$0.2$0.2
$0.4
$0.05
ReasoningToolsJSONCacheVisionWeb

Performance

Measured on real Eden AI traffic over the last 30 days. Endpoints with too little traffic are not shown.

Fastest first token
1.19 s
Google Cloud
Lowest latency
2.48 s
Greenference (Europe)
Highest throughput
46.29 tok/s
FlexAI
Best availability
99.99%
DeepInfra (United States)
Endpoint
First token
Latency p50
Latency p95
Throughput
Availability
Cache hits
Tool errors
JSON errors
Google Cloud
1.19 s
41.68 s
115.04 s
35.17 tok/s
99.29%
10.88%
72.65%
73.14%
FlexAI
1.92 s
3.58 s
20.8 s
46.29 tok/s
99.62%
27.7%
0.2%
0.71%
Greenference (Europe)
1.99 s
2.48 s
66.1 s
18.35 tok/s
99.84%
43.62%
0.57%
0%
DeepInfra (United States)
1.65 s
24.3 s
64.27 s
16.59 tok/s
99.99%
0.01%
6.49%
0%

Estimate your monthly cost

Enter your expected volume to compare every endpoint at once.

Call Gemma 4 31B IT with Eden AI

Use the model ID to let Eden AI pick an endpoint, or an endpoint ID from the table above to pin one.

import requests

response = requests.post(
    "https://api.edenai.run/v3/chat/completions",
    headers={"Authorization": "Bearer YOUR_EDENAI_API_KEY"},
    json={
        "model": "",
        "messages": [{"role": "user", "content": "Hello!"}],
    },
)
print(response.json())

Questions about Gemma 4 31B IT

On Eden AI, prices for this model start at $0 per million input tokens and $0 per million output tokens. Prices differ by endpoint; the providers table lists each one. .

Up to 262K tokens, depending on the endpoint.

4 endpoints from 4 providers, available in: Global, Europe, United States. All of them are called with the same Eden AI API key.

Yes. At least one endpoint is hosted in Europe; pin it with its endpoint ID from the providers table to keep requests in the EU.

Over the last 30 days, the fastest first token was 1.19 s (Google Cloud) and the highest throughput was 46.29 tokens per second (FlexAI).