GLM-4.7 Flash

glm-4.7-flash by Z.ai Released January 20, 2026
ReasoningFunction callingStructured outputPrompt cachingWeb searchComputer useImage inputFile inputAudio inputVideo inputEU hostingFree endpointUp to % off
Input from
$0.06 / 1M
Output from
$0.4 / 1M
Cached input from
$0.01 / 1M
Context window
203K
Endpoints
5
from 4 providers
Regions
Global, Europe, United States

About GLM-4.7 Flash

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency.

GLM-4.7 Flash is a chat model by Z.ai, available on Eden AI through 5 endpoints from 4 providers, from $0.06 per million input tokens and $0.4 per million output tokens, with a context window of up to 203K tokens.

Providers and prices

5 endpoints from 4 providers, cheapest first, in USD per million tokens. Every one is called with the same Eden AI key: pin one with its endpoint ID, or let Eden AI route for you. .

Endpoint
Region
Context
Input / 1M
Output / 1M
Cached / 1M
Supports
DeepInfra (United States)Official−%deepinfra/zai-org/GLM-4.7-Flash
United States
203K
$0.06$0.06
$0.4
$0.01
ReasoningToolsJSONCacheVisionWeb
Cloudflare (CF/ZAI-ORG/GLM-4.7-FLASH)Official−%cloudflare/@cf/zai-org/glm-4.7-flash
CF/ZAI-ORG/GLM-4.7-FLASH
131K
$0.0605$0.0605
$0.4
$
ReasoningToolsJSONCacheVisionWeb
Greenference (Europe)Official−%greenference/glm-4.7-flash
Europe
131K
$0.07$0.07
$0.4076
$0.0143
ReasoningToolsJSONCacheVisionWeb
Amazon Web Services (Europe)Official−%amazon/zai.glm-4.7-flash
Europe
203K
$0.07$0.07
$0.4
$
ReasoningToolsJSONCacheVisionWeb
Amazon Web Services (United States)Official−%amazon/zai.glm-4.7-flash@us
United States
203K
$0.07$0.07
$0.4
$
ReasoningToolsJSONCacheVisionWeb

Performance

Measured on real Eden AI traffic over the last 30 days. Endpoints with too little traffic are not shown.

Fastest first token
0.46 s
Amazon Web Services (Europe)
Lowest latency
0.41 s
Amazon Web Services (Europe)
Highest throughput
127.66 tok/s
Greenference (Europe)
Best availability
99.95%
DeepInfra (United States)
Endpoint
First token
Latency p50
Latency p95
Throughput
Availability
Cache hits
Tool errors
JSON errors
DeepInfra (United States)
1.42 s
5.58 s
21.54 s
24.07 tok/s
99.95%
13.99%
11.11%
0.05%
Cloudflare (CF/ZAI-ORG/GLM-4.7-FLASH)
s
s
s
tok/s
%
%
%
%
Greenference (Europe)
0.92 s
2 s
66.91 s
127.66 tok/s
95.1%
72.14%
1.8%
55.34%
Amazon Web Services (Europe)
0.46 s
0.41 s
1.37 s
57.75 tok/s
99%
0%
0%
0.42%
Amazon Web Services (United States)
s
s
s
tok/s
%
%
%
%

Estimate your monthly cost

Enter your expected volume to compare every endpoint at once.

Call GLM-4.7 Flash with Eden AI

Use the model ID to let Eden AI pick an endpoint, or an endpoint ID from the table above to pin one.

import requests

response = requests.post(
    "https://api.edenai.run/v3/chat/completions",
    headers={"Authorization": "Bearer YOUR_EDENAI_API_KEY"},
    json={
        "model": "",
        "messages": [{"role": "user", "content": "Hello!"}],
    },
)
print(response.json())

Questions about GLM-4.7 Flash

On Eden AI, prices for this model start at $0.06 per million input tokens and $0.4 per million output tokens. Prices differ by endpoint; the providers table lists each one. .

Up to 203K tokens, depending on the endpoint.

5 endpoints from 4 providers, available in: Global, Europe, United States. All of them are called with the same Eden AI API key.

Yes. At least one endpoint is hosted in Europe; pin it with its endpoint ID from the providers table to keep requests in the EU.

Over the last 30 days, the fastest first token was 0.46 s (Amazon Web Services (Europe)) and the highest throughput was 127.66 tokens per second (Greenference (Europe)).