GLM-5.1

glm-5.1 by Z.ai Released April 8, 2026
ReasoningFunction callingStructured outputPrompt cachingWeb searchComputer useImage inputFile inputAudio inputVideo inputEU hostingFree endpointUp to 35% off
Input from
$0.91 / 1M
Output from
$2.86 / 1M
Cached input from
$0.169 / 1M
Context window
203K
Endpoints
5
from 5 providers
Regions
Global, Asia, Europe, United States

About GLM-5.1

GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks.

GLM-5.1 is a chat model by Z.ai, available on Eden AI through 5 endpoints from 5 providers, from $0.91 per million input tokens and $2.86 per million output tokens, with a context window of up to 203K tokens.

Providers and prices

5 endpoints from 5 providers, cheapest first, in USD per million tokens. Every one is called with the same Eden AI key: pin one with its endpoint ID, or let Eden AI route for you. .

Endpoint
Region
Context
Input / 1M
Output / 1M
Cached / 1M
Supports
QwenOfficial−35%qwen/glm-5.1
Global
203K
$1.4$0.91
$2.86
$0.169
ReasoningToolsJSONCacheVisionWeb
DeepInfra (United States)Official−%deepinfra/zai-org/GLM-5.1
United States
203K
$1.05$1.05
$3.5
$0.205
ReasoningToolsJSONCacheVisionWeb
TensorX (Europe)Official−%tensorx/z-ai/glm-5.1
Europe
203K
$1.4$1.4
$4.4
$0.35
ReasoningToolsJSONCacheVisionWeb
NebiusOfficial−%nebius/zai-org/GLM-5.1
Global
203K
$1.4$1.4
$4.4
$1.4
ReasoningToolsJSONCacheVisionWeb
Z.ai (Asia)Official−%zai/glm-5.1
Asia
203K
$1.4$1.4
$4.4
$0.26
ReasoningToolsJSONCacheVisionWeb

Performance

Measured on real Eden AI traffic over the last 30 days. Endpoints with too little traffic are not shown.

Fastest first token
1.29 s
TensorX (Europe)
Lowest latency
2.28 s
TensorX (Europe)
Highest throughput
65.99 tok/s
TensorX (Europe)
Best availability
100%
TensorX (Europe)
Endpoint
First token
Latency p50
Latency p95
Throughput
Availability
Cache hits
Tool errors
JSON errors
Qwen
2.66 s
4.92 s
30.1 s
37.27 tok/s
99.43%
86.39%
0%
0%
DeepInfra (United States)
2.27 s
5.38 s
72.59 s
27.14 tok/s
100%
90.87%
0%
0%
TensorX (Europe)
1.29 s
2.28 s
67.14 s
65.99 tok/s
100%
96.86%
0%
0%
Nebius
2.57 s
96.31 s
244.45 s
34.52 tok/s
100%
14.44%
0%
0%
Z.ai (Asia)
9.03 s
60.44 s
138.72 s
21.09 tok/s
100%
0.05%
0%
0%

Estimate your monthly cost

Enter your expected volume to compare every endpoint at once.

Call GLM-5.1 with Eden AI

Use the model ID to let Eden AI pick an endpoint, or an endpoint ID from the table above to pin one.

import requests

response = requests.post(
    "https://api.edenai.run/v3/chat/completions",
    headers={"Authorization": "Bearer YOUR_EDENAI_API_KEY"},
    json={
        "model": "",
        "messages": [{"role": "user", "content": "Hello!"}],
    },
)
print(response.json())

Questions about GLM-5.1

On Eden AI, prices for this model start at $0.91 per million input tokens and $2.86 per million output tokens. Prices differ by endpoint; the providers table lists each one. .

Up to 203K tokens, depending on the endpoint.

5 endpoints from 5 providers, available in: Global, Asia, Europe, United States. All of them are called with the same Eden AI API key.

Yes. At least one endpoint is hosted in Europe; pin it with its endpoint ID from the providers table to keep requests in the EU.

Over the last 30 days, the fastest first token was 1.29 s (TensorX (Europe)) and the highest throughput was 65.99 tokens per second (TensorX (Europe)).