GLM-5

glm-5 by Z.ai Released February 12, 2026
ReasoningFunction callingStructured outputPrompt cachingWeb searchComputer useImage inputFile inputAudio inputVideo inputEU hostingFree endpointUp to % off
Input from
$0.6 / 1M
Output from
$2.08 / 1M
Cached input from
$0.12 / 1M
Context window
1.05M
Endpoints
6
from 5 providers
Regions
Asia, Europe, United States

About GLM-5

GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows.

GLM-5 is a chat model by Z.ai, available on Eden AI through 6 endpoints from 5 providers, from $0.6 per million input tokens and $2.08 per million output tokens, with a context window of up to 1.05M tokens.

Providers and prices

6 endpoints from 5 providers, cheapest first, in USD per million tokens. Every one is called with the same Eden AI key: pin one with its endpoint ID, or let Eden AI route for you. .

Endpoint
Region
Context
Input / 1M
Output / 1M
Cached / 1M
Supports
DeepInfra (United States)Official−%deepinfra/zai-org/GLM-5
United States
203K
$0.6$0.6
$2.08
$0.12
ReasoningToolsJSONCacheVisionWeb
TensorX (Europe)Official−%tensorx/z-ai/glm-5
Europe
203K
$1$1
$3.2
$0.25
ReasoningToolsJSONCacheVisionWeb
Amazon Web Services (Europe)Official−%amazon/zai.glm-5
Europe
200K
$1$1
$3.2
$
ReasoningToolsJSONCacheVisionWeb
Amazon Web Services (United States)Official−%amazon/zai.glm-5@us
United States
200K
$1$1
$3.2
$
ReasoningToolsJSONCacheVisionWeb
Z.ai (Asia)Official−%zai/glm-5
Asia
203K
$1$1
$3.2
$0.2
ReasoningToolsJSONCacheVisionWeb
Mistral AI (Europe)Official−%mistral/zai-glm-5
Europe
1.05M
$1.4$1.4
$4.4
$0.14
ReasoningToolsJSONCacheVisionWeb

Performance

Measured on real Eden AI traffic over the last 30 days. Endpoints with too little traffic are not shown.

Fastest first token
0.48 s
Mistral AI (Europe)
Lowest latency
4.48 s
Amazon Web Services (Europe)
Highest throughput
185.68 tok/s
TensorX (Europe)
Best availability
100%
Mistral AI (Europe)
Endpoint
First token
Latency p50
Latency p95
Throughput
Availability
Cache hits
Tool errors
JSON errors
DeepInfra (United States)
2.72 s
9.71 s
124.29 s
46.62 tok/s
100%
45%
0%
0%
TensorX (Europe)
2.32 s
23.26 s
104.57 s
185.68 tok/s
100%
10.27%
0%
0%
Amazon Web Services (Europe)
0.89 s
4.48 s
35.21 s
75.88 tok/s
89.36%
0%
100%
0%
Amazon Web Services (United States)
s
s
s
tok/s
%
%
%
%
Z.ai (Asia)
8.66 s
43.87 s
128.25 s
35.12 tok/s
100%
38.95%
0%
0%
Mistral AI (Europe)
0.48 s
17.76 s
42.61 s
181 tok/s
100%
20.89%
0%
0%

Estimate your monthly cost

Enter your expected volume to compare every endpoint at once.

Call GLM-5 with Eden AI

Use the model ID to let Eden AI pick an endpoint, or an endpoint ID from the table above to pin one.

import requests

response = requests.post(
    "https://api.edenai.run/v3/chat/completions",
    headers={"Authorization": "Bearer YOUR_EDENAI_API_KEY"},
    json={
        "model": "",
        "messages": [{"role": "user", "content": "Hello!"}],
    },
)
print(response.json())

Questions about GLM-5

On Eden AI, prices for this model start at $0.6 per million input tokens and $2.08 per million output tokens. Prices differ by endpoint; the providers table lists each one. .

Up to 1.05M tokens, depending on the endpoint.

6 endpoints from 5 providers, available in: Asia, Europe, United States. All of them are called with the same Eden AI API key.

Yes. At least one endpoint is hosted in Europe; pin it with its endpoint ID from the providers table to keep requests in the EU.

Over the last 30 days, the fastest first token was 0.48 s (Mistral AI (Europe)) and the highest throughput was 185.68 tokens per second (TensorX (Europe)).