Provider

Gradium

Gradium builds real-time voice models, with text to speech, speech to text and speech translation served from the EU region.

summary
  • Real-time voice is the whole point: Gradium competes on latency, and publishes a sub-50 millisecond figure for its text to speech model.
  • Three capabilities on Eden AI: text to speech, speech to text, and speech translation which transcribes and translates in one call.
  • EU region: all three models are listed in the EU, and Gradium documents region-pinned processing with no cross-region failover.
  • Two billing units: per character for speech at about $0.058 per 1,000, per second of audio for transcription at about $0.62 an hour.
  • Check the mode: text to speech is synchronous, transcription runs as an asynchronous job, so it fits recorded audio rather than a live stream.

What is Gradium?

Gradium is a Paris-based company building real-time voice models. Its focus is latency: it publishes a sub-50 millisecond figure for its text to speech model, which is the range where a spoken reply stops feeling like a wait and starts feeling like a conversation.

That number is Gradium's own published claim for the model. What you measure through any gateway, including Eden AI, includes a network hop on top of it, so treat it as the model's floor rather than an end-to-end promise.

Gradium at a glance

AttributeDetails
ProviderGradium
CategoryReal-time voice models
HeadquartersParis
Capabilities on Eden AIText to speech, speech to text, speech translation
Models on Eden AI3
Region on Eden AIEU
ExecutionText to speech is synchronous. Speech to text runs as an async job
Billing unitPer character for speech, per second of audio for transcription

Why latency is the whole argument in voice

Text generation can afford to be slow. You watch tokens appear, and a second of thinking reads as care. Voice cannot. A pause before a spoken answer reads as a broken product, and the tolerance is measured in a few hundred milliseconds for the entire round trip, not for one component of it.

That is why voice providers compete on a metric nobody discusses for text models. Once you are building an agent that speaks, the budget is spent before you notice: transcription, then the model, then synthesis, then the network. A speech component that is fast enough leaves room for the rest of the chain to exist.

When should you choose Gradium?

Choose Gradium if:

You are building something that speaks back and the round trip is your constraint rather than the quality of any single component.

Your data has to stay in the EU. Gradium models are served from the EU region on Eden AI, and Gradium documents region-pinned processing with no cross-region failover for organisations enrolled in residency.

You need transcription and translation together. Speech translation handles both in one call, which removes a stage and its latency from the chain.

Consider another provider if:

You need live streaming transcription. Speech to text runs as an asynchronous job through Eden AI, so it fits recorded audio rather than a live stream.

You need voice cloning or on-device models. Gradium builds them, but they are not exposed through Eden AI today.

You are generating long-form narration. Speech is billed per character, so an audiobook costs by the length of the script. For that shape of work, compare against per-second pricing before committing.

Gradium pros and cons

ProsCons
Built for real-time voice, where latency is the constraint that decides whether a product feels usableThree capabilities only, narrower than the general-purpose speech platforms
Served from the EU region on Eden AI, and Gradium documents region-pinned processingSpeech to text is asynchronous, so it is not a fit for live transcription through this route
Transcription and translation in one call, rather than transcribe then translatePer-character pricing on speech means long-form narration costs scale with the script, not the audio
A Paris company, so contracting and support sit in the same jurisdiction as the dataVoice cloning and on-device models are not exposed through Eden AI today

Gradium models and capabilities on Eden AI

Available Gradium models

CapabilityModel stringPrice on Eden AIMode
Text to Speechaudio/tts/gradium$0.058 per 1,000 charactersSync
Speech to Textaudio/speech_to_text_async/gradium$0.62 per hour of audioAsync job
Speech Translationaudio/speech_to_text_async/gradium/stt-translate$0.83 per hour of audioAsync job

Text to speech is billed per character of input, so you know the cost of a line before you synthesise it. Transcription and translation are billed per second of audio, which works out at roughly sixty two cents and eighty three cents an hour respectively.

Transcribe, or transcribe and translate

The two speech to text models share an endpoint and differ in what comes back. The base model returns the transcript in the spoken language. The translation model returns it in the target language, in the same call.

The second is not simply a convenience. Running transcription and translation as two steps means two round trips and two chances for the first output to degrade the second. One call removes both.

Where Gradium runs

All three models are listed in the EU region on Eden AI. Gradium's own documentation describes region-pinned processing for enrolled organisations: inference stays in the region, there is no cross-region failover, requests are refused rather than processed outside the region, and session data is written in the region. Its control plane, which holds account data, is documented as EU-hosted for all customers.

That is a stronger position than most speech providers can document. As with any provider, verify it against your own compliance requirement rather than taking a region label as proof.

What can you build with Gradium?

Voice agents that hold a conversation

The whole category depends on the round trip. A voice agent that answers in under a second feels like a product, and one that takes three seconds feels like a form with extra steps. The speech components are where most of that budget is either spent or saved.

Multilingual support without a translation stage

Speech translation turns a call in one language into text in another, in a single request. For support tooling, that is the difference between a pipeline and a feature.

Accessibility and transcription workflows

Per-second pricing on transcription makes the arithmetic simple. An hour of audio costs about sixty two cents, so a back catalogue of recordings has a cost you can work out before you start rather than discover afterwards.

Why use Gradium through Eden AI?

Speech sits beside the models that consume it. A voice agent needs transcription, a language model and synthesis, and on Eden AI all three are the same key, the same bill and the same request shape.

That also makes the components comparable. Swapping one speech provider for another to see which one your users actually prefer is a change of model string rather than a second integration, which is the only way anyone ever measures this honestly.

Frequently asked questions about Gradium on Eden AI

Gradium builds real-time voice models. Through Eden AI it provides text to speech, speech to text, and speech translation, which returns a transcript in a target language in a single call.

Gradium publishes a sub-50 millisecond latency figure for its text to speech model. That is the vendor’s published claim for the model itself. What you measure through a gateway includes a network hop on top of it, so treat the figure as the model’s floor rather than an end-to-end guarantee.

All three Gradium models are listed in the EU region. Gradium documents region-pinned processing for enrolled organisations, with no cross-region failover and requests refused rather than processed outside the region, and an EU-hosted control plane for all customers. Verify it against your own compliance requirement rather than treating a region label as proof.

Text to speech is billed per character, at about $0.058 per 1,000 characters. Speech to text and speech translation are billed per second of audio, which works out at roughly $0.62 and $0.83 per hour.

Not through Eden AI today. Speech to text runs as an asynchronous job, so it fits recorded audio. Text to speech, by contrast, is synchronous and returns in the request.

They are using Gradium

No items found.

Alternatives to Gradium

Deepgram is primarily about fast and accurate speech recognition, especially when audio volume, streaming or voice-product latency matter.

Speech

ElevenLabs should be evaluated through voice quality, speaker realism, latency and the type of audio experience the product needs.

Speech

Gladia should be compared on transcription speed, multilingual coverage and what happens after the transcript is produced.

Speech
let’s start

Start building with Eden AI

A single interface to integrate the best AI technologies into your products.