Voice agents often stumble on the details that matter most in a call—the order number, the callback digits, or the email address the caller must jot down. Gradium AI has introduced a new text-to-speech (TTS) model, now set as the default across its API and Studio platform. The company reports an 81.0% human-rated pass rate on a 500-sentence hard-case evaluation set spanning five languages, outperforming Cartesia Sonic 3.6 (75.1%) and ElevenLabs v3 Conversational (65.4%). Additionally, the model achieves a time-to-first-audio of 216 milliseconds at P50 on Coval, which is 170 ms faster than its predecessor.
Deployment and Compatibility
Yes, it's deployable today with no migration required. Gradium activated the model as the default on August 31, 2026, across both its API and Studio. Existing voices, including custom clones, continue to function without any changes.
Benchmarking Accuracy
Gradium constructed a 500-sentence evaluation dataset, which it open-sourced on Hugging Face under the CC BY 4.0 license. The set includes 100 items across 10 criteria in five languages (EN, DE, FR, ES, PT). Seven atomic criteria focus on spelling, acronyms, alphanumeric tokens, dates, regular numbers, large and floating numbers, and email addresses. Three composite criteria—Orders, IT Ticket, and Claims—stack several of these elements into realistic, single-turn agent interactions.
Scoring is human-based and strict. A sentence passes only if an independent native-speaker rater hears every element pronounced correctly and completely; a single dropped digit fails the entire sentence. To ensure fairness, audio was loudness-normalized, the order of presentations was randomized, and raters were limited to 40 comparisons per session with enforced breaks.
Pooling results across all ten criteria and averaging equally over the five languages, the scores are: Gradium TTS at 81.0%, Cartesia Sonic 3.6 at 75.1%, and ElevenLabs v3 Conversational at 65.4%.
via MarkTechPost
