SpaceXAI Releases Grok Voice Transcribe 2.0: Speech-to-Text API

SpaceXAI Releases Grok Voice Transcribe 2.0: Speech-to-Text API Claims 2x Accuracy Over 1.0 at $0.10 per Hour


SpaceXAI has released Grok Voice Transcribe 2.0, its newest speech-to-text (STT) model. The development team claims it is twice as accurate as Grok Voice Transcribe 1.0 at the same price. The model targets difficult audio: noisy phone lines, competing voices, local accents, and spoken credentials. It runs in batch and real-time streaming modes through the Speech to Text API.


Is it deployable? Yes, as a hosted API. It is live today under the model ID grok-voice-transcribe-2.0. SpaceXAI has not announced open weights, so self-hosting is not an option.


What Is Grok Voice Transcribe 2.0?


Grok Voice Transcribe 2.0 is built on the audio foundation model behind Grok Voice. The SpaceXAI team states that Grok Voice already handles tens of thousands of customer-support calls per day. It also transcribes millions of hours of video narration and runs the Grok assistant in Tesla vehicles.


The training data consists of live, noisy, multilingual audio recorded across diverse environments. SpaceXAI then refined the model with post-training.


Benchmarks: What SpaceXAI Reports


SpaceXAI reports a first-place accuracy rank among 32 streaming models on the public Artificial Analysis leaderboard. That benchmark, AA-WER Streaming, uses about 8 hours of audio. It weights AA-AgentTalk at 50%, VoxPopuli at 25%, and Earnings22 at 25%. See the methodology for details.


SpaceXAI also measures word error rate (WER) on four internal sets drawn from production traffic:


  • Telephony (8 kHz): English customer support calls
  • Conversational: English conversations with Grok
  • Credentials: phone numbers, emails, and add...

via MarkTechPost

Related