Grok Voice Think Fast 2.0 has been released, supporting both transcription and voice interaction.



SpaceXAI has announced its voice AI model, ' Grok Voice Think Fast 2.0 .' This model can interact with users via voice and significantly improves upon its predecessor in terms of voice inference capabilities, conversational ability, and reliability when using tools.

Introducing Grok Voice Think Fast 2.0 | SpaceXAI

https://x.ai/news/grok-voice-think-fast-2

The following table compares Grok Voice Think Fast 2.0, its predecessor Grok Voice Think Fast 1.0, OpenAI's real-time voice dialogue AI ' GPT Realtime 2.1 (High)', and Google's high-speed inference model 'Gemini 3.1 Flash (High)' using Artificial Analysis benchmarks. The evaluation items from top to bottom are 'Overall Voice-Related Quality Indicators', 'Voice Inference', 'Conversation Dynamics', 'Agent Performance', and 'Speed to First Voice'. Grok Voice Think Fast 2.0 achieved the highest score or fastest results among the comparison in four of the items, excluding Conversation Dynamics.




According to Artificial Analysis, Grok Voice Think Fast 2.0 ranks second in the Speech-to-Speech Quality Index, behind ' Qwen Audio 3.0 TTS Plus, ' and is ranked first in agent performance, while also boasting one of the fastest start times for voice input at 0.70 seconds.




Grok Voice Think Fast 2.0 excels not only in voice interaction but also as a transcription model. In evaluations using thousands of short phrases in 24 different languages, Grok Voice Think Fast 2.0 recorded a 1.5 to 2.0 times improvement compared to ' Deepgram Nova 3 ,' which specializes in speech recognition and speech synthesis, and ' Scribe v2 ,' a speech recognition model that ElevenLabs touts as having 'the most accurate real-time transcription with a low latency of 150ms.' Compared to its predecessor, Grok Voice Think Fast 1.0, it achieved a 1.4 times improvement. Furthermore, according to SpaceXAI, Grok Voice Think Fast 2.0 was developed with a focus on performing exceptionally well in real-world environments such as background noise and degraded audio over the phone, and in noisy environments, the difference compared to other speech recognition models widens to approximately 10 times.

The following graphs show the word error rate for each language, with lower numbers indicating higher transcription accuracy. Grok Voice Think Fast 2.0 recorded low word error rates for all languages.



A key feature of the Grok Voice Think Fast model is that it infers queries while speaking, making it far more intelligent than other speech dialogue models while minimizing latency. Furthermore, Grok Voice Think Fast 2.0 has been trained to be significantly more efficient at processing inference tokens compared to the previous model, reducing the number of inference tokens by approximately 60%. As a result, tool calls are much faster in production environments and are usually executed before the agent finishes its first sentence.

If you are using the 'grok-voice-latest' option, your device will automatically switch to Grok Voice Think Fast 2.0 starting August 5, 2026. You do not need to do anything to upgrade, but if you wish to continue using the previous model, you will need to specify 'grok-voice-think-fast-1.0' to fix the version. Grok Voice Think Fast 2.0 costs $0.08 (approximately 13 yen) per minute of voice.

in AI, Posted by log1e_dh