Articles

Grok Voice Think Fast 2.0: Pricing, Specs, Default Date

xAI's Grok Voice Think Fast 2.0 becomes the default grok-voice-latest on Aug 5, with an 82.9% speech-quality score and $0.08/min pricing. What changed.

Chisato Chisato · · 6 min read
A glowing blue waveform of particles rippling across a dark field

The voice-AI race gets a new default this week. On August 5, 2026, xAI flips its grok-voice-latest alias from Grok Voice Think Fast 1.0 to Grok Voice Think Fast 2.0, the company’s most capable speech-to-speech model. xAI (which brands the effort SpaceXAI) announced the model on July 29 and made it available immediately through its Voice API; tomorrow’s switch means every developer pointed at the -latest alias upgrades automatically, with no action required. Teams that want to stay on the older model have to pin grok-voice-think-fast-1.0 before the cutover.

The upgrade is a bet that the next competitive surface in AI is not just how smart a model is, but how naturally it can listen and talk — the same wager OpenAI made with its full-duplex GPT-Live models a few weeks earlier.

What changed in 2.0

xAI describes Think Fast 2.0 as a next-generation voice model with gains across three axes: speech reasoning, conversational ability, and tool-use reliability. In practice, that means a model meant to hold a fluid back-and-forth, follow instructions mid-conversation, and call external tools without the awkward pauses and dropped turns that plague earlier voice assistants.

Two efficiency claims stand out. xAI says the new model uses roughly 60% fewer reasoning tokens than version 1.0 to reach its answers, and that its tool-call timing is snappier — both of which matter for latency, the quiet killer of voice interfaces. A half-second gap between a user finishing a sentence and the assistant beginning its reply is enough to make a conversation feel mechanical, so cutting the internal work required per turn is as much a product decision as a cost one.

On accuracy, xAI makes an unusually aggressive claim: that Think Fast 2.0 outperforms dedicated, state-of-the-art transcription models on its own evaluation. The company reports transcription-error improvements of 1.5 to 2 times over Deepgram Nova 3 and ElevenLabs Scribe v2 across 24 languages, with the gap widening to roughly tenfold in noisy environments — precisely the conditions (bad connections, background noise, crosstalk) where prior voice modes tend to break.

The benchmark number

xAI anchored the release to the AA Speech-to-Speech Quality Index, where Think Fast 2.0 scored 82.9%. The company says that places it ahead of its own predecessor as well as two of the most visible rivals in the category — GPT-Realtime-2.1 and Gemini 3.1 Flash.

As always, a vendor’s headline benchmark is chosen on the vendor’s ground, and independent testing will refine the picture. Speech-to-speech quality is also notoriously hard to reduce to a single figure: latency, interruption handling, prosody, and accuracy in the wild all shape whether a voice model feels natural, and none of them is fully captured by one index. Still, the direction is clear enough. xAI is positioning Grok Voice not as a niche add-on but as a front-runner in a category every major lab is now contesting.

Pricing and the cost trade-off

The upgrade is not free. xAI priced Think Fast 2.0 at $0.08 per minute of speech-to-speech audio — up from $0.05 per minute for version 1.0 — plus $0.004 per text input. That is a 60% increase in the per-minute rate, and it puts a real number on the value xAI believes the new model delivers.

The pricing choice is notable because it runs against xAI’s broader playbook. The company has spent 2026 competing hardest on cost, shipping Grok 4.3 as the cheapest US frontier reasoning model and leaning on price in a market where raw capability is converging and frontier pricing keeps compressing. Raising the voice rate signals that xAI sees speech-to-speech as a premium capability worth charging for — a place where quality, not just price, can win the deal.

For developers, the automatic -latest switch makes the math immediate. Anyone building on the alias for cost predictability should note that their per-minute voice bill rises tomorrow unless they pin the older model, and should weigh the 2.0 accuracy and latency gains against that increase for their specific workload. Voice products with thin margins or high call volumes will feel the difference; those where transcription accuracy in noisy conditions is the make-or-break factor may find the upgrade pays for itself.

Why full-duplex voice is the new battleground

Grok Voice arrives into a market where every major lab is racing on the same two fronts: making models feel natural to interact with, and making them capable of taking action. Voice sits at the intersection. It is the interface most likely to pull AI assistants off the screen and into ambient, hands-free use — in cars, on phones, through earbuds, and inside customer-facing systems that field millions of calls.

The technical challenge is a genuine trade-off. A model fast enough for real-time speech struggles to also reason deeply or use tools well; a model smart enough to reason is usually too slow to hold a natural conversation. Different labs are attacking that tension differently. OpenAI’s GPT-Live splits the problem, running a nimble conversational layer that delegates hard questions to a heavier frontier model behind the scenes. xAI’s “Think Fast” branding points at the opposite instinct — squeezing enough reasoning into a single low-latency model that it can both converse and think without handing off, which is what the 60%-fewer-tokens claim is really about.

Tool-use reliability is the other half of the story, and the reason voice matters for large language model providers beyond consumer chat. A voice assistant that can reliably trigger the right function — book the appointment, pull the record, execute the workflow — is an agent with a spoken interface. Improvements in tool timing and reliability are what move a voice model from a demo that sounds impressive to a system a business will put in front of customers.

Where xAI fits

xAI’s approach across its lineup has been consistent: ship capable models quickly, distribute them where developers already work, and compete on a mix of price and speed of iteration. Grok Voice fits that pattern, and it extends the company’s push beyond text and its Opus-class Grok 4.5 release into the real-time, multimodal surfaces where usage is growing fastest.

The competitive picture is crowded. OpenAI, Google, and a wave of specialized voice startups are all shipping speech models on a compressed calendar, and dedicated transcription vendors like Deepgram and ElevenLabs — the ones xAI is now benchmarking against — have built businesses on accuracy alone. By claiming to beat purpose-built transcription models with a general speech-to-speech system, xAI is arguing that the specialized-vendor moat is narrowing, the same way general frontier models have absorbed capability after capability that once required a dedicated tool.

What it means

Grok Voice Think Fast 2.0 is a bet that the returns in AI are shifting toward how a model feels to use, and that voice is where that competition plays out next.

Who wins. Developers building voice products get a more accurate, lower-latency model with better tool use — and the automatic -latest upgrade means they inherit those gains without lifting a finger. Businesses that operate in noisy, real-world conditions stand to benefit most from the transcription claims, if independent testing bears them out.

Who feels the pressure. Dedicated transcription vendors are the direct target of xAI’s benchmarks; a general speech-to-speech model that matches or beats them on accuracy erodes the case for a separate STT layer. Rival voice models from OpenAI and Google now have a concrete quality number to answer. And developers watching their bills feel a subtler pressure: the per-minute price went up, and the default switch makes it the path of least resistance.

What to watch next. Three things. First, independent evaluations — the 82.9% index score and the transcription claims are xAI’s own, and the category is hard to measure, so third-party testing across accents, languages, and noisy conditions is the real verdict. Second, real-world latency and interruption handling, which determine whether a voice model holds up outside curated demos. Third, whether the price increase sticks as the voice market compresses the way text pricing has — xAI has chosen to charge more for quality here, and the market will decide whether speech-to-speech is the premium capability the company is betting it is.

Chisato Chisato · · 4 min read

The ReAct Pattern: How AI Agents Reason and Act

ReAct interleaves an LLM's reasoning with tool calls and their results, letting an agent adjust its plan after each observation instead of reasoning blind.

#AI #Agents #LLMs
Chisato Chisato · · 4 min read

What Is an LLM Router?

An LLM router sends each request to the cheapest or fastest model that can handle it, instead of routing every call to one model regardless of difficulty.

#AI #LLMs #Agents