OpenAI GPT-Live: Full-Duplex Voice, Features, Access
OpenAI launched GPT-Live and GPT-Live-1 mini, full-duplex voice models that listen and speak at once and delegate hard questions to a frontier model. What's new.
OpenAI has replaced the halting, turn-based experience of talking to ChatGPT with something closer to a phone call. On July 8, 2026, the company introduced GPT-Live, a new family of full-duplex voice models that can listen and speak at the same time, and began rolling it out globally across iOS, Android, and ChatGPT.com the same day. A day later, on July 9, OpenAI followed with ChatGPT Work, an agent-based product that delegates real tasks across a user’s applications. Together, the two launches landed in the same week that OpenAI opened its GPT-5.6 model family to the public, capping one of the busiest release stretches the company has run.
What GPT-Live is
GPT-Live ships in two variants: GPT-Live-1, which becomes the default voice model for paid ChatGPT users, and GPT-Live-1 mini, which serves the free tier. Both replace Advanced Voice Mode, the previous voice experience, rather than sitting alongside it.
The defining feature is a full-duplex architecture. Where earlier voice assistants worked in strict turns — you speak, it processes, it replies — GPT-Live can process incoming audio and generate speech simultaneously. In practice that means the model can interject with a quick “mhmm” or “yeah” to signal it is following along, jump into a fast back-and-forth, or simply stay silent when a user pauses to think. OpenAI’s framing is that the model behaves less like a command line for your voice and more like a conversation partner that understands the rhythm of human speech.
The company describes the design as a deliberate split between conversation and cognition. GPT-Live handles the real-time listening and speaking, and for anything that needs web search, deeper reasoning, or heavier work, it delegates to OpenAI’s latest frontier model behind the scenes — reported as GPT-5.5 — and folds the result back into the conversation when it is ready. That routing is what lets the voice layer stay fast and low-latency while still reaching for a large reasoning model when a question demands it.
Why full-duplex matters
Latency has been the quiet killer of voice interfaces. A half-second gap between a user finishing a sentence and an assistant beginning its reply is enough to make a conversation feel mechanical, and the workaround — waiting for a clear end-of-turn signal before responding — makes interruptions clumsy. Full-duplex removes that constraint by treating listening and speaking as concurrent streams rather than a strict handoff.
The behavioral details OpenAI highlighted — backchannel cues like “mhmm,” the ability to be interrupted mid-sentence and recover gracefully, and knowing when to stay quiet — are the difference between a demo and something people will actually use hands-free. These are the moments where prior voice modes broke character, talking over users or freezing when interrupted. By building the turn-taking behavior into the model rather than bolting it on with voice-activity detection, OpenAI is betting it can make spoken interaction feel native instead of adapted.
The delegation model is the other half of the story. A single monolithic model fast enough for real-time speech would struggle to also do multi-step reasoning or search; a model smart enough to reason deeply would be too slow to hold a natural conversation. Splitting the two — a nimble conversational layer that calls a heavier reasoning model only when needed — is an architectural answer to that tension, and it mirrors how agentic systems increasingly route work to the right tool for each subtask.
ChatGPT Work: the agent layer
The day after GPT-Live, OpenAI shipped ChatGPT Work, positioned as an agent that “delegates real work across your apps.” It is powered by Codex — OpenAI’s coding and task-execution engine — and by the newly public GPT-5.6, and it is aimed squarely at the workplace: connecting to the applications a knowledge worker uses and carrying out multi-step tasks on their behalf rather than just answering questions.
ChatGPT Work slots into a broader industry push toward agents that act rather than advise. If you want to understand the mechanics beneath products like this, our guide to building your own AI agent walks through the loop of planning, tool calls, and execution that underpins them, and the emergence of shared standards like the Model Context Protocol is what lets an assistant reach into a user’s calendar, documents, and internal tools without a bespoke integration for each one.
The two launches are complementary. GPT-Live changes how you talk to ChatGPT; ChatGPT Work changes what ChatGPT can do once you have asked. A natural, low-latency voice interface is far more useful when the thing on the other end can actually execute a task across your apps rather than just describe how you might do it yourself.
Availability and the API question
GPT-Live began rolling out to ChatGPT users globally on July 8, across iOS, Android, and the web, with GPT-Live-1 as the paid default and GPT-Live-1 mini on the free tier. That is the part developers should read carefully: at launch, GPT-Live is a ChatGPT feature, not a developer-facing API.
OpenAI’s commitment to developers is a single sentence — it plans to bring the models to the API “soon,” and enterprises can sign up to be notified. As of launch there were no published endpoints, model IDs, per-minute audio pricing, or rate limits for GPT-Live. For teams that need programmatic voice today, OpenAI pointed to its existing Realtime API as the available option. Anyone considering building a production voice product on GPT-Live should wait for concrete pricing and latency numbers to appear on the documentation rather than a launch post.
The competitive picture
GPT-Live arrives into a market where every major lab is racing on the same two fronts: making models feel more natural to interact with, and making them capable of taking action. The voice push follows a wave of frontier model releases — OpenAI’s own GPT-5.6, Anthropic’s Sonnet 5, and SpaceXAI’s Grok 4.5 — that has compressed the industry’s release calendar into a matter of weeks. Voice is the next surface where that competition plays out, because it is the interface most likely to pull AI assistants off the screen and into ambient, hands-free use.
There is also a strategic logic to shipping voice and an agent product back-to-back. The companies that own the primary interface to AI — the place users go by default — accrue enormous advantages in data, habit, and distribution. A conversational voice layer that feels genuinely natural, wired to an agent that can complete tasks, is OpenAI’s pitch to make ChatGPT that default surface before a rival does.
What it means
GPT-Live is a bet that the next competitive battleground is interface, not just raw capability. The frontier labs have spent two years leapfrogging each other on benchmarks; GPT-Live signals that the returns are increasingly in how the model feels to use, and full-duplex voice is the most tangible step yet toward assistants that fit into the flow of a conversation instead of interrupting it.
The delegation architecture is the part worth watching. Splitting a fast conversational model from a slower frontier reasoner — and routing between them automatically — is a pragmatic pattern that other labs are likely to copy. It sidesteps the impossible trade-off of building one model that is both instant and deeply capable, and it hints at how future systems will be composed: not a single giant model, but a fabric of specialized ones handing work to each other.
For developers, the story is “wait and see.” The absence of a launch-day API, pricing, or rate limits means GPT-Live is not yet something to build on. The Realtime API remains the tool for voice today, and the smart move is to prototype the interaction while treating GPT-Live’s eventual API terms — especially per-minute audio cost — as the number that will decide whether a voice product is viable.
The winners and losers are starting to sort. Pairing GPT-Live with ChatGPT Work is a clear move to make ChatGPT the default AI surface for both consumers and the workplace. If it works, it deepens OpenAI’s distribution moat at the exact moment rivals are matching it on model quality. The thing to watch next is whether the voice experience holds up outside curated demos — in noisy rooms, over bad connections, across accents — because that, more than any benchmark, is what will determine whether people actually talk to their AI.
Keep reading
Chisato · · 6 min read OpenAI Astra Solves 10 Open Math Problems With Lean Proofs
OpenAI says an internal version of Astra, its next major model, solved ten long-open math problems — each shipped with a machine-checkable Lean proof.
Chisato · · 4 min read What Is Prompt Chaining? Multi-Step LLM Pipelines
Prompt chaining splits a task into a sequence of smaller LLM calls, each one feeding the next, instead of asking one giant prompt to do everything.
Chisato · · 6 min read OpenAI GPT-5.6-Cyber: What It Is and Who Gets Access
OpenAI launched GPT-5.6-Cyber and split its Daybreak security program into Blue and Red tiers. What the model does, its benchmarks, and who can use it.