OpenAI unveils advanced voice models for more natural real-time conversations

OpenAI has launched new conversational models, GPT-Live-1 and GPT-Live-1 mini, designed to sound more natural and manage turn-taking more effectively. These full-duplex models can speak and listen simultaneously, allowing users to interrupt naturally and enabling features such as live translation.
The company is also rolling out GPT-Live-1 mini as the default replacement for the current Advanced Voice Mode in ChatGPT. Paid-tier users will have access to the larger GPT-Live-1 model. Previously, the system relied on a combination of a speech-to-text model, a large language model for generating responses, and a text-to-speech model to produce the final output.
During a press briefing, OpenAI stated that the new models address issues like interrupting users while they are speaking and lacking sufficient intelligence to answer questions. The updated models will route queries to the latest text models, such as GPT-5.5, for search, reasoning, or agentic tasks while maintaining the conversation flow.
OpenAI also demonstrated that the model can remain silent for extended periods, absorbing conversational context until it is addressed. Additionally, since the new voice mode integrates with newer GPT models, it can present information visually. Other startups, such as Monogram—which raised $40 million in seed funding from DST and Lux Capital—are also focusing on visual responses to make assistants more interactive.
The company noted that the new voice mode in ChatGPT is built for longer conversations. At the briefing, ChatGPT Voice’s product lead, Atty Eleti, mentioned holding 30- to 40-minute conversations with the voice feature during walks.
OpenAI believes voice could become the primary interface for complex computing tasks. Reports suggest the company may launch AI-capable earbuds this year, though no details on hardware products were provided.
“Over time, we think this will also unlock the ability to use voice as a kind of primary interface to computing, and to manage increasingly complex long-running agentic work. The kind of amazing use cases that we see people using Codex and ChatGPT to accomplish, we think voice can be the future interface to all kinds of work,” Eleti said.
OpenAI has been enhancing voice-based features over the past few years to make ChatGPT’s voice mode sound more natural. The company reports that over 150 million people use Voice and Dictation features to talk to ChatGPT.
Competitors are also working to make their assistants more expressive.
Both Apple and Amazon have updated their assistants to be more conversational, with improved context handling. Startups like Sesame, co-founded by Oculus co-founder Brendan Iribe and Ankit Kumar, have launched AI assistants that engage in more natural conversation while completing background tasks.
OpenAI is following a similar path, aiming to let users interact with its assistant hands-free for longer periods. Despite claiming the new voice mode sounds more natural, the company emphasized that it is not designed as an AI companion. It noted that the new models include safeguards to provide age-appropriate responses to teens and offer resources if conversations touch on topics like self-harm.
The new voice mode still has room for improvement. During a demo of the live translation feature for Hindi, the assistant spoke with a heavy American accent and used unnatural, slightly bookish Hindi. The company said the mode is optimized for “most spoken languages” but did not specify which ones.
Related article
Google Tests Remy AI Agent for Gemini as Focus Shifts to User Control
According to Business Insider, Google is testing Remy, a new AI personal agent for Gemini. This tool aims to execute tasks on behalf of users, streamlining both professional workflows and daily routines.Currently, Remy is undergoing testing in an int
Ollie bets privacy focus to win AI assistant race
To be genuinely helpful, an AI assistant must understand its user deeply. Ollie, a personal assistant designed for daily life, operates on the premise that this doesn’t require surrendering your data or compromising your privacy.While certain enterpr
How AI LIVE: London Will Explore AI & Industrial Automation
The summit will convene C-suite executives from around the globe to address pressing challenges in global industries, ranging from AI-driven disruption to economic volatility.AI LIVE: The London Summit will gather over 2,000 international leaders und
Related Special Topic Recommendations
Comments (0)
0/500

OpenAI has launched new conversational models, GPT-Live-1 and GPT-Live-1 mini, designed to sound more natural and manage turn-taking more effectively. These full-duplex models can speak and listen simultaneously, allowing users to interrupt naturally and enabling features such as live translation.
The company is also rolling out GPT-Live-1 mini as the default replacement for the current Advanced Voice Mode in ChatGPT. Paid-tier users will have access to the larger GPT-Live-1 model. Previously, the system relied on a combination of a speech-to-text model, a large language model for generating responses, and a text-to-speech model to produce the final output.
During a press briefing, OpenAI stated that the new models address issues like interrupting users while they are speaking and lacking sufficient intelligence to answer questions. The updated models will route queries to the latest text models, such as GPT-5.5, for search, reasoning, or agentic tasks while maintaining the conversation flow.
OpenAI also demonstrated that the model can remain silent for extended periods, absorbing conversational context until it is addressed. Additionally, since the new voice mode integrates with newer GPT models, it can present information visually. Other startups, such as Monogram—which raised $40 million in seed funding from DST and Lux Capital—are also focusing on visual responses to make assistants more interactive.
The company noted that the new voice mode in ChatGPT is built for longer conversations. At the briefing, ChatGPT Voice’s product lead, Atty Eleti, mentioned holding 30- to 40-minute conversations with the voice feature during walks.
OpenAI believes voice could become the primary interface for complex computing tasks. Reports suggest the company may launch AI-capable earbuds this year, though no details on hardware products were provided.
“Over time, we think this will also unlock the ability to use voice as a kind of primary interface to computing, and to manage increasingly complex long-running agentic work. The kind of amazing use cases that we see people using Codex and ChatGPT to accomplish, we think voice can be the future interface to all kinds of work,” Eleti said.
OpenAI has been enhancing voice-based features over the past few years to make ChatGPT’s voice mode sound more natural. The company reports that over 150 million people use Voice and Dictation features to talk to ChatGPT.
Competitors are also working to make their assistants more expressive.
Both Apple and Amazon have updated their assistants to be more conversational, with improved context handling. Startups like Sesame, co-founded by Oculus co-founder Brendan Iribe and Ankit Kumar, have launched AI assistants that engage in more natural conversation while completing background tasks.
OpenAI is following a similar path, aiming to let users interact with its assistant hands-free for longer periods. Despite claiming the new voice mode sounds more natural, the company emphasized that it is not designed as an AI companion. It noted that the new models include safeguards to provide age-appropriate responses to teens and offer resources if conversations touch on topics like self-harm.
The new voice mode still has room for improvement. During a demo of the live translation feature for Hindi, the assistant spoke with a heavy American accent and used unnatural, slightly bookish Hindi. The company said the mode is optimized for “most spoken languages” but did not specify which ones.
Ollie bets privacy focus to win AI assistant race
To be genuinely helpful, an AI assistant must understand its user deeply. Ollie, a personal assistant designed for daily life, operates on the premise that this doesn’t require surrendering your data or compromising your privacy.While certain enterpr
How AI LIVE: London Will Explore AI & Industrial Automation
The summit will convene C-suite executives from around the globe to address pressing challenges in global industries, ranging from AI-driven disruption to economic volatility.AI LIVE: The London Summit will gather over 2,000 international leaders und





Home






