ChatGPT's voice mode has been rewritten: It will make you forget you are talking to AI
OpenAI has introduced the GPT-Live-1 and GPT-Live-1 mini models, aiming to make the ChatGPT voice experience more natural.
12punto
OpenAI has announced new voice models for ChatGPT's voice mode. The models, named GPT-Live-1 and GPT-Live-1 mini, aim to allow users to speak with artificial intelligence in a more natural and uninterrupted manner.
According to information provided by the company, the new models are more successful at managing pauses and transitions during conversation. This is intended to make voice conversations with ChatGPT feel more like a two-way dialogue rather than a classic command-response experience.
ABLE TO SPEAK AND LISTEN AT THE SAME TIME
One of the standout features of the new voice models is support for full-duplex communication. Thanks to this structure, the assistant can continue speaking while simultaneously listening to the user's voice. This aims to provide a more fluid experience, especially during long conversations.
It has been stated that the current voice mode in ChatGPT will be replaced by the GPT-Live-1 mini model by default. Users on paid subscription tiers will have access to GPT-Live-1, which is offered as a larger and more advanced model.
OpenAI's previous voice infrastructure consisted of multiple stages, such as a speech-to-text system that converted the user's voice into text, a language model that generated the response, and a text-to-speech system that voiced that response. It is stated that the new models transform this structure into a more direct, voice-based experience.
The company emphasizes that the new voice mode was developed for long-form conversations. According to reports, an OpenAI executive said they were able to have uninterrupted conversations lasting 30 to 40 minutes through this feature during their walks. The new structure is also expected to reduce issues such as the assistant unnecessarily interrupting users.
For more complex queries requiring search, reasoning, or agent capabilities, the voice mode can consult the company's current text model, GPT-5.5, in the background. The goal is for the voice chat flow to continue without interruption during this process.
However, it is not yet clear whether the new system will offer the same level of performance in all languages. In a live Hindi translation demo, it was reported that the assistant spoke with a distinct American accent and that the language sounded artificial. Although OpenAI states that the new mode is optimized for "most of the most spoken languages," it did not share a detailed list of languages.