OpenAI Unveils GPT-Live to Bridge the Gap Between AI and Human Speech
Communicating with a virtual assistant was traditionally akin to interacting on a digital walkie-talkie: you would talk, pause, and then hear an artificial response. Going against the conventional flow, OpenAI has launched GPT-Live, a duo of voice models from the future generation that enables ChatGPT to simultaneously listen and speak.
The main novelty behind this innovation is the so-called “full-duplex” approach used by engineers. While earlier solutions required complete silence to start answering a user’s question, GPT-Live continuously analyzes incoming audio input several times each second.
This allows people to interrupt the assistant midway through its speech in order to redirect it without causing any confusion. Moreover, the AI throws in casual human verbal expressions such as “mhmm” and “got it” in order to demonstrate its comprehension of the speech, or remains silent if it detects a mere pause.
Two versions of OpenAI are deployed globally through the iOS, Android, and web interfaces. Free users get an automatic update to the GPT-Live-1 mini version, while those paying for the Go, Plus, and Pro subscriptions get access to the advanced GPT-Live-1 version.
Notably, the system is capable of conducting web searches live in the middle of the vocal conversation. Upon receiving a challenging question, the voice interface uses background models such as GPT-5.5 to engage in the deep reasoning process.
Although the initial tests showed minor language peculiarities, such as the unusually bookish style of expression and a strong American accent during live tests in languages such as Hindi, this update represents a major breakthrough in design.
With the voice being transformed into a highly flexible, hands-free user interface that can operate for 40-minute-long discussions, OpenAI is setting the stage for future interactions in which humans will be able to easily control complicated computing processes via natural speech.