Voice is increasingly becoming the command through which work is handed over to artificial intelligence. As long as systems only answered a question, typing remained the most precise method.
Now that agents can act on a file and carry out a task for several minutes, speaking offers a different advantage. It makes it possible to explain quickly what is wanted and to correct the software while it works.
A Shift in the Architecture of Voice Models
What made this shift possible was a change in the architecture of voice models. Older assistants split every exchange into three steps: they transcribed the voice, passed the text to a language model, and finally converted the answer into audio.
The process worked for asking about the weather or setting an alarm, but it introduced a noticeable delay. In the move to text it often left out elements such as intonation and rhythm, and it could interpret second thoughts and self-corrections as separate fragments.
GPT Live, presented by OpenAI on 8 July, processes the conversation in real time. Thanks to a full-duplex architecture, the system keeps listening even while it responds, as in a phone call between two people. The user can interrupt it or change the instruction.
A Fast-Growing Market
According to Grand View Research, voice agents are already worth 3.5 billion dollars and by 2033 could reach 35.2 billion, with an average annual growth of 39 per cent.
Data released by Google and reported by the Financial Times help explain the interest of big tech firms. Voice conversations last on average five times longer than text ones, and the use of the live mode doubled in the year ending in April. OpenAI states that more than 150 million people use voice or dictation in ChatGPT every week.
Cost remains a constraint. A conversation occupies the system for its whole duration, and latency has to be reduced. Boson AI claims that its speech-to-speech model can run at about a tenth of the cost of competitors, while Smallest.ai separates the conversation from the more demanding reasoning.
Applications and Limits
Coding is one of the first areas where the mode has found concrete use, because speaking makes it easier to explain how a problem arose. Fonio.ai, which claims more than two million calls a month for 7,500 companies, uses agents for support and appointments.
Regulatory constraints remain. From 2 August 2026, Article 50 of the AI Act applies: a voice agent must inform the person that they are interacting with an AI system, unless this is already evident.
There are also practical limits. Speaking to a computer in a shared office disturbs those nearby and can make confidential information audible. For this reason the hardware industry is working on directional microphones and dedicated devices.
Original Article: linkiesta.it



