Google, OpenAI, and Microsoft are actively investing in technologies that enable communication with AI agents via voice, banking on the increasing popularity of oral communication with AI compared to text queries. This direction is increasingly referred to as voice AI. Investor interest is confirmed by PitchBook data cited by the Financial Times: voice AI startups raised about $7 billion in the first quarter of 2026 – seven times more than in the same period of 2025.
The calculation is based on the expectation that oral speech will gradually displace text input as the primary way to communicate with AI. Indicators from major market players confirm this.
Google Records Growth in Voice and Visual Queries
According to Google, the number of queries in AI Mode more than doubled every quarter since the service's launch up until April-May 2026. In the US, voice or images are now used in more than every sixth search query, and the number of image-based inquiries is growing by more than 40% per month. The company links this dynamic to users increasingly choosing a communication format close to natural speech instead of typing text.
OpenAI's Bet on Voice
Out of roughly 900 million weekly active ChatGPT users – a figure the company announced in late February 2026 – more than 150 million people use voice and dictation to communicate with the service every week.
On July 8, 2026, OpenAI launched the GPT-Live family of models: GPT-Live-1 and the lightweight GPT-Live-1 mini. These are full-duplex models capable of simultaneously listening to the interlocutor and responding, without the familiar "query-response" pause. They form the basis of the updated ChatGPT Voice.
In addition to the software, OpenAI is preparing a separate hardware device. According to Bloomberg sources, the company is developing a screenless smart speaker with a camera and sensors based on GPT-Live. An announcement could happen as early as 2026, with a market release expected in early 2027.
Microsoft Develops Expressive Speech Synthesis
Microsoft presented the MAI-Voice-2 model in June 2026 – an expressive speech synthesizer supporting 15 languages and allowing control over the emotional tone of voice. The technology is already being integrated into the company's products, including Dynamics 365 Contact Center and Azure Voice Live.
Key areas where companies are betting on voice AI include:
Search and interaction with AI agents via oral speech and images
Full-duplex voice models for more natural dialogue
Separate hardware devices for voice interaction
Speech synthesis with emotion control for corporate services
The growth in investments and simultaneous announcements from Google, OpenAI, and Microsoft show that voice interfaces are transitioning from experimental features to a standalone direction for AI product development. The further spread of such technologies will depend on how quickly users adopt these new interaction methods in everyday scenarios.
AI Opinion
From the perspective of regulatory dynamics, the voice AI boom carries a risk that investors and developers have overlooked: the user's trust in the "humanity" of the interface. While Google, OpenAI, and Microsoft compete in dialogue naturalness, State Duma deputies have already submitted a proposal to the Federal Antimonopoly Service (FAS) to oblige voice robots to disclose their nature from the first seconds of a call. The gap between the technological race for the "presence effect" and regulators' attempts to preserve the boundary between human and machine could become as significant a market factor as the volume of investments.
Historical patterns suggest caution: voice assistants already went through a cycle of inflated expectations during the early versions of Siri and Alexa, when mass adoption did not justify monetization forecasts. The question remains open as to whether the current wave of full-duplex models will be a real shift in technology perception or just another turn of the same cycle.
end-content




