From Text to Voice: Google, OpenAI, and Microsoft Invest Billions in AI Verbal Communication

cryptonews.ru2026-08-07 tarihinde yayınlandı2026-08-07 tarihinde güncellendi

Özet

Major tech companies Google, OpenAI, and Microsoft are heavily investing billions into AI-powered voice communication, anticipating a shift from text-based to spoken interaction with AI. Data shows a sevenfold increase in funding for voice AI startups in early 2026 compared to the same period in 2025. Google reports that voice and image queries now constitute over one-sixth of searches in the US, with visual searches growing monthly. OpenAI revealed that more than 150 million of its weekly ChatGPT users engage via voice. The company launched its full-duplex GPT-Live models, enabling simultaneous listening and responding, and is reportedly developing a screenless smart speaker. Microsoft introduced MAI-Voice-2, an expressive speech synthesis model with emotional control for 15 languages, integrated into its enterprise services. Key investment areas include voice/image-based search, full-duplex dialogue models, dedicated hardware for voice interaction, and emotionally controlled speech synthesis for business applications. This trend indicates voice interfaces are evolving from experimental features into a core AI development focus. However, rapid advancement brings regulatory challenges, such as potential requirements for AI to disclose its non-human nature immediately. The history of voice assistants also suggests caution, as previous cycles of high expectations, like with early Siri and Alexa, were followed by disappointing monetization. The success of this new wave hinges...

Google, OpenAI, and Microsoft are actively investing in technologies that enable communication with AI agents via voice, banking on the growing popularity of verbal interaction with AI compared to text queries. This field is increasingly referred to as voice AI. Investor interest is confirmed by PitchBook data cited by the Financial Times: voice AI startups raised about $7 billion in the first quarter of 2026—seven times more than in the same period in 2025.

The bet is that spoken language will gradually displace text input as the primary way to interact with AI. Metrics from major market players confirm this trend.

Google Observes Growth in Voice and Visual Queries

According to Google, the number of queries in AI Mode has more than doubled every quarter since the service's launch up until April–May 2026. In the US, voice or images are now used in more than every sixth search query, and image-based requests are growing by over 40% per month. The company attributes this trend to users increasingly choosing a communication format closer to natural speech instead of typing.

OpenAI's Bet on Voice

Out of roughly 900 million weekly active ChatGPT users—a figure the company announced in late February 2026—more than 150 million people use voice and dictation to communicate with the service each week.

On July 8, 2026, OpenAI launched its GPT-Live model family: GPT-Live-1 and the lightweight GPT-Live-1 mini. These are full-duplex models capable of simultaneously listening to the interlocutor and responding without the typical 'query-response' pause. They form the foundation of the updated ChatGPT Voice.

Beyond software, OpenAI is also preparing a dedicated hardware device. According to Bloomberg sources, the company is developing a screenless smart speaker with a camera and sensors based on GPT-Live. An announcement is possible as early as 2026, with a market release expected in early 2027.

Microsoft Advances Expressive Speech Synthesis

In June 2026, Microsoft introduced the MAI-Voice-2 model—an expressive speech synthesizer supporting 15 languages and allowing control over the emotional tone of the voice. The technology is already being integrated into the company's products, including Dynamics 365 Contact Center and Azure Voice Live.

Key areas where companies are focusing on voice AI include:

  • search and interaction with AI agents through spoken language and images

  • full-duplex voice models for more natural dialogue

  • dedicated hardware devices for voice interaction

  • speech synthesis with emotion control for corporate services

The rise in investments and simultaneous announcements from Google, OpenAI, and Microsoft indicate that voice interfaces are transitioning from experimental features to a standalone direction for AI product development. The further spread of such technologies will depend on how quickly users adopt new interaction methods in everyday scenarios.

AI Opinion

From the perspective of regulatory dynamics, the voice AI boom carries a risk overlooked by investors and developers: the user's trust in the 'humanity' of the interface. While Google, OpenAI, and Microsoft compete in dialogue naturalness, Russian State Duma deputies have already submitted a proposal to the Federal Antimonopoly Service (FAS) to oblige voice robots to disclose their nature from the first seconds of a call. The gap between the technological race for the 'presence effect' and regulators' attempts to preserve the boundary between human and machine could become as significant a market factor as investment volume.

A historical pattern suggests caution: voice assistants have already gone through a cycle of inflated expectations during the early versions of Siri and Alexa, when mass adoption did not meet monetization forecasts. The question remains open whether the current wave of full-duplex models will be a real shift in technology perception or just another turn of the same cycle.

end-content

İlgili Sorular

QWhat are the key areas where major tech companies are investing in voice-based AI interaction?

AThe key areas include search and working with AI agents via voice and images, full-duplex voice models for more natural dialogue, dedicated hardware devices for voice interaction, and speech synthesis with emotional control for corporate services.

QAccording to the article, how has user adoption of voice and visual queries changed on Google's platform?

AGoogle reports that queries in AI Mode more than doubled each quarter from its launch until April-May 2026. In the US, voice or images are now used in more than every sixth search query, with image-based searches growing over 40% monthly.

QWhat is significant about OpenAI's GPT-Live models announced in July 2026?

AOpenAI's GPT-Live-1 and GPT-Live-1 mini are full-duplex models, meaning they can listen and respond simultaneously without the typical 'query-response' pause, which forms the foundation for the updated ChatGPT Voice.

QWhat regulatory concern is raised regarding the development of highly natural voice AI?

AThe regulatory concern is that as AI voices become more human-like, there is a risk of users mistaking them for humans. Russian lawmakers, for example, have proposed requiring voice robots to disclose their artificial nature from the first seconds of a call.

QWhat cautionary historical pattern does the article mention in relation to the current voice AI boom?

AThe article mentions the cycle of inflated expectations experienced by early voice assistants like Siri and Alexa, where mass adoption did not meet monetization forecasts, questioning if the current full-duplex model wave represents a real shift or another phase of the same cycle.

İlgili Okumalar

İşlemler

Spot
活动图片