ChatGPT Voice Enters the Desktop Arena: You Talk, a Team of AIs Get to Work

marsbitPubblicato 2026-07-27Pubblicato ultima volta 2026-07-27

Introduzione

Recently, AI pioneer Andrej Karpathy shared his preferred method for interacting with AI: using voice input to convey complex, unstructured thoughts. Instead of carefully typing out requests, he simply speaks freely for minutes at a time. He finds that AI excels at untangling these "stream-of-consciousness" monologues, returning clearer, more organized outputs than the original spoken input. This approach, he argues, improves "mind-merging" with the model by providing high-bandwidth context. Following this trend, OpenAI has integrated advanced voice capabilities into the desktop version of ChatGPT for Work and Codex scenarios, available globally for macOS and Windows. Powered by the new GPT-Live model, the feature allows real-time, interruptible conversation—users can speak to start tasks, check progress, manage multiple AI agents simultaneously, and change directions mid-task. The system can leverage project context, connected documents, calendars, and communication tools to continue unfinished work. Industry leaders like Elon Musk and Sam Altman have also emphasized the shift toward voice, noting that typing is a bottleneck for conveying complex intent to increasingly capable AI agents. Speech offers a much higher throughput, allowing users to dump extensive context quickly, even if it's messy, letting the AI handle the structuring. Beyond mere input, OpenAI's desktop voice feature enables project management through speech. Users can prioritize tasks, coordinate multiple...

Recently, Andrej Karpathy shared a lazy method he loves to use: when you want an AI to do something but can't be bothered to type out the requirements clearly, what should you do?

He said, just lean back in your chair, switch to voice, open the floodgates of consciousness, and talk for 10 minutes—rambling, messy, however you like.

Sometimes he even starts by telling the AI, "Switched to voice recognition, don't mind the typos."

The magic is, he found AI is particularly good at piecing back together these long, messy, disjointed monologues. The version it returns to you is often cleaner than what you said: that tangled mess of thoughts in your head gets untangled by it, and the number of back-and-forth corrections also decreases.

This is exactly the smoothest way Andrej Karpathy interacts with AI himself.

In his words, this improves the "mind meld" between human and model. People often can't be bothered to type out the entire context of something word for word, so why not just open their mouth and dump the whole tangled mess out.

The value of voice isn't in freeing your hands, but in dumping high-bandwidth context to the AI.

Just these past few days, OpenAI has directly moved this "command AI with your mouth" playstyle onto the desktop.

Desktop ChatGPT Can Now Be Commanded by Voice

On July 23, OpenAI officially brought ChatGPT Voice into the Work and Codex scenarios of the desktop app.

Starting that day, a global phased rollout began on macOS and Windows, covering Plus, Pro, Business, Edu, and Enterprise tiers. Android is coming soon, and iOS can access it remotely via pairing.

Under the hood is the new-generation voice model GPT-Live, launched just on July 8th, which can listen and speak simultaneously: you can interrupt at any time, and it will pick up from your latest utterance, no longer the old clunky method of "record first, transcribe, then reply."

The key is that this time, voice isn't just for casual chat.

Within Work and Codex, you can start a task, ask about its progress, check the current state of a particular agent, and even coordinate multiple Agents within the same conversation, letting them work on their own things.

A task runs in the background, and you can interrupt at any moment to change its direction; if it gets stuck or finishes, it will proactively speak up or give you a visual prompt on the screen.

It can even pick up where you left off on a current task: leveraging existing project context, connected to documents, calendar, contacts, and communication logs, to finish a half-done job.

Official examples from daily life include: having it check your calendar for conflicts, scan your inbox to see if a flight changed, and prepare meeting notes, all "while you grab coffee or attend to other things."

On macOS, there's also the added capability of Appshots and screen context, allowing it to directly read your currently focused window, combined with local files and code repositories for understanding. Hotkeys are also customizable.

Codex plus ChatGPT Work has already surpassed 10 million weekly active users.

There's a very interesting shot in OpenAI's promotional video: several engineers standing in the same room, facing the same desktop session, each giving their own commands: one AI, a room of people ordering it around.

This is probably their vision of a "group programming party."

OpenAI demonstrates ChatGPT Voice in a home scenario: speak what you want done to the desktop Codex, and it continues the work in the background.

Typing is a Bottleneck for AI Input, But Voice Isn't

This, of course, isn't a spur-of-the-moment idea from OpenAI alone.

Almost in the same week, three industry big shots almost unanimously cast their vote for "voice."

Karpathy posted on X, saying when there's too much context and he's too lazy to type, he just rambles and lets the AI organize it.

A Grok fan retweeted Karpathy's post, saying they were already used to Grok's speech-to-text, "Now typing feels like torture," and casually吐槽ed Anthropic's voice input as "actually still pretty basic."

Musk retweeted this, adding: You can use Grok Build to assign tasks to Grok just like talking to a person.

Altman didn't forget to promote his new-generation voice model GPT-Live either.

He admitted he's always preferred typing over talking to AI, but now he "talks more to ChatGPT than he types."

He specifically pointed out: Voice is most useful when you have a lot of context to dump to the AI at once.

What excites the three big shots isn't simply "finally being able to use your mouth."

Behind this lies this logic: as agents become more capable, the intentions you need to feed them become more complex, and the keyboard pipeline starts to congest:

The more concisely you type, the drier the information the model receives.

Voice output, however, is a naturally wide pipeline: you can blurt out the whole story, cause and effect, concerns, and preferences all at once. Even if it's messy, the model can handle it and then sort out that stream of consciousness for you.

Community members have already tested: voice throughput is roughly 240 words per minute, while typing maxes out at 70-80 words—a difference of three to four times.

GPT-Live offloads the heavy lifting to the backend GPT-5.5 and GPT-5.6 to think about, while the frontend just focuses on keeping the conversation flowing and catching what you say. This decoupling is precisely to let people speak freely.

Ultimately, the thoughts in the human brain are inherently parallel and jumpy, but the keyboard forces you to compress them into lines of text.

Voice loosens this wall: it's transitioning from a "lazy input method" to a "high-bandwidth context entry point."

In a word, typing is a bottleneck for AI input, but voice isn't.

Behind Voice Input Lies "Managing Projects with Your Mouth"

What OpenAI has truly packed into the desktop version this time is using voice to "manage AI."

According to the official FAQ, within Work and Codex, you can do these things by voice: start a task, prioritize several tasks, interrupt midway and change direction, while the work continues to run in the background.

Going further, it can coordinate multiple intelligent agents to work together across multiple conversations and projects. You say something like "A does this first, B stays on hold, C notify me when there's a result," and the backend will rearrange accordingly.

It can also pick up where you left off last time. This relies on project context, plus connected tools—your documents, calendar, contacts, communication logs.

When a task gets stuck or is finished, it will notify you with a voice prompt or a visual prompt on the screen.

This is no longer the job of a voice input method; it's more like a command center where you're directing a team of agents.

It's still opening your mouth to talk to AI, but previously it was about making the AI understand you; now it's about making the AI work for you.

What ChatGPT Wants to Be is Your "AI Work Operating System"

Behind a single voice feature lies a larger roadmap: OpenAI wants ChatGPT to be your "AI work operating system."

Over the past few months, OpenAI has been reworking the desktop ChatGPT, putting Chat, Work, and Codex into the same entry point, with one-click switching, running persistently in the background.

Codex, too, has long expanded into a universal platform capable of running multiple agents in parallel, handling long tasks, and understanding entire projects.

Why put this capability on the desktop, instead of keeping it on the smartphone everyone uses?

Mainly because the tiny chat window on a phone can't handle complex tasks; it's suitable for asking a couple of quick questions.

To truly let an AI connect to your files, code, and software, and see a task through from start to finish, it needs to reside on your computer:

The desktop has files, integrated development environments (IDEs), browsers, a whole suite of enterprise software—this is the main battlefield for agents to work.

Behind this is a flip in how humans interact with AI.

In the past, you used AI by "operating software": open, input, wait, one turn at a time.

Now you're more like a manager leading a team: open your mouth and delegate tasks. The relationship between you and AI has quietly shifted from "you ask, it answers" to "you say, it does."

A power user bluntly said that using this suite of things feels "basically like J.A.R.V.I.S."

He even had Codex help him configure a hotkey, allowing him to switch between three machines and several screens just by speaking.

Returning to the opening scene, what truly excites big shots like Musk, Altman, and Karpathy isn't being able to throw away the keyboard, but rather that when input is no longer the bottleneck, people can return the brainpower saved to what's truly worth pondering: decision-making.

So the next question isn't really "Can you open your mouth?" but rather, when you have a room full of on-demand AIs in front of you, what exactly do you want them to do for you?

References:

https://x.com/OpenAI/status/2080378182469857576

https://aihot.virxact.com/items/cmrxxi2qg03kcroxpcugdaldy

This article is from the WeChat public account "New Zhiyuan," author: ASI Apocalypse

Domande pertinenti

QWhat is the main advantage of using voice to interact with AI according to the article?

AThe main advantage is that voice provides a high-bandwidth channel for dumping complex, messy, and parallel streams of consciousness to the AI. It allows users to convey more context, intentions, and details effortlessly compared to typing, enabling the AI to better understand and structure the user's thoughts.

QWhat is the name of OpenAI's new speech model mentioned in the article, and what is its key feature?

AThe new speech model is called GPT-Live. Its key feature is its ability to listen and speak simultaneously, allowing users to interrupt and have the model respond fluidly based on the latest input, rather than using an older 'record, transcribe, then respond' method.

QIn which specific scenarios within the desktop ChatGPT app has OpenAI integrated Voice functionality?

AOpenAI has integrated Voice functionality into the Work and Codex scenarios within the desktop ChatGPT app.

QWhat is the broader strategic goal that OpenAI is pursuing with the integration of advanced voice capabilities into the desktop ChatGPT app?

AOpenAI's broader strategic goal is to position ChatGPT as an 'AI work operating system.' This means creating a central platform on the user's desktop where AI agents can manage tasks, coordinate with software and files, and be directed via voice commands, shifting the interaction from a question-and-answer format to a managerial 'command-and-execute' paradigm.

QAccording to the article, what is the fundamental shift in the human-AI interaction paradigm that voice-enabled task management represents?

AThe fundamental shift is from a user 'operating software' (inputting, waiting, receiving an answer) to a user acting like a 'manager with a team.' The user verbally delegates tasks, and the AI executes them, often coordinating multiple agents in the background. This changes the relationship from 'you ask, it answers' to 'you say, it does.'

Letture associate

Will Changxin Technology Continue to Rise Today?

Changxin Technology made a historic debut on the stock market, with its share price soaring 465.82% to close at 49 yuan. Its market capitalization reached 3.28 trillion yuan, surpassing Industrial and Commercial Bank of China to become the largest company by market cap on the A-share market. Daily trading volume exceeded 140 billion yuan, a first in A-share history. This created a moment of realization for 7.7 million investors who won the lottery for its shares. On the first day, investor strategies varied: some sold immediately and later regretted missing intraday highs, others secured profits to avoid future volatility, while a third group held or even bought more shares, betting on long-term growth. The staggering IPO, massive public enthusiasm, and debut during a peak industry cycle led some to compare Changxin to PetroChina's 2007 listing, which was followed by a long decline. Key similarities noted include comparable fundraising scales (approx. 666 billion yuan for Changxin vs. 668 billion for PetroChina) and both companies listing at a perceived high point in their respective commodity cycles (oil then, memory chips now). However, analysts caution against over-simplifying the comparison. They highlight core differences: Changxin operates in the high-growth semiconductor sector with strong "domestic substitution" tailwinds. Brokerages like Huaxi Securities project significant revenue and profit growth from 2026 to 2028, driven by DDR5 adoption, product mix optimization, and economies of scale. Nomura Securities issued a "buy" rating with a 116 yuan target price, citing AI-driven demand for DRAM, tight supply as major players shift to HBM production, and Changxin's vast room for market share growth. Some analysts position the current memory cycle, fueled by AI, as just beginning, contrasting with the mature energy cycle PetroChina entered. The article concludes that for investors, monitoring the memory cycle's progression and Changxin's breakthroughs in high-end technologies like HBM will be crucial, rather than relying on superficial historical parallels.

marsbit10 min fa

Will Changxin Technology Continue to Rise Today?

marsbit10 min fa

Claude Designer Lags Behind Engineers, Fires Back by Creating a Million-User Tool

The article tells the story of Nate Parrott, a designer at Anthropic who created Claude Design as a side project to keep pace with his engineering teammates. When Anthropic released Opus 4.5 in November 2025, the two engineers on Parrott's Claude Code team significantly increased their output using the new AI capabilities. Parrott, the sole designer, found himself struggling to match their speed, becoming a bottleneck in the workflow. To catch up, he began experimenting in his spare time. He initially tried prompting Claude to generate designs from text descriptions and screenshots, with limited success. His breakthrough came when he shifted focus from asking Claude to "design" to asking it to generate HTML. He realized HTML could be a rich visual canvas for creating everything from slides and interactive prototypes to full web pages. He built a simple interface with a chat panel on the left and a live HTML preview on the right. The key to making the output useful was incorporating Anthropic's brand system—fonts, colors, assets, and design principles—into the prompts. This ensured generated designs were immediately on-brand. He shared an internal prototype with his team, and other product designers quickly adopted it for creating clickable prototypes, a task traditionally requiring manually drawing every state. The project's "official" turning point came during an Anthropic Labs offsite, where Parrott noticed many attendees were using his tool to build presentation slides on the fly, sometimes right before their turn to speak. This organic adoption convinced the Labs team to formally staff the project, turning the side project into a real product. Claude Design is positioned as a "pre-production" tool for visual communication and exploration—handling slides, landing pages, PDFs, emails, and social media graphics. It integrates with tools like Canva, Adobe, and Vercel. Its core value is accelerating the early stages of design: exploring directions, building consensus, and establishing systems. For actual production code, Anthropic still recommends Claude Code. The story highlights how AI disrupts workflows unevenly and how individuals can respond by building new tools to create their own advantages. Parrott's tool, born from necessity, eventually gained over a million users in its first week.

marsbit21 min fa

Claude Designer Lags Behind Engineers, Fires Back by Creating a Million-User Tool

marsbit21 min fa

Fields Medalist Warns: AI Could Kill Mathematics

The 2026 Fields Medal award was followed by a startling announcement: laureate Jacob Tsimerman joined OpenAI to pursue AI safety research, predicting AI would surpass humans in all mathematical proof areas within two years. Soon after, Fields Medalists Terence Tao and Timothy Gowers expressed deep concern at ICM 2026. Gorges warned that AI might "kill" mathematics not through stagnation, but through an overwhelming surplus of proofs, likening it to a lake dying from eutrophication. This concern is echoed in the "Leiden Declaration," signed by over 3,000 mathematicians including Tao, Peter Scholze, and Kevin Buzzard, advocating for mathematics as a profoundly human endeavor. However, Gowers, who did not sign, fears a future where AI-generated mathematics proliferates while human expertise and the shared intuition vital to the field vanish, turning math into an unvisited "cemetery of thought." Gowers' perspective shifted dramatically after testing ChatGPT 5.5 Pro. The AI solved a doctoral-level number theory problem and later produced a counterexample for the "unit distance problem," achievements Gorges considered publishable in top journals. He now concedes that large language models can handle advanced research, a realization that left him feeling the "rug pulled out from under" him when AI solved problems he personally contemplated. The debate extends to the nature of mathematical discovery. As noted by Peter Woit, AI agents have no interest in the "credit game" of academia. If theorems cease to be attributed to individual mathematicians, truth may simply return to its impersonal state in the universe. The central question remains: in an age of potentially limitless AI-generated discovery, what is the role and purpose of the human mind in mathematics?

marsbit52 min fa

Fields Medalist Warns: AI Could Kill Mathematics

marsbit52 min fa

NVIDIA's 20-Year CUDA Moat Collapsed Over a Weekend, Claude Single-Handedly Got AMD's New GPU Running

In a single weekend, Claude, an AI agent from Anthropic, successfully ported and optimized its cutting-edge model to run on a brand-new AMD MI355X server rack without any manual code intervention. This feat demonstrates a potential breakthrough in overcoming NVIDIA's long-established CUDA software ecosystem dominance, built over two decades. Anthropic's team simply instructed Claude to get the AMD machine running. By Monday, it not only worked but was showing a continuously improving performance curve. The achievement impressed AMD CEO Lisa Su and accelerated a major deployment partnership: Anthropic plans to deploy up to 2GW of AMD Instinct GPUs starting in 2027. The key enabler is AMD's new ROCm.AI platform, a toolbox designed specifically for AI agents like Claude. It provides AI-readable documentation, including chip instruction sets (ISA), and tools like the Hyperloom service that allows agents to autonomously profile performance, identify bottlenecks, test configurations, and generate optimized kernels. In a demo, Hyperloom boosted the output speed of a model by 38%. This represents a fundamental shift. While CUDA's strength lies in its vast, human-expert-driven ecosystem of tools and tacit knowledge, AMD's strategy is to make its hardware and software stack directly accessible and optimizable by AI agents. An agent can parallelize tasks—debugging, profiling, coding—that would take human engineers years to master, compressing the traditional software adaptation timeline from years to tasks. The competition is no longer just about peak hardware specs but also about how well AI can read, utilize, and tune a platform.

marsbit52 min fa

NVIDIA's 20-Year CUDA Moat Collapsed Over a Weekend, Claude Single-Handedly Got AMD's New GPU Running

marsbit52 min fa

Just Now, Peking University Alumna Lilian Weng Announces Resignation: The Best 'Alignment' in the AI Era Is Taking Care of Yourself

Lilian Weng, a Peking University alumnus and former OpenAI executive, has announced her departure from Thinking Machines Lab, the AI startup she co-founded with ex-OpenAI CTO Mira Murati 20 months ago. Her resignation comes shortly after the company released its first open-source model, Inkling. In her farewell message, Weng cited health reasons as the primary factor, stating that the past seven months involved more illness than any other period in her life. The intense pressure and guilt of being unable to work during critical periods, like the Inkling launch, became unsustainable. She expressed that she could not continue in a role where she felt unable to give her full effort. Weng joined OpenAI in 2018, contributing to projects like the Dactyl robotic hand and later building the Safety Systems team from scratch. She is also widely known for her influential technical blog, Lil'Log. Her departure follows a period of high expectations and challenges for Thinking Machines, which secured a record $20 billion seed round at a $120 billion valuation in 2025 but has since seen fluctuations in its valuation and the departure of other key co-founders. Weng emphasized that she remains passionate about AI but needs a more predictable and defined role without the relentless pressure of a co-founder. Her exit highlights the immense personal toll and speed of competition in the AI industry, where the drive for progress often clashes with human limits. Her decision underscores that stepping back to prioritize well-being can be an act of courage.

marsbit1 h fa

Just Now, Peking University Alumna Lilian Weng Announces Resignation: The Best 'Alignment' in the AI Era Is Taking Care of Yourself

marsbit1 h fa

Trading

Spot
活动图片