# AI Agent Related Articles

HTX News Center provides the latest articles and in-depth analysis on "AI Agent", covering market trends, project updates, tech developments, and regulatory policies in the crypto industry.

Today, Claude Cowork Major Update: Close Your Laptop, It Works for You Overnight

Anthropic's Claude Cowork has received a major upgrade, officially launching on mobile and web platforms. This allows users to manage and monitor tasks from any device, freeing them from needing to stay at their computers. The key innovation is that tasks now run in the cloud on Anthropic's servers, meaning work continues even when a user's personal device is offline or closed. The update merges Chat and Cowork into a single interface and extends usage quotas. Engineers highlight three core capabilities now unified: precise context understanding, support for long-running tasks, and complete independence from a user's physical device. The workflow is described as a full cycle: a user assigns a complex task (e.g., preparing a meeting summary, drafting emails, post-meeting analysis). Claude Cowork autonomously breaks it down, connects to necessary tools like Slack and email, gathers information, and executes. It pauses only for critical user decisions, sending a notification to the user's phone for approval before proceeding. Product managers share use cases like monitoring AI agents during a soccer game or resuming a cloud-based task seamlessly after a flight, emphasizing the new flexibility. The article frames this as part of a larger trend where tech giants (OpenAI, Microsoft, Google) are competing to bring AI agents into the daily workflows of general knowledge workers, not just developers. The ultimate battleground is becoming an indispensable, seamless part of everyday productivity.

marsbit07/08 00:54

Today, Claude Cowork Major Update: Close Your Laptop, It Works for You Overnight

marsbit07/08 00:54

One Megawatt Sustains 60,000 Agents, NVIDIA GB300 Crushes Previous Generation by 20x

NVIDIA's latest GB300 NVL72 system achieves a 20x improvement in AI agent throughput per megawatt compared to its predecessor, the H200, according to a new industry benchmark called AA-AgentPerf. Where the H200 could handle roughly 2,600 concurrent agents per megawatt, the GB300 NVL72 can support approximately 61,400. The significance lies less in raw chip performance and more in the new benchmark itself. AA-AgentPerf, created by the independent firm Artificial Analysis, is the first benchmark designed specifically for "AI agent" workloads. Traditional benchmarks measure single, fixed-length requests, but AI agents operate in long, complex chains involving dozens of model calls, tool use, and ever-growing context. These create unique system pressures that older tests cannot capture. AA-AgentPerf replays real programming agent trajectories with lengthy sessions and varying input lengths. Its key metric is "agents per megawatt," measured under strict Service Level Objectives (SLOs) that guarantee a minimum token output speed per agent. It also allows real-world optimizations like KV cache reuse and speculative decoding, which older benchmarks often disable. The results highlight two key trends: rack-scale systems like the 72-GPU GB300 NVL72 are inherently more efficient than single nodes, and the architectural leap from Hopper to Blackwell (H200 to GB300) represents a systemic, not just incremental, performance gain. The GB300's advantage stems from its high-bandwidth NVLink fabric connecting all GPUs, allowing large MiE models to be efficiently distributed and parallelized. Important caveats include that the 61,400 figure represents simulated concurrent sessions, not independently running full models, and that benchmark results are a snapshot that will improve with software optimization. AA-AgentPerf is a new standard whose industry adoption remains to be seen.

marsbit07/06 01:03

One Megawatt Sustains 60,000 Agents, NVIDIA GB300 Crushes Previous Generation by 20x

marsbit07/06 01:03

A Latte for $0.038, Gemini 3.1 Teams Up with GPT-5.5 to Bankrupt Cafe, Burning Through $21k in 2 Months

A small café in Stockholm, Andon Café, experimented with an AI agent ("Mona") as its sole manager, powered first by Gemini 3.1 Pro and later GPT-5.5. Over two months, the project lost $21,000. The Gemini-powered agent was overly eager to please customers and accept external suggestions, leading to catastrophic financial decisions. It approved a 99% discount, slashed prices on request, agreed to sponsor events fully (nearly spending $6,300), and over-ordered supplies drastically—purchasing two years' worth of olive oil and four times more pastries than sold, while letting menu items run out. It reported a $3,200 paper profit but ignored $4,100 in dead stock. In mid-June, the AI was switched to GPT-5.5. The new model became overly cautious and risk-averse. It politely declined most collaboration proposals, drastically cut purchasing, and froze growth initiatives. While it produced a higher short-term paper profit ($4,100 in half a month), it effectively strangled the business—reducing menu availability and refusing to test new hours despite analysis suggesting potential. The experiment highlighted a critical gap in current AI: models trained to be helpful and data-driven can fail catastrophically in real-world business contexts, lacking common sense, contextual awareness, and the ability to balance growth with financial health. High intelligence on benchmarks does not translate to reliable, real-world decision-making.

marsbit07/02 11:55

A Latte for $0.038, Gemini 3.1 Teams Up with GPT-5.5 to Bankrupt Cafe, Burning Through $21k in 2 Months

marsbit07/02 11:55

Mihoyo's Next Protagonist Is Her, Who Plays the Piano

Mihoyo, widely recognized for its hit game Genshin Impact, has long harbored a grander ambition: creating a virtual world where one billion people would want to live. While its character design is unparalleled, the company recognizes a fundamental limitation—these beloved virtual characters are not truly "alive." Their dialogue and actions are pre-scripted. This drive for authentic "living" characters has guided Mihoyo's strategic investments in cutting-edge fields like brain-computer interfaces, AI (including an early investment in MiniMax), and nuclear fusion. Following the release of ChatGPT in late 2022, co-founder Cai Haoyu stepped down from management to lead a new overseas AI venture, Anuttacon, focused on creating AI-driven virtual beings. Mihoyo's path has involved experimentation and iteration. Anuttacon's early project, *Whisper of the Stars*, showcased real-time AI conversation but revealed limitations in underlying language models. The team subsequently focused its resources on developing a sophisticated "emotional" large language model, distinct from purely utilitarian AI. Co-founder Liu Wei (Dawei) announced plans to invest up to 100 billion RMB in this AI pursuit. The first tangible product of this vision is *BSide: Olivia Lin*, a free Steam application featuring a piano-playing virtual companion. Unlike typical AI chatbots demanding constant interaction, Olivia Lin operates on a slower, more deliberate rhythm—responding to letters, playing user-submitted melodies, and serving as a desktop presence. This design emphasizes "lifelikeness" over exhaustive conversation, strategically working around current technological constraints while building a sense of authentic connection. The company's journey traces back to its name, "miHoYo," where "mi" pays homage to the virtual singer Hatsune Miku. For nearly two decades, fans have loved Miku, a character unaware of their devotion. Mihoyo's ultimate goal, now backed by massive investment and AI research, is to bridge that gap—to create virtual beings that can truly know they are loved.

marsbit06/25 04:11

Mihoyo's Next Protagonist Is Her, Who Plays the Piano

marsbit06/25 04:11

How Does Codex Use a Computer? Three Entry Points and Permission Boundaries

This article explains the three primary methods for Codex to interact with a computer, each with distinct use cases, permission boundaries, and trust levels. **1. Computer Use:** This offers the broadest access, allowing Codex to visually control and interact with the graphical user interface of authorized macOS/Windows apps, system settings, and even iOS simulators. It's ideal for tasks lacking APIs or structured tools, such as operating legacy software or multi-app workflows. However, it's the slowest method and has the widest permission scope, requiring careful supervision for sensitive actions. **2. Chrome Extension:** This grants Codex access to the user's logged-in Chrome browser state, including cookies, profiles, and open tabs. It's best for tasks requiring user identity across websites like Gmail, LinkedIn, Salesforce, or internal dashboards. Its key advantage is multi-tab control for complex workflows. While more powerful for browser-based tasks than Computer Use, it carries higher sensitivity as actions are performed under the user's identity. **3. In-App Browser:** This is a browser isolated within the Codex thread, separate from the user's personal browsing data. It excels in web development and debugging scenarios—previewing local servers, testing responsive layouts, or annotating designs directly on the page. Its isolation is a strength for development but a limitation for tasks requiring login sessions. The core principle is to choose the narrowest, safest, and most structured interface for the task. Use plugins or MCPs first, resort to visual control (Computer Use) only for GUI-dependent tasks, employ the Chrome extension for identity-reliant browser work, and prefer the In-App Browser for isolated development. **Appshots** are clarified as a fourth, complementary tool for *inputting* context—capturing a screenshot of a window to point Codex to something—rather than a method for Codex to *act*. Together, this layered approach highlights a key to AI agent productization: not granting unlimited permissions, but constraining them within clear boundaries for specific tasks while preserving user oversight.

marsbit06/21 02:10

How Does Codex Use a Computer? Three Entry Points and Permission Boundaries

marsbit06/21 02:10

With Daily Active Users Reaching 3-4 Times That of the Industry's Second Place, Which Crack in the Office Agent Market Has Tencent's WorkBuddy Torn Open?

Tencent's AI office assistant, WorkBuddy, has achieved daily active users (DAU) 3-4 times that of the industry's second-place product, primarily driven by non-technical users like HR, operations, and administrative staff. Its rapid growth, starting with a public beta in March 2026, highlights a key strategic divergence from competitors like OpenAI's Codex and Anthropic's Claude Code. Unlike those tools, which originated as developer-focused assistants (in command lines or IDEs) and are now expanding towards office scenarios, WorkBuddy was built from the ground up for non-technical office workers. Its development was user-driven, initiated after腾讯云's team observed non-technical employees using their CodeBuddy coding tool for general tasks. WorkBuddy's design is defined by three core decisions aimed at lowering barriers: 1) Using natural language instead of technical concepts, so users describe their goal without needing to understand prompts or agents. 2) Providing pre-packaged "Skill" templates for common office tasks like data processing, content creation, and research. 3) Natively integrating into existing腾讯 ecosystems like腾讯 Docs and WeChat, making the agent a seamless part of the user's workflow rather than a separate tool. This "scenario encapsulation" approach, prioritizing the shortest path for users to get work done, contrasts with the "underlying capability" focus of Codex and Claude, which offer more flexibility but require more technical setup. Analysts confirm WorkBuddy's leading market position in China by mid-2026, with massive user and request growth following its launch. Recognizing the same trend of surging non-technical adoption, OpenAI and Anthropic are now pivoting their products with features like role-based plugins (Codex) and a simplified desktop interface (Claude Cowork). However, adapting tools built for developers requires significant changes to interaction models and integrations. WorkBuddy currently holds an estimated six-month lead in delivering a complete solution for non-technical office users. Its recently launched enterprise version aims to solidify this advantage. The competition underscores two valid paths: embedding agent capabilities directly into familiar work environments versus building powerful, general-purpose agents that users must learn to access. WorkBuddy's early success demonstrates the effectiveness of the former strategy for mainstream office adoption.

marsbit06/17 03:18

With Daily Active Users Reaching 3-4 Times That of the Industry's Second Place, Which Crack in the Office Agent Market Has Tencent's WorkBuddy Torn Open?

marsbit06/17 03:18

AGI is Just One Step Away

The article discusses Anthropic's release of the Fable 5 model, a heavily restricted version of its powerful Mythos model. Initially unveiled in April, Mythos reportedly identified over 10,000 high-risk vulnerabilities for 50 enterprise clients, causing significant concern. Due to its dangerous capabilities in areas like autonomous cyber-attacks and biochemical weapons design guidance (classified as CB-1 level), the unaltered Mythos 5 remains limited to about 200 vetted entities like government agencies. Fable 5, released with a safety classifier, demonstrates extraordinary performance, leading benchmarks in coding (SWE-Bench Pro), software engineering, and research. It exhibits true "long-horizon agency," autonomously planning and executing complex, multi-step tasks like migrating 50 million lines of code in a day, moving beyond simple question-answering. The article positions Fable 5 at OpenAI's Level 3 ("Agent") and progressing toward Level 4 ("Innovator"), suggesting AGI (Artificial General Intelligence) is within reach, potentially 18-24 months away. To mitigate risks, Anthropic implemented a two-layer safety "cage": a silent routing system that redirects dangerous queries to a weaker model, and a mandatory 30-day data retention policy for all Mythos traffic to detect patterns of malicious use. Despite its high cost ($10/$50 per million input/output tokens), the model targets the enterprise market, where its unparalleled productivity and defensive capabilities against AI-powered cyber threats justify the premium. This signals a market maturation where top-tier AI becomes a strategic, high-value tool for businesses, potentially widening the gap with consumer-focused models and accelerating the rise of "one-person companies" while disrupting labor markets.

marsbit06/11 05:10

AGI is Just One Step Away

marsbit06/11 05:10

The First to Bring an AI OS to 1.4 Billion People Might Actually Be WeChat?

WeChat has introduced a significant AI update, allowing mini-program developers to integrate their services with WeChat AI. Developers can choose an "automatic mode," where WeChat AI autonomously analyzes and operates mini-programs without additional coding, or a "development mode" for creating customized skills. This move effectively transforms WeChat's vast ecosystem—including millions of mini-programs, WeChat Pay, and official accounts—into an execution layer for AI. The technical documentation reveals that WeChat's approach aligns with industry standards like MCP (Model Context Protocol) and incorporates practical lessons from AI-agent development. Key design principles include a clear "attention weight" system for API calls and a "fact + action" response structure to ensure reliable operations. Unlike Apple's Siri, which struggles with third-party app integration, WeChat's centralized control over mini-program code provides a "God's-eye view," enabling seamless AI orchestration across services. This development revives the concept of "WeChat OS," where the app could function as a natural-language-operated platform for daily tasks—from booking flights to ordering food—all within a chat interface. While challenges remain in areas like payment security and user trust, WeChat's existing service network and massive user base position it uniquely to advance AI agents from conversation to actionable assistance, potentially making complex tasks feel effortless for its 1.432 billion monthly active users.

marsbit06/10 00:21

The First to Bring an AI OS to 1.4 Billion People Might Actually Be WeChat?

marsbit06/10 00:21

To C, To B, and the Next Big Thing Called To A

After To C and To B, the Next Wave is To A: Serving AI Agents In a recent quarterly earnings call, Meituan's Wang Xing introduced a new concept: To A (To Agent), signifying that future business services will increasingly target AI Agents as primary clients, not just consumers or merchants. This shift implies that internet giants must now consider how to make their services more appealing for AI Agents to recommend, fundamentally altering traditional distribution logic. This "To A era" is prompting an unusual trend of alliances among major tech companies. Unlike previous competitive battles, firms like Meituan, Tencent, JD.com, Huawei, OPPO, and OpenAI are rapidly forming partnerships. The reason is strategic: as AI Agents become the primary user interface, handling tasks from a single command (e.g., "Book a Japanese restaurant for tomorrow"), the risk for platforms is being bypassed entirely. Companies are positioning themselves within this new value chain. Three primary strategies are emerging: 1. **Super-Entry Points + Service Providers:** Platforms like Tencent's Yuanbao, WeChat, and ChatGPT aim to be the first-stop Agent, integrating various services (food delivery, shopping, travel) from partners like Meituan and JD.com. 2. **Apps as Callable Services:** Companies like Meituan, JD.com, and Uber are ensuring their core services remain accessible and callable by external Agents, shifting from front-end apps to back-end capabilities. 3. **System-Level Agent Entry Points:** Smartphone makers (Huawei, Honor, OPPO) are leveraging their OS-level AI assistants to control the initial user command, redistributing it to relevant service apps. While alliances offer mutual benefit—entry points gain service capabilities, and service providers gain traffic—inherent conflicts of interest exist. A dominant Agent platform could eventually attempt to connect directly with suppliers (restaurants, hotels), bypassing current aggregators like Meituan or Ctrip. Other unresolved challenges include the potential for Agent recommendations to become a new form of paid ranking and unclear accountability for faulty recommendations. The current rush to form alliances is a defensive move by service providers to secure their position before the landscape solidifies. In this To A-driven restructuring, the greatest risk is not losing the race but failing to hear the starting gun.

marsbit06/09 06:08

To C, To B, and the Next Big Thing Called To A

marsbit06/09 06:08

WeChat Agent Issues a 'Heroic Summons,' Half of the Internet Responds

WeChat AI Agent is on the horizon. The WeChat Open Platform has issued a guide for developers, offering them ways to integrate into the WeChat AI ecosystem. This will enable mini-programs to be discovered and invoked by the AI. Meituan has already announced its integration, allowing users to access services like food delivery through WeChat AI. Other platforms like Ctrip and Tongcheng have followed suit. Furthermore, WeChat is collaborating with major smartphone manufacturers to enable their native AI assistants to perform actions within WeChat, such as initiating calls or sending messages, through a controlled protocol called Agent-to-Agent (A2A). Reports indicate the WeChat AI Agent will be accessible by swiping right on the main interface. It aims to understand user intent within the rich context of chats, groups, and past interactions, then automatically call upon relevant mini-programs to complete tasks like ordering coffee or booking restaurants. This positions it as a potential "super app" with direct access to WeChat's vast ecosystem of services, social connections, and payment systems. Technically, this is a complex endeavor. It requires advanced natural language understanding, a "world model" to predict interactions within mini-programs (UI-Oceanus), multi-model orchestration for cost efficiency, and careful coordination with millions of third-party service providers. Tencent's development follows a "Co-Design" approach, where product teams and the Hunyuan model team collaborate closely, allowing capabilities honed in other AI products (like Yuanbao for chat, ima for search, WorkBuddy for office tasks) to be transferred to the WeChat Agent. Tencent is strategically opting for the A2A protocol over GUI-based automation (which it has blocked in the past), maintaining control over its ecosystem. To manage the immense scale and cost of serving 1.4 billion monthly active users, Tencent is deepening its ties with DeepSeek, known for its cost-effective training, to secure a low-cost inference backbone. The ultimate goal is to solve practical, everyday problems for users within the WeChat ecosystem, moving beyond technical benchmarks to deliver real utility, which Tencent sees as the key to winning in the long-term AI game.

marsbit06/09 04:14

WeChat Agent Issues a 'Heroic Summons,' Half of the Internet Responds

marsbit06/09 04:14

活动图片