Anthropic Data: Nearly Half of AI Agent Calls Concentrated in Software Engineering, These 16 Vertical Domains Remain Blue Oceans

marsbitPublicado a 2026-02-24Actualizado a 2026-02-24

Resumen

According to Anthropic's comprehensive study on real-world AI Agent usage, nearly 50% of all tool usage by AI agents is concentrated in software engineering. In contrast, 16 other sectors—including healthcare, legal, finance, and education—each account for less than 5% of total usage, representing significant untapped opportunities. A key insight is the "trust deficit": while models like Claude are capable of working autonomously for nearly five hours, the 99.9th percentile of user sessions lasts only about 42 minutes. This gap highlights a major product opportunity. Over time, user trust grows—experienced users shifting from pre-approval to proactive monitoring—but overall adoption still lags behind technical capability. The report suggests that vertical AI applications in underserved domains could spawn hundreds of unicorns, mirroring the rise of SaaS. Success requires deep domain expertise, proprietary data integration, context-aware engineering, and effective change management. Regulatory approaches should enable—not hinder—human-AI collaboration by focusing on monitoring and intervention rather than mandatory step-by-step approvals. In summary, the AI agent landscape remains early-stage, with vast potential in verticals where domain-specific agents can automate complex, high-value workflows.

Author: Garry's List

Compiled by: Deep Tide TechFlow

Deep Tide Introduction: Anthropic has released the most comprehensive study to date on the real-world usage of AI Agents. The core data shows: software engineering accounts for nearly 50% of AI Agent tool calls, while 16 vertical domains including healthcare, legal, and education combined account for less than half of the remainder, with each domain's share below 5%.

This is not a sign of market saturation, but a map of 300 vertical AI unicorns—more valuable is a counterintuitive finding cited in the article: models can already work independently for nearly 5 hours, but users only let them work for 42 minutes. This "trust deficit" itself is the next product opportunity.

Full Text Below:

Software engineering accounts for nearly 50% of all AI Agent tool calls. Sixteen domains including healthcare, legal, and finance are almost untouched, each below 5%. This means there are 300 vertical AI unicorns waiting to be built.

If I were to start a business today, I would stare at the red area in the bar chart above until I saw my future.

Box founder Aaron Levie said:

This chart is a great reminder of how much opportunity there is in the AI Agent space right now.

There will certainly be a lot of horizontal Agent opportunities, but there is also a lot of workflow that requires deep domain expertise to truly help users automate the unique processes in their vertical.

The template is: build Agent software that integrates proprietary data to effectively bridge users and Agent collaboration in handling workflows, while possessing deep domain-specific contextual engineering capabilities and the ability to drive change management on the client side.

Many domains still have huge gaps.

Software engineering occupies half of all AI Agent activity. The other half is scattered across 16 vertical domains, none exceeding 9%. Healthcare accounts for 1%, legal for 0.9%, and education for 1.8%. These are not saturated markets; they are markets that barely exist.

Anthropic just released the most comprehensive study to date on real AI Agent usage. The core finding: software engineering accounts for 49.7% of Agent tool calls on its API. The core conclusion buried deeper: everything else is a blue ocean.

Deployment Lag

One data point should excite entrepreneurs: the model's capabilities far exceed the boundaries of what users are willing to trust it with.

METR's capability assessment shows that Claude can solve tasks that would take a human nearly five hours to complete. But in actual use, the 99.9th percentile session duration is only about 42 minutes. This gap—between what AI can do and what we allow it to do—is a huge opportunity.

Figure: The maximum duration Claude Code was trained on nearly doubled in three months. This not only improved capabilities but also enhanced trust.

Source:x.com

From October 2025 to January 2026, the 99.9th percentile single-session duration almost doubled, growing from less than 25 minutes to over 45 minutes. Growth was steady across model versions. This isn't just the model getting stronger; it's users learning through repeated use, gradually extending their trust in the Agent.

"From August to December, Claude Code's success rate on internal users' most challenging tasks doubled, while the number of human interventions per session decreased from 5.4 to 3.3."

The capability is already there; deployment hasn't caught up. This isn't a problem; it's a product opportunity.

How Trust Evolves

20% of new users automatically approve Claude Code's actions. By 750 sessions, over 40% of sessions run in full auto-approval mode. But there's a counterintuitive finding: experienced users intervene more, not less. New users intervene in 5% of turns, while experienced users intervene in 9%.

Figure: Trust is a skill that accumulates. New users automatically approve 20% of sessions. By 750 sessions, this exceeds 40%.

Image: Anthropic

Source: x.com

This isn't a contradiction but a shift in supervision strategy. Beginners approve step-by-step before actions occur; experienced users authorize first and intervene only if problems arise—they've moved from pre-approval to active monitoring.

Here's a safety-relevant finding: on complex tasks, Claude Code proactively requests clarification more than twice as often as humans proactively intervene. The Agent pauses to confirm rather than charging ahead. This is a feature, not a bug.

"The core insight of this study is: the autonomy Agents exercise in practice is co-constructed by the model, the user, and the product. Claude pauses to ask questions when uncertain, thereby limiting its own independence. Users build trust through collaboration with the model and adjust their supervision strategies accordingly."

Levie's Vertical AI Playbook

Aaron Levie points to the immense wealth and value waiting to be unlocked: build Agent software that integrates proprietary data, make it truly solve real people and problems, pack it with context to maximize intelligent output, and—this is the part most entrepreneurs miss—drive change management on the client side.

This last point is why vertical AI is so hard to replicate. Anyone can build an API wrapper, but few can truly navigate the workflows, regulatory constraints, and organizational resistance unique to medical billing, legal discovery, or building permit approvals.

SaaS grew tenfold every decade over the past few decades. Over 40% of venture capital in the past 20 years flowed to SaaS companies. This industry spawned over 170 SaaS unicorns. The logic is simple: each of these unicorns has a vertical AI version waiting to emerge. And the AI version could be ten times larger because it replaces not just software but also operators.

The Nature of Co-Construction

Anthropic's core finding deserves serious attention from anyone involved in AI policy making. Autonomy is not an inherent property of the model but is co-constructed by the model, the user, and the product. Pre-deployment evaluations cannot capture this; you must measure it in real use.

Anthropic stated officially:

Software engineering accounts for about 50% of Agent tool calls on our API, but we are also seeing emergence in other industries. As the boundaries of risk and autonomy continue to expand, post-deployment monitoring becomes critical. We encourage other model developers to expand on this research.

The safety numbers are reassuring: 73% of tool calls have a human in the loop, and only 0.8% of operations are irreversible. The highest-risk deployment scenarios—such as API key exposure or autonomous crypto trading—are mostly security assessments, not real production environments.

"Regulatory requirements that mandate specific interaction patterns—for example, requiring human approval for every action—will only create friction without necessarily delivering safety benefits."

Policies mandating "approve every action" kill productivity gains without increasing safety. A better goal is to ensure humans can monitor and intervene, not to mandate specific approval workflows.

Where the Unicorns Are Hidden

The map is drawn. Software engineering is already being done. Healthcare, legal, finance, education, customer service, logistics—16 vertical domains, each with single-digit market share—are waiting for someone to truly embed domain expertise into Agents.

300 SaaS unicorns were born before; the next 300 vertical AI unicorns are about to emerge. The founders who pick a vertical, embed domain expertise into Agents, and figure out how to drive change management will own the enterprise software market for the next decade.

The model can work for five hours; users only let it work for 42 minutes. That's the signal: we are still in the very early stages, there is so much left to build, and in countless places that haven't seen even a minute of intelligence at work.

Preguntas relacionadas

QWhat percentage of AI Agent tool usage is concentrated in software engineering according to Anthropic's data?

ASoftware engineering accounts for nearly 50% (49.7%) of all AI Agent tool usage.

QWhat is the key opportunity identified in the gap between AI's capabilities and user trust?

AThe gap between AI's ability to work for nearly 5 hours and users only allowing it to work for about 42 minutes represents a major product opportunity to build trust and increase deployment.

QHow did user intervention behavior change as they gained more experience with Claude Code?

AWhile more experienced users (after 750 sessions) ran over 40% of sessions in auto-approval mode, they actually intervened more frequently (9% of turns) compared to new users (5% of turns), shifting their strategy from pre-approval to active monitoring.

QAccording to Aaron Levie, what are the key components for building a successful vertical AI agent?

AThe key components are: building agent software that integrates proprietary data, effectively bridging user and agent collaboration, possessing deep domain-specific contextual engineering capabilities, and driving change management on the customer side.

QWhat does the article suggest is the future market potential for vertical AI compared to SaaS?

AThe article suggests that while the SaaS industry produced over 170 unicorns, there are potentially 300 vertical AI unicorns waiting to be built, and the AI versions could be ten times larger because they replace not just software but also the operators.

Lecturas Relacionadas

STAR 50 Soars 10.73%, Why Did A-Shares Stage a "V-Shaped Reversal"?

After a prolonged decline, the Chinese A-share market staged a strong rally on July 21. The STAR 50 index surged 10.73%, its largest single-day gain in nearly a year, leading a broad-based "V-shaped" reversal. The Shanghai Composite Index rose 1.79%, the Shenzhen Component Index gained 4.81%, and the ChiNext Index jumped 7.05%. Total market turnover reached 2.97 trillion yuan, an increase of 256.1 billion yuan from the previous session, with over 3,100 stocks advancing. The semiconductor sector spearheaded the rebound, with related ETFs posting significant gains. Analysts attribute the surge to three converging factors. First, coordinated capital inflows from "national team" institutions, insurance funds, listed company buybacks, and fund house self-purchases have bolstered market liquidity and confidence. Second, supportive policy signals, including commitments from regulators to ensure stable market operations, provided a favorable backdrop. Third, a stabilization and recovery in overseas markets, notably South Korea, created a positive external environment. Institutions suggest the most severe panic selling phase for the tech sector has likely passed, following a significant digestion of crowded positions and leveraged funds. While short-term volatility may persist, the medium to long-term outlook remains underpinned by enduring trends like AI computing demand expansion and semiconductor localization. The market's focus now shifts to the sustainability of supportive fund flows, earnings reports, and upcoming catalysts from the global AI industry chain.

marsbitHace 39 min(s)

STAR 50 Soars 10.73%, Why Did A-Shares Stage a "V-Shaped Reversal"?

marsbitHace 39 min(s)

U.S. Tech Momentum Stocks Post Largest Single-Day Gain Ever, But Is the Plunge Over?

US tech momentum stocks staged a sharp rebound on Tuesday (July 21st). Morgan Stanley's TMT Momentum Factor surged over 12%, marking its largest single-day gain on record, exceeding even peaks from the 2000 dot-com bubble. Key momentum indices from Goldman Sachs also posted their strongest daily performances in years. The rally was led by semiconductors, with the Philadelphia Semiconductor Index jumping 4.6%. This rebound followed three consecutive down days and a cumulative 33% plunge in momentum stocks, one of the steepest drawdowns since the dot-com era. Analysts attribute the surge largely to a short squeeze. Heavy selling had pushed high-beta momentum stocks into deeply oversold territory, forcing many short sellers, particularly in Asia, to cover their positions, creating a self-reinforcing buying spiral. However, the rebound's internals appear weak. Trading volume was notably low, and advancing stocks still lagged decliners on the S&P 500, indicating a narrow, concentrated rally rather than broad market participation. Diverging views emerge on the outlook. BTIG warns the bounce has hit key resistance and recommends selling into strength, citing extreme volatility and historical parallels to past market tops. Conversely, Goldman Sachs and UBS believe the momentum unwind is nearing its end, suggesting it may be time to gradually add exposure, as positioning has been significantly reduced. They caution, however, that high volatility warrants a measured approach, potentially using defined-risk strategies. The upcoming earnings season, particularly reports from major tech firms like Alphabet, is seen as a critical test for the rally's sustainability. Simultaneously, bond markets flashed a warning, with yields rising partly due to spiking oil prices. Analysts note that if long-term Treasury yields break decisively higher, it could pose a significant headwind for equities, especially growth stocks.

marsbitHace 47 min(s)

U.S. Tech Momentum Stocks Post Largest Single-Day Gain Ever, But Is the Plunge Over?

marsbitHace 47 min(s)

U.S. Tech Momentum Stocks Record Largest Single-Day Gain Ever, but Has the Rout Ended?

U.S. tech momentum stocks staged a dramatic rebound on Tuesday, July 21st. Key momentum indices like the Morgan Stanley TMT Momentum Factor and Goldman Sachs' High Beta Momentum Long Index posted historic or near-historic single-day gains, fueled largely by semiconductor stocks. This sharp rally followed a severe three-day sell-off that saw momentum stocks plunge 33%, marking one of the steepest pullbacks since the dot-com bubble. Analysts attribute the bounce primarily to a short squeeze, as forced covering from over-leveraged traders, particularly in Asia, created a buying spiral. However, the rally's health is questioned due to weak market breadth—overall trading volume was low, and decliners outnumbered advancers in the S&P 500 despite the index's gain—suggesting a narrow, concentrated surge rather than broad recovery. Opinions on the sustainability diverge. BTIG strategists warn the rebound has hit key resistance levels, citing extreme volatility and historic stock dispersion as signs of an ongoing broader correction, and recommend selling into strength. Conversely, Goldman Sachs and UBS view the aggressive momentum unwinding as nearing its end, noting reduced positioning and a lack of new fundamental catalysts. They suggest the sell-off presents a selective opportunity to add exposure, albeit cautiously and gradually using defined-risk strategies. The immediate trajectory hinges on the ongoing earnings season, with market focus on Alphabet's capital expenditure guidance for AI investment clarity. Meanwhile, bond markets present a risk, with rising Treasury yields—potentially heading toward 5.5%—and widening credit spreads for mega-cap tech companies posing a threat to equity valuations. The combination of technical factors, earnings results, and macro conditions leaves the durability of the rebound in doubt.

链捕手Hace 49 min(s)

U.S. Tech Momentum Stocks Record Largest Single-Day Gain Ever, but Has the Rout Ended?

链捕手Hace 49 min(s)

Long-Divided Must Unite, Long-United Must Divide: When L1 Becomes Its Own Rollup, What Is Ethereum's Endgame?

"The Inevitable Cycle: When L1 Becomes Its Own Rollup – What is Ethereum's Endgame?" For years, the Ethereum community grappled with concerns that L2s were fragmenting the ecosystem and eroding L1's value. While L2s provided cheaper execution, they also splintered liquidity and the unified user experience of a single chain. This has prompted a fundamental reassessment of the relationship between L1 and L2. Ethereum's roadmap is evolving. The "Scale" initiative merges L1 and L2 expansion into a holistic framework. L1 itself is advancing with higher gas limits, statelessness, and zkEVM verification, no longer content to be just a low-throughput settlement layer. Consequently, the primary value proposition of L2s is shifting from merely providing cheap blockspace to offering L1 cannot easily provide: application-specific optimizations, privacy features, and flexible governance models. L2s are becoming a spectrum of execution environments with varying degrees of security inheritance from Ethereum. A critical challenge in this multi-chain future is interoperability. The vision is to make Ethereum "feel like one chain again." This relies on advancements in native account abstraction (like EIP-7702) and intent-based architectures (Open Intents Framework), where users declare desired outcomes, and solvers handle the complex cross-chain execution. Furthermore, shortening Ethereum's finality time from minutes to seconds is crucial, as it underpins trust between chains for bridges, stablecoins, and cross-chain applications. Perhaps the most provocative idea is that Ethereum L1 itself could become a form of "its own Rollup." As zkEVM and proof systems mature, high-performance nodes could execute transactions and generate validity proofs. Regular validators would then verify these proofs instead of re-executing all transactions. This blurs the traditional L1/L2 hierarchy, making "Rollup" more of a general execution-verification architecture. Native Rollup aims to integrate L2 validation more directly into the Ethereum protocol, allowing L2s to inherit L1's security more fully and move away from reliance on security councils. In the end, L2s are not destined to replace L1 or be made obsolete by it. The likely future is a unified system where diverse execution environments—each optimized for specific use cases like DeFi, gaming, or privacy—coexist. They will share a common foundation of security, liquidity, and verifiable state, seamlessly connected to restore a cohesive user experience. The next phase for Ethereum is not just about scaling through separation, but about intelligently reintegrating what was separated back into a coherent whole.

链捕手Hace 1 hora(s)

Long-Divided Must Unite, Long-United Must Divide: When L1 Becomes Its Own Rollup, What Is Ethereum's Endgame?

链捕手Hace 1 hora(s)

Trading

Spot
活动图片