AI Investors' 2026 Anxiety: When Models Devour Everything, What Moat Is Left for Startups?

marsbitPublished on 2026-06-11Last updated on 2026-06-11

Abstract

In 2026, a wave of investor anxiety questions the defensibility of AI startups as models improve, fearing that most companies are just "thin wrappers" destined to be absorbed by foundation models or chipmakers. The author argues against this despair, positing that true moats lie not in benchmark performance but in areas models cannot easily reach. The logic of despair is that if models excel at all measurable tasks, only compute and cutting-edge model weights hold lasting value. However, the essay contends that the most valuable work is inherently "untrainable." Benchmarks measure what can be measured and thus optimized for, but real-world correctness often resides in private, complex systems. Examples include legacy codebases, intricate legal transactions, or hospital workflows. This kind of correctness is proprietary, costly to establish, and cannot be validated quickly—it requires time and trust within an organization. As models commodify visible, measurable tasks from both above (labs absorbing scaffolding) and below (saturation by cheaper models), value shifts to "untrainable ground." This encompasses work where correctness is a private truth, locked behind integration barriers, licenses, liability frameworks, and entrenched user habits. Trust and adoption are slow, human-centric processes that smarter models cannot accelerate. Successful companies defend their position by embedding deeply into client operations, owning the definition of "good" within a specific domai...

Author: Sarah Guo

Translation: Deep Tide TechFlow

Deep Tide Introduction: When large models begin to crush humans on all leaderboards, investors are falling into a kind of despair: what is worth investing in besides Anthropic and Nvidia? This top Silicon Valley investor explains with data and case studies that the real moat isn't on the leaderboard—it's hidden in places that cannot be measured by benchmarks.

Mid-2026, the investor version of AI psychosis is despair: There's nothing worth investing in, we should just put all our money in Anthropic and Nvidia and go home.

I've never felt this way. I'm already convinced the model is several minor versions smarter than me, I'm happy to buy Anthropic and Nvidia at market price, and all my smartest friends are fairly convinced self-improvement will succeed soon—but I still don't feel this despair.

This despair is not stupid. The logic is this: if models keep getting better at everything, then every company built on them is just a thin layer of wrapping, waiting to be absorbed, and the only value that survives is compute power and frontier weights.

Take software as an example, the case despair theorists rely on most. When Devin launched in 2024, it could only solve 13% of tasks on standard software benchmarks, basically ignored. A year and a half later, the best agents score in the 80s, and they're doing real work inside Goldman Sachs and the U.S. Army. Almost everyone draws the same wrong lesson: the model ate software engineering. But as the model devours the most measurable parts of software engineering, we're rediscovering what many teams have long known—engineering has always resisted measurement, and the easiest part to measure may not be the only important part.

MIT's Mert Demirer and his collaborators finally put numbers to it: among over 100,000 developers, the latest coding agent increased the amount of code written by about 180%, while the amount of code actually shipped increased by about 30%. Writing code got cheaper. The remaining parts still have to go through people, and they matter. Of course, the net impact is still staggering.

Benchmarks are things you can measure, and what you can measure is what you can train for. Therefore, coding agents matured first: compilers are free verifiers, test suites are free verifiers, when the answer checks itself for free, you can grind against the check until you beat it. But passing the tests never tells you whether that change is the right one for a decade-old codebase with three undocumented modules whose reason for existing, and a deployment pipeline held together by a cron job nobody will admit to writing.

That kind of correctness cannot be read from a leaderboard, and in fact cannot be read from anything. You learn by running in the real world long enough to discover whether such a complex system works, and smarter models don't make the world run faster. Nobody runs unit tests on something Google-scale and trusts the green check; you trust it because it has weathered years of real load. Such correctness isn't just private, it's the slow kind of moat that capital can't bulldoze. Even optimists admit the clock cannot be skipped: Noam Brown, pioneer of reasoning models at OpenAI, recently wrote that the only reliable way to evaluate an agent over a year-long timescale might just be... to run it for a year.

As Gabe Pereyra says, true automation is not just the model getting better. It's the product, model, workflow, and company all moving together, and three of those four move at the speed of organizations.

The people-moving part is what benchmarks don't touch: getting a skeptical partner to change how she handles matters, keeping the team together through the rebuild. That's why when we hire a CEO, the ability to handle people matters at least as much as analytical skill, and smarter models won't change that weighting. Feedback is fuzzy, the timescale is years, trust is personal. Every company I know gave every engineer access to frontier coding models, but none of them changed their engineering org anywhere near that speed. Adoption took a quarter, what a magical token growth quarter that was! But the rebuild is taking years.

What's visible is what's leaving. Valuable work is structurally invisible: anything you can put on a leaderboard, you can train against, so anything measurable is already on its way to commoditization. This process takes time and never fully completes, but the direction never reverses. In the monetary terms of my friend Matt MacInnis at Rippling: tokens spent answering generic questions are almost worthless because anyone's model can answer it, while tokens spent reasoning over your company's data are worth much more because they do what you actually want, not just what looks plausible.

Visible work gets eaten from two sides. From below, task saturation: once a job can be cheaply checked, buyers stop asking which model did it and start asking how much it costs, and the job falls to the cheapest open-source or distilled model that week. Wherever they can have impact, margins matter in the end. From above, labs are trying to get the model to devour its own scaffolding. Retrieval, routing between cheap and expensive calls, tool use, even reasoning strategies, all the apparatus that used to wrap the model gets pulled into the weights until the wrapper is the model. That's frontier absorption. Margin pressure cuts the other way too: a general-purpose agent has to be ready for anything, which is expensive, while a focused app can tune a workflow until it runs on a fraction of the token spend, and unlike the lab selling those tokens, it keeps the spread.

So, we can ask two things about any type of work. Is its correctness private and expensive to build, the kind of truth that exists only inside someone's data? Is it isolated, locked inside systems you cannot enter? Contrast these against how saturated the task is, and you get a 2x2 matrix. Saturated work with public answers is commodity tokens, owned by open-source models. Frontier work with public answers, where coding benchmarks live, is where labs win, because when evaluation is free, owning it costs nothing. The prize is in the last corner, the untrainable one: frontier work whose correctness exists only in private domains. You can see this in the inference clouds hosting AI-native pioneers, where the vast majority of tokens are generated by bespoke models, not general-purpose open-source ones.

The walls into that last corner vary in height. A single developer's toy codebase is portable and standardized, so the climb is short. A bank's production system is neither, and you don't get root access by being 2% smarter on SWE-Bench Verified.

Capability eats many things, but a better model does not make a private ground truth public. It doesn't hold the license, sign the liability, or own the firm's documents, and it cannot be the party sued when the answer is wrong. Intelligence isn't the bottleneck here. Licensing is, and so is liability. You can imagine a model far smarter than anyone, and it still must be allowed through the door, and someone still must sign their name to what it does.

That door has a lock and a bolt. The lock is context: you only get to validate whether the AI did useful things after being trusted inside the system, after security reviews, integration, the contract where you sign your name to the result. The bolt is the user. Most doctors in the U.S. now open OpenEvidence every day, and no amount of compute can buy that. A lab could train a perfect medical model tomorrow and still fail to get into a doctor's habit, or into UCSF's decision flow, because trust is built slowly, on relationships, requiring the user's acquiescence, not erasing it with gradient descent.

This too is work. An app earns its place in the untrainable corner by doing unglamorous work: arranging the company's private reality so the model can act on it, giving the model tools to act, working with the customer to change the reality of their employees. A company that brings the translation is hard to copy—and the translation never ends. Integration and maintenance last as long as the relationship, won by teams that put domain-specialist engineers and tools next to the customer.

For example, at a top white-shoe law firm, the M&A practice alone runs nearly a thousand deals a year. For confidentiality and many other reasons, you can't have hundreds of associates each downloading client files to their desktop and asking a general-purpose agent to sift through them, and even if you could, what you'd learn would be piecemeal, one associate's correction at a time, missing how the whole deal flows. The important signal exists at the deal level, and a deal has a shape: for M&A, it's NDA, term sheet, due diligence, purchase agreement, ancillary documents, closing checklist; for IP litigation, it's motions, discovery, prior art, more motions. Each practice area has its own, and lawyers and tools are not interchangeable across them. And the problem the law firm actually solves sits a level above all this: running every practice area in parallel, like a top partner running hundreds of matters at once while onboarding new ones and training associates. Transforming such a law firm is not a single task you can write an eval for. It requires an operator to do what the analytics firm does, with incredibly fuzzy goals, incomplete feedback, long timescales, in an environment that won't stay still.

Unfortunately, invisible value is also hard to sell, for the same reason it's hard to commoditize: companies can't tell from the outside whether AI will transform their ops, just as a benchmark can't tell. So the strongest enterprises stop trying to prove it from the outside and go inside, pricing on outcome. Sierra charges when its agent resolves a customer issue, not when it kicks it to a human, so price becomes the evaluation, which only works because Sierra owns the definition of "resolved." Cognition's Devin does the same move in software, offering a "performance guarantee," which you can only give for outcomes in systems you are trusted inside.

Even serving tokens, the layer everyone loves to call a pure commodity, doesn't act like one. The best AI-native companies concentrate their serving on one or two providers (Baseten or Fireworks) because per-token cost commoditizes on schedule, while reliability at real traffic and guaranteed access to scarce compute do not. Where you serve is a separate choice from which models you use. Price is the only part of inference that acts like a commodity.

A common objection raised is that the lab is your supplier—why wouldn't it run its own first-party product below cost to bleed you dry, or revoke your API access and take the market itself? This is the real version of despair theory, and it only works if the model layer is a single-player game. It's obviously not—it looks more like a deathmatch between three and a half parties, with a pack of international players six months behind on training, and a G League five times the size of last year's. Customers want competition among suppliers, and labs want market share more than they want any single app dead.

You can see this in markets where labs go head-to-head. In consumer chat, the best model never simply wins. ChatGPT held the lead for years in real competition, and the share it's losing now is going to Gemini, on the strength of Android and search, not a better model. Anthropic, currently rated by prediction markets (and internet vibes) as having the best model, is barely a factor in consumer chat but built its business in enterprise and coding instead. If a better model cannot take users from a competitor in the most core application, it won't make it through a hospital's records or a bank's liability by integration. The public's choice today is not based on coding alone. If the frontier stays crowded, the layer above it will be valuable.

If work can't be scored from the outside, someone inside has to decide what even counts as a good answer, and that decision is the whole game. Enough of those decisions, written down, becomes a benchmark. Harvey released one for law, Sierra for voice agents. You earn the right to define what good means for a domain by being the one that domain already uses, and these companies won that right through the fight of real adoption.

The evaluations that decide real money are private and vary by company: this firm, on this matter, will accept what as good work, and it's nowhere near done because the depth of law dwarfs any public test. OpenEvidence is establishing what a safe clinical answer looks like. These are not really measurement, this is judgment about what's true and what's good, written down until it becomes the standard by which everyone else is measured, and the underlying lab, however smart, cannot write it because that status exists only inside the domain. That authority tends to land where it already sits. Senior lawyers write the law benchmark. Defining safe clinical answers falls to doctors. And what resolved means is whatever the company that already has the customer says it means.

The frontier of absorption keeps rising because we keep learning to measure more work, and the measurable gets eaten. The untrainable ground shrinks under the feet of whoever stands on it, so you cannot find a defensible point and rest. You keep moving toward whatever still can't be scored, you keep re-underwriting. On a narrow task, with your private data and your own evaluations, you can fine-tune to the frontier and beat the general-purpose model where it matters, and that specialized model becomes part of the moat. On the other hand, competing on the general-purpose model is a capital war you lose to whoever has the most compute, the trap for companies with shallow access and visible tasks. The day it promises survival by out-training the frontier on general tasks, the winner seems most determined by datacenter scale, and the end is usually not an independent champion but a sale to someone compute-rich.

All this is defense. The harder part is offense, choosing what to build in the first place. This is what I spent a year looking for, and I might have found it three times. The model doesn't help here. It will do whatever you point it at, but can't tell you what's worth pointing at, you cannot benchmark that, so you cannot train it. This is also why incumbents won't take everything: they hold the ground they have, and the next thing comes from whoever spots a use before the rest of us. Perhaps intention is a scarcer input than compute.

The despair theory is half right. The thin wrapper layers are indeed being absorbed, and a lot of what looks like a company today is a thin wrapper. It's wrong about what's left. The mechanism is clear; the destination is not. What I'd bet on is the direction: intelligence keeps getting cheaper, and value keeps sliding toward the few places the model cannot reach. The untrainable is value with history. So get into one, do the unglamorous translation, start writing down what good means there, because someone will. This year's most cited benchmark score is a map of territory about to become worthless, and a notice of who's about to lose the right to say what counts as good.

Related Questions

QWhat is the main anxiety described among AI investors by 2026 according to the article?

AThe main anxiety is a feeling of despair that there is nothing left to invest in except for the leading model providers like Anthropic and hardware leaders like Nvidia, as models seemingly commoditize and absorb all value built on top of them, leaving no defensible moat for startups.

QWhy does the author argue that benchmarks are misleading indicators of a company's defensibility?

AThe author argues that benchmarks measure only what is publicly measurable and trainable. Therefore, any task that can be benchmarked is on a path to commoditization. True defensibility lies in 'untrainable ground'—private, hard-to-measure work involving integration, domain-specific knowledge, trust, and organizational change, which cannot be captured by a public score.

QWhat two factors create a 'wall' protecting valuable, 'untrainable' work from being absorbed by general AI models?

AThe two factors are: 1) The 'lock' of context—gaining trusted access to a private system requires security reviews, integrations, and contracts, which is a slow process. 2) The 'latch' of the user—establishing user habits and trust within an organization (like doctors using a specific tool) is based on relationships and slow adoption, not just superior model intelligence.

QHow do leading AI-native companies like Sierra and Cognition change their business models to align with the concept of 'untrainable' value?

AThey shift from selling based on inputs (like tokens) to pricing based on outcomes and guarantees. For example, Sierra charges only when its agent solves a customer's problem, and Cognition's Devin offers performance guarantees for software tasks. This is only possible because these companies have earned the trust to define what 'solved' or 'good' means within a specific client context.

QWhat is the author's final investment thesis or recommended direction in the face of AI models becoming universally capable?

AThe author advises betting on moving into areas of 'untrainable' value—where correctness is private, expensive to establish, and isolated within specific systems or organizations. The strategy is to do the 'unremarkable work of translation': integrate deeply into a domain, start defining what 'good' means within that private context, and build a moat based on relationships, trust, and proprietary workflow orchestration that models cannot easily replicate.

Related Reads

Only 153 Venture Capital Firms Invested in July: Is the Crypto VC Industry Experiencing a 'Mass Extinction'?

In July 2026, only 153 unique venture capital firms participated in disclosed crypto funding rounds, marking the lowest monthly count since November 2020. This figure represents an 87% decline from the peak of 1,177 firms in 2022. Overall, the first seven months of 2026 saw crypto projects raise approximately $11.78 billion across 481 rounds. This crypto VC contraction contrasts sharply with the broader venture capital landscape, where global VC investment reached a record $560.4 billion in H1 2026, heavily fueled by major AI company financings. This shift in capital allocation has drawn funds away from the crypto sector. Within crypto, funding is highly concentrated. Trading platforms, prediction markets, and payment sectors absorbed 53% of the total capital. While early-stage deals remain frequent, the largest sums flow to a few late-stage rounds and mergers & acquisitions, which surged to $7.23 billion in Q2 2026. The market is consolidating around top funds like a16z crypto and Dragonfly, which successfully raised new multi-billion dollar funds, while many smaller firms have retreated. Analysts describe this as a "great extinction" for crypto VCs, where capital is becoming more selective, favoring proven business models and assets over early-stage speculation. This raises the bar for project quality, funding efficiency, and viable exit paths.

marsbit20m ago

Only 153 Venture Capital Firms Invested in July: Is the Crypto VC Industry Experiencing a 'Mass Extinction'?

marsbit20m ago

Strategy's Loss in the Second Quarter Reaches $8.22 Billion Amid Bitcoin Decline

Strategy, the largest corporate holder of Bitcoin, reported a net loss of $8.22 billion for the second quarter. This loss was primarily driven by an $8.32 billion unrealized loss on its Bitcoin holdings due to a decline in the asset's price during the period. Despite these paper losses, the company increased its Bitcoin holdings to 843,775 BTC, a 25% growth since the start of the year. As part of a new monetization strategy, Strategy sold approximately $218.4 million worth of Bitcoin, mainly to fund dividends for preferred shareholders, with $216 million of that sold after Q2 ended. The company also built a $3.75 billion cash reserve, which it claims is sufficient to cover over two years of dividend and interest payments, aiming to insulate itself from Bitcoin's volatility while meeting obligations. Following the earnings release, Strategy's stock (MSTR) rose 4.7% in regular trading but corrected slightly after-hours. This pattern reflects how the company's accounting results are heavily tied to Bitcoin's price swings, even as its long-term strategy remains unchanged. The report indicates that Strategy is maintaining its core strategy of accumulating Bitcoin while building a financial buffer. This quarterly loss follows a recognizable pattern, with the company posting significant unrealized losses in previous quarters (e.g., $12.4 billion in Q4 2025 and ~$12.5 billion in Q1 2026) due to fair-value accounting. A key technical shift is its new monetization program, which introduces periodic selling pressure on the market, transitioning Strategy from a pure accumulator to a participant that occasionally adds supply. A critical question remains: how long can the cash reserve cover dividend obligations if a Bitcoin price downturn persists beyond two years?

cryptonews.ru40m ago

Strategy's Loss in the Second Quarter Reaches $8.22 Billion Amid Bitcoin Decline

cryptonews.ru40m ago

Will Terrorist Durov Ban Russian Officials?

Telegram founder Pavel Durov publicly reacted to being labeled a "terrorist" by Russian authorities, stating the designation came after he refused demands for mass surveillance and censorship on the platform. In a Telegram post, he highlighted that this status formally bans him from "publishing information online." Durov concluded with a statement widely circulated: Russian officials "clearly don't understand who can ban whom on the internet." This remark suggests Durov could potentially restrict official Russian government and officials' channels on Telegram, which continue to operate on the platform despite its formal blocking in Russia. The situation parallels previous, slow-moving state directives, like switching officials to domestic cars, contrasted with the current push to migrate all government communication to the Russian-made messenger MAX by 2030. However, reports indicate many officials still use Telegram via workarounds, fearing surveillance on MAX, while alternatives like BiP and KakaoTalk recently became inaccessible in Russia without a VPN. Durov has not specified any immediate actions against state channels. His statement is an initial response, with further developments depending on the authorities' reaction. The dynamic differs from 2020 when Russian regulators lifted a block on Telegram; now, Durov implies control from within the platform itself over the official accounts that persisted through that earlier blockade.

cryptonews.ru40m ago

Will Terrorist Durov Ban Russian Officials?

cryptonews.ru40m ago

DeepSeek V4 Official Version Arrives, New Capabilities Emerge, Value-for-Money King Enters the Fray

On July 31st, DeepSeek officially launched the public API beta for its DeepSeek-V4-Flash model. A key highlight is its performance on multiple Agent benchmark tests, reportedly nearing or even surpassing the level of the V4-Pro preview version from three months ago. Notably, the Flash model achieves this with significantly smaller scale (130B active parameters vs. Pro's 490B), suggesting that post-training optimization and data quality may be as crucial as raw model size. DeepSeek emphasized that the V4-Flash-0731 uses the same model architecture and size as its preview version, with improvements attributed solely to "re-trained post-training." The update also marks the official debut of DeepSeek's self-developed Agent framework, "Harness." The move signals DeepSeek's strategic push to position its cost-effective Flash model as a competitive base for Agent applications—scenarios requiring autonomous planning, tool usage, and complex task execution—where inference speed and cost are critical. By natively supporting OpenAI's Responses API format and adapting for code-generation scenarios, DeepSeek aims not just to be a cheaper alternative but to establish its own ecosystem in the Agent era. This release follows DeepSeek's record-breaking ~$50 billion fundraising round roughly two months prior, underscoring market confidence in its technology and commercialization prospects. The company is reportedly preparing for another funding round at a valuation of approximately $71 billion. The Flash model's advancement represents a step in fulfilling the high expectations that come with this valuation, setting the stage for the impending release of the V4-Pro official version and intensifying competition in the global Agent landscape.

marsbit44m ago

DeepSeek V4 Official Version Arrives, New Capabilities Emerge, Value-for-Money King Enters the Fray

marsbit44m ago

Trading

Spot
活动图片