Dragonfly Partner Haseeb: The Fastest-Growing Companies of the Future May All Get Stuck at 149 Employees

marsbitPublié le 2026-06-24Dernière mise à jour le 2026-06-24

Résumé

Dragonfly partner Haseeb explores the distorted economics of AI model pricing, drawing parallels to tax policy. He notes that startups and small teams (under 150 users) enjoy heavily subsidized, fixed-price AI subscriptions (like Claude Code), where the marginal cost of an additional token is effectively zero. This creates a powerful incentive for them to maximize token usage ("token-maxxing") and innovate aggressively with AI automation. In contrast, large enterprises (over 150 users) are forced onto "Enterprise" plans, paying per-token API fees with high (~75%) markups. This acts like a steep "tax" on AI-powered labor, disincentivizing marginal automation and experimental use, and encouraging them to retain more human workers. Haseeb argues this pricing creates a "150-person cliff," a regulatory notch similar to labor laws in France that discourage firms from growing past 50 employees. He predicts the fastest-growing future companies may deliberately cap their headcount at 149 to avoid the punitive enterprise pricing. This would foster an "AI-first" management philosophy obsessed with automation and outsourcing to stay lean. While not intentionally designed, this bifurcated pricing could become one of the most influential de facto tax policies, shaping how AI replaces labor—not through mass layoffs at big firms, but through agile, AI-native startups outcompeting them.

Author: Haseeb

Compiled by: Jiahuan, ChainCatcher

@SemiAnalysis_ recently discovered an incredible phenomenon in the economics of AI programming subscriptions. If you max out usage, the fee you pay is actually 20 to 70 times cheaper than buying tokens through the API.

Many see this and say: My God, look at how much these LLM companies are subsidizing tokens. The bubble is bound to burst soon.

This reaction is wrong. LLM companies are willing to offer such generous packages precisely because most users rarely hit the ceiling. The product is like a gym membership: generous allowances exist because the vast majority of people hardly use them.

But I've spent a long time pondering this; there is something odd here.

We don't know their actual blended margins on subscriptions, but according to SemiAnalysis, at 20% average utilization, Anthropic's Max 5x plan barely breaks even. A 20% utilization rate might even be optimistic, especially in organizations where everyone (including non-programmers) has a subscription but only uses it occasionally. Most institutions I know, including Dragonfly, are generous with Claude Code subscriptions and encourage non-technical staff to try them.

But what SemiAnalysis didn't delve into is that this is purely a small business phenomenon. Large enterprises cannot use this subscription pricing.

Here's why: When you reach 150 people or more, you are forced out of the "Team" subscription model. You must switch to "Enterprise," priced at a base of $20 per seat plus API fees based on actual token usage. Enterprises pay linearly for token costs, and SemiAnalysis estimates API token gross margins to be around 75%. This is a massive price hike that kicks in suddenly at 150 people.

So, if you are a small business or startup (or an individual), your perception of AI spending is distorted. Your token pricing is actually heavily subsidized; Anthropic likely maintains extremely low or even negative margins on you.

You might wonder why Microsoft and Uber fret so much about token spend and talk about "token-mining." That's the reason. Their structural cost per token is much higher than for startups and individuals.

But Anthropic doesn't care! For a B2B company, extracting maximum value from small companies or individuals isn't very meaningful. Look at companies like Datadog or Cloudflare; 80% to 90% of their revenue comes from large contracts (annual recurring revenue over $100,000). Making zero profit on the long tail is just a customer acquisition cost.

This is classic B2B sales thinking.

But there's another way to view the same situation: through the lens of tax policy.

Because if tokens are replacing labor, then the gross margin OpenAI and Anthropic charge on tokens is essentially a tax on AI labor.

Viewing token pricing this way leads to two major consequences.

Token Pricing as Tax Policy

Assuming the profit margins in the SemiAnalysis article hold: subscriptions break even, large enterprise API gross margin is 75%. The initial reaction is to call this a 75% AI labor tax on large organizations, and a 0% tax on startups.

Standard tax analysis would say this discourages large companies from using AI labor internally, pushing them on the margin to reduce automation and retain more human labor. (Obviously, it also encourages using smaller or open-source models, but the net effect is both are incentivized. Remember, we're talking about the margin here.)

However, what drives behavior more strongly is not the average tax rate. In tax policy, it never is. What we really care about is the marginal tax rate.

For startups on flat-rate subscriptions, the marginal price of the next token is zero until they hit the cap. And a zero marginal price creates the maximum possible distortion a policy can create.

For startups, the subscription model is basically an innovation subsidy. The overwhelming incentive is to figure out how to spend the entire token budget as efficiently as possible. This means running Ralph loops, filling screens with Claude Code sessions, scheduling swarms of agents to work together.

Until the cap is hit, exploration is free. So startups are essentially racing to squeeze the last drop of value from their subscription, out-producing each other. Paradoxically, the more you use, the lower the average token price. Every startup wants to be the one making Anthropic lose the most on subscriptions.

Large enterprises face the opposite incentive. If you exceed 150 seats, every exploratory token is charged at full mark-up (plus a 75% surcharge!). So every step they take exploring the frontier is linearly punished.

Big companies will still automate large, obvious bulk tasks. But marginal, experimental, risky automation will never be discovered because the cost of discovery is too high. This tax structure ultimately encourages them to retain more human labor, preserving their overall organizational structure.

This is the opposite of Japan. Due to a declining population, Japan faces a huge labor shortage. Historically, this meant Japan pursued intense automation, as high human costs incentivized it. That's why Japan has robots in restaurants, factories, hotels, hospitals.

But, strangely, large enterprises find themselves in the opposite dilemma of Japan: if they have to pay an extremely high tax to use AI, it weakens the incentive to automate, strengthening the motivation to retain existing employees (even more so if wages stagnate during this period).

So where does labor substitution flow in this model?

Everyone is staring at big companies, waiting for AI layoffs. But with a 75% tax, aggressively replacing your own employees with AI may simply be uneconomical; token budgets would explode.

But this doesn't mean substitution won't happen; it just manifests differently.

When big firms lose market share to AI-native startups with minimal blended human labor costs, the big firms' revenue and stock price declines trigger layoffs. But those eliminated jobs never reappear at the winning startups. The net reduction effect is the same; this employment gap is just transferred to a lower-taxed part of the economy.

This is also why "AI-washing" (portraying ordinary layoffs as newfound AI efficiency) may not be a flash in the pan. AI-washing is when a company attributes layoffs to AI efficiency, but is actually just masking ordinary business weakness.

Many think this is just a blip in the current AI hype cycle. But even though everyone is primed to witness big companies doing real AI layoffs, "replacing jobs" with AI, that may never happen at scale.

Labor substitution might unfold a different way: startups beat incumbents, incumbents disguise their decline with AI-washing all the way to death, and the startups never rebuild those old jobs. Job substitution still happens, just not where everyone is looking.

That's the first consequence of this model. But there's a second, even weirder consequence.

The 150-Person Cliff

A regulatory notch is a regulatory boundary that induces huge behavioral jumps. Example: the 30-hour-a-week full-time employment threshold, which created tons of jobs that exactly clocked 29 hours a week.

France famously has extremely rigid labor laws that kick in once a company hits 50 employees (works councils, mandatory profit-sharing, firing protections). Smaller companies are exempt. This gives employers a huge incentive to desperately stay below 50 people.

Source: Garicano, Luis, Claire Lelarge, and John Van Reenen, 2016, Firm Size Distortions and the Productivity Distribution: Evidence from France.

Extend this analogy to AI. LLM companies have established a tax threshold that punishes companies exceeding 150 seats. This means you must stay small to keep that wonderful subsidized subscription price, taxing tokens at ~0% (or even negative) rather than 75%.

This could spawn an entirely new management philosophy for companies. Startups will become increasingly obsessed with solving everything with agents, smaller teams, more frequent layoffs, more outsourcing, doing everything possible to keep human-touch points to an absolute minimum.

Not because it's the "optimal" level of automation, but because the incentives force them there. If the magic number is 149, then every seat is precious; you can't afford to waste a single person outside the core joints of the company.

This cliff might be touted by Harvard Business School types as "the new wave of AI-first management." But properly understood, it's just a rational response to enterprise pricing schemes.

This might sound exaggerated. But you can already see the behavioral divergence across organizations. Talk to developers at big companies; they are carefully counting tokens, increasingly anxious about leadership cutting token budgets. Meanwhile, developers at startups are furiously maxing out usage (tokenmaxxing), launching swarms of agents overnight, and checking logs in the morning. I expect this trend to accelerate.

No one designed this intentionally. No committee decided to subsidize innovation for startups and tax incumbents. It all follows directly from tried-and-true traditional enterprise pricing strategies.

But this is how tax laws have always been: a bunch of ancillary rules that ultimately determine which companies can be built, and how those companies distort themselves to minimize their tax burden.

You might argue this is temporary; LLM companies will eventually meter everyone. GitHub Copilot already made this shift. Maybe, maybe not. But before pricing normalizes, 149-person companies and this new wave of AI-first management may have already exploded, gobbled up market share, and written the playbook for the next generation of startups.

Tax policy is crucial. The entire "gig economy" concept only exists because of the legal line between W-2 and 1099. As more and more labor is eaten by AI, token pricing may be the most impactful tax policy of the next decade. Yet no one will ever vote on it.

(Don't be surprised if the fastest-growing companies of the next cycle are all conspicuously stuck at 149 seats.)

Questions liées

QWhat is the core observation about AI pricing models that Haseeb discusses in the article?

AHaseeb discusses how AI model companies (like Anthropic) have tiered pricing that creates a significant 'cliff effect.' Small startups and individuals can access a generous 'Team' subscription model with a fixed price, effectively subsidizing their token usage. However, once an organization grows beyond approximately 150 seats, it is forced onto an 'Enterprise' pricing model where tokens are billed at full cost with high (75% estimated) margins. This acts like a tax, creating vastly different economic incentives for small companies versus large enterprises.

QHow does the 'Team' subscription model affect the behavior of startups, according to the article?

AThe 'Team' subscription model offers a fixed price for a large token allowance, making the marginal cost of an additional token effectively zero until the cap is reached. This creates a powerful incentive for startups to 'tokenmaxx'—meaning they are encouraged to experiment freely, use as many tokens as possible, and aggressively automate tasks to extract maximum value from their subscription. It acts as a subsidy for innovation, pushing startups to find the most efficient ways to consume their entire token budget.

QWhy might large enterprises not rapidly replace human workers with AI automation, based on the article's analysis?

ALarge enterprises are subject to the 'Enterprise' pricing model, where each additional token incurs a high marginal cost (with an estimated 75% margin for the provider). This makes the exploration and marginal application of AI for automation expensive. Therefore, while they may automate obvious, high-volume tasks, the high 'tax' on AI labor discourages them from pursuing riskier, experimental, or extensive automation. The economic incentive is to retain existing human labor, as aggressively replacing it with AI would cause their token budget to 'explode' in cost.

QWhat is the '150-person cliff' and what potential consequence does the article suggest it could lead to?

AThe '150-person cliff' refers to the pricing threshold where a company must switch from the subsidized 'Team' plan to the costly 'Enterprise' plan. The article suggests this creates a powerful incentive for AI-native companies to deliberately stay below this size (at or near 149 employees) to maintain their low-cost AI access. This could lead to a new management philosophy where companies are 'obsessed' with using AI agents for everything, minimize human headcount through frequent layoffs and outsourcing, and structure themselves to be extremely lean to avoid crossing the pricing cliff.

QWhat broader economic parallel does the article draw between AI token pricing and traditional policy?

AThe article draws a parallel between AI token pricing and traditional tax policy. The tiered pricing structure (subsidized for small companies, high-margin for large ones) functions as a de facto tax on AI labor. Just as tax codes create incentives and distortions (like companies limiting employee hours to avoid benefits), the AI pricing model creates 'regulatory notches' that shape corporate behavior. The author argues this unlegislated 'tax policy' could become one of the most influential forces shaping the economy in the coming decade by determining which types of companies thrive and how they are structured.

Lectures associées

Reportage : La Banque du Japon pourrait relever ses taux dès septembre, le rythme des hausses pourrait s’accélérer

La Banque du Japon pourrait procéder à une hausse des taux dès septembre, accélérant ainsi son cycle de resserrement monétaire entamé en 2024. Selon des sources informées citées par Reuters, la banque centrale envisage de dépasser le rythme actuel d'environ deux augmentations par an, une accélération qui pourrait la conduire à agir trimestriellement. La prochaine réunion des 17-18 septembre est désormais considérée comme un moment clé, le marché évaluant la probabilité d'une hausse à près de 80%. Cette évolution est motivée par des pressions inflationnistes multiformes : la faiblesse persistante du yen, qui renchérit les importations, une inflation de gros élevée, la remontée des anticipations d'inflation des ménages et des entreprises au-dessus ou près de l'objectif de 2%, ainsi que des chocs externes comme les conflits géopolitiques et la demande mondiale en IA. Lors de sa dernière réunion en juillet, la BoJ a maintenu ses taux mais a émis son signal le plus fort à ce jour en faveur d'un resserrement plus précoce et rapide. Des membres du comité de politique ont souligné la nécessité d'éviter de "prendre du retard", une préoccupation partagée par le gouverneur Kazuo Ueda. Face à la montée des risques inflationnistes, la banque centrale semble déterminée à ne pas trop attendre pour relever davantage le coût du crédit.

marsbitIl y a 8 mins

Reportage : La Banque du Japon pourrait relever ses taux dès septembre, le rythme des hausses pourrait s’accélérer

marsbitIl y a 8 mins

Pour faire face à la menace quantique, Ethereum abandonne Poseidon et se tourne vers les fonctions de hachage traditionnelles

Auteur : ChandlerZ, Foresight News Le 13 août, le chercheur d'Ethereum Justin Drake a annoncé que la Fondation Ethereum abandonnera l'algorithme de hachage compatible SNARK Poseidon au niveau L1, au profit de fonctions de hachage traditionnelles comme SHA2 ou BLAKE2. Cette décision marque un ajustement majeur dans la feuille de route de la cryptographie post-quantique, après huit ans de recherche et des investissements de dizaines de millions de dollars. Poseidon, lancé en 2019, était considéré comme idéal pour les applications comme zkRollup et zkVM, offrant une efficacité supérieure dans les circuits SNARK. Cependant, face à l'exigence de sécurité post-quantique, ses limites sont apparues. Les progrès récents en conception SNARK, notamment via l'utilisation de « domaines binaires », permettent désormais aux fonctions de hachage traditionnelles d'atteindre des performances comparables à Poseidon dans les circuits SNARK, avec environ 1 million de vérifications de hachage par seconde sur un ordinateur portable. La transition est également motivée par l'accélération de la menace quantique. Un rapport de 2026 souligne que les ordinateurs quantiques pourraient compromettre les cryptosystèmes actuels comme ECDSA d'ici 2030-2033, mettant en danger des milliers de milliards d'actifs. Face à cela, Ethereum mise sur des schémas basés sur le hachage, réputés plus résistants aux attaques quantiques. La feuille de route prévoit le déploiement d'une machine virtuelle minimale (leanVM) en 2027, suivie de déploiements sur les couches de consensus, d'exécution et de disponibilité des données en 2028. Cette approche vise à compresser massivement les signatures via leanVM, réduisant la taille des données tout en garantissant la sécurité. D'autres blockchains, comme Solana, se préparent également. Solana a choisi le schéma de signature post-quantique Falcon, normalisé par le NIST, tandis que Starknet prévoit de remplacer Pedersen par BLAKE2 et d'introduire des signatures post-quantiques. Ethereum opte ainsi pour des primitives cryptographiques plus matures et largement analysées, privilégiant la robustesse à long terme.

marsbitIl y a 36 mins

Pour faire face à la menace quantique, Ethereum abandonne Poseidon et se tourne vers les fonctions de hachage traditionnelles

marsbitIl y a 36 mins

Explication du rapport de recherche de Nomura : Les résultats de Lumentum confirment la pénurie persistante de puces optiques, offrant des opportunités structurelles aux fournisseurs chinois

Lumentum a annoncé une croissance de 109% de son chiffre d'affaires au trimestre de juin, atteignant 1,01 milliard de dollars, portée par une forte demande pour tous les types de lasers. Nomura souligne dans un rapport que cette performance confirme que le déséquilibre entre l'offre et la demande pour les puces optiques (EML et lasers CW) devrait persister au moins jusqu'en FY26-FY27, créant une opportunité structurelle pour les fournisseurs chinois. La pénurie concerne notamment les lasers EML, avec une demande soutenue pour les composants 100G et une accélération pour le 200G. Les lasers CW sont essentiels pour les applications silicium-photoniques 1,6T. Les transitions technologiques vers le NPO (une étape vers le CPO) et la croissance des commutateurs optiques (OCS) représentent de nouvelles sources de croissance. Nomura identifie des opportunités pour les entreprises chinoises : Source Photonics pour les puces, InnoLight pour la gestion de la chaîne d'approvisionnement et la transition 800G/1,6T, et TFC Optical Communication pour l'essor du NPO. Les performances de Lumentum valident une tendance : la demande explosive des centres de données pour l'IA dépasse la capacité de production de l'industrie des composants optiques, ouvrant une fenêtre temporelle pour la chaîne d'approvisionnement chinoise.

marsbitIl y a 49 mins

Explication du rapport de recherche de Nomura : Les résultats de Lumentum confirment la pénurie persistante de puces optiques, offrant des opportunités structurelles aux fournisseurs chinois

marsbitIl y a 49 mins

Trading

Spot
活动图片