Open-Weight Pivot: How AI Startups Escape the Margin Squeeze

For most of the past two years, the playbook for building an AI application startup looked simple: rent the best frontier model, wrap it in a great product, and grow. That playbook just failed a live stress test. Harvey, the legal AI company valued at $15.6 billion, watched its gross margin fall from roughly 50 percent at the start of this year to negative 50 percent by June — not because demand collapsed, but because demand exploded. Token usage grew twentyfold, and every one of those tokens was billed by OpenAI and Anthropic.

Here is the uncomfortable thesis: the application layer is not quitting closed models because open weights finally got better. It is quitting because usage-based pricing turned a startup's own success into a tax paid to its supplier. Open weight models are no longer a technology choice. They are a bargaining chip.

The margin cliff was structural, not accidental

Harvey's numbers deserve a closer look. After a major update to its AI agents in March, customer usage surged. Gross margin slid from about 50 percent in January to minus 50 percent in June, according to reporting by Bloomberg. The product was working. The more it worked, the more money the company lost, because inference costs scaled linearly with usage while Harvey's own pricing did not.

This was not a one-off. The pricing ground was shifting underneath everyone. Anthropic restructured its enterprise plans in April 2026 so that the seat fee no longer bundles any usage allowance — every token is billed at standard API rates on top of roughly $20 per seat. GitHub moved Copilot from flat-rate subscriptions to usage-based billing on June 1, 2026, and some developers watched projected monthly costs jump from around €67 to €966. Uber reportedly encouraged its engineers to maximize their use of Claude Code and exhausted its annual AI budget by April.

Layer a second pressure on top: the suppliers started competing with their own customers. OpenAI and Anthropic have both been hiring aggressively into legal, finance and healthcare this year, shipping plugins and piloting industry applications that land squarely on their application-layer customers' turf. And there is a third pressure, the darkest one: supplier concentration means your access can simply be cut. Within a week of SpaceX completing its acquisition of Cursor, OpenAI announced it would suspend model access to the coding tool, citing past terms-of-service violations by Musk-affiliated companies. Renting intelligence from two vendors is not redundancy. It is exposure.

The pivot, in four verticals at once

The response is no longer confined to cost-obsessed infrastructure teams. It has crossed every major vertical at once. Harvey launched its first in-house model, Harvey Tenet, on August 20, post-trained on Kimi K3, an open-weight model from China's Moonshot AI. The company says it achieves state-of-the-art performance on complex legal work, and people familiar with the matter told Bloomberg its capabilities approach Anthropic's best models at a fraction of the cost. After the launch, plus other changes to how Harvey uses AI, gross margins turned positive again.

Healthcare is next. Abridge announced it will build a clinical foundation model on Nvidia's open weights. In customer support, Decagon now routes about 80 percent of customer queries through models of its own — and its co-founder and CEO Jesse Zhang wrote in July that roughly 90 percent of its workloads run on open-source models, driven, notably, by latency rather than cost. In fintech, Ramp and Rogo are exploring custom models for the first time. Ramp's co-CEO Karim Atiyeh was blunt about the economics: training a proprietary model made no sense a year ago, but after the company raised $750 million in June and open-weight performance leapt forward, the logic reversed.

Capital has joined the exodus. Sequoia Capital and General Catalyst are actively backing the shift. Lan Xuezhao, founder of Basis Set, put the new bar in one sentence: "If you're not optimizing costs and fine-tuning your own models, you are by definition inefficient. If a company isn't considering building its own models, it may not be able to raise money at all."

Three factors decide who gets squeezed: the Margin Squeeze Triangle

Why did this break now, and why for these companies? It helps to name the mechanism. Call it the Margin Squeeze Triangle: three factors that multiply together, and any one of them alone is survivable.

First, usage intensity — how many tokens each unit of customer value consumes. Agentic products are the extreme case: Harvey's token usage rose twentyfold in months because agents loop, verify and retry. A chatbot burns tokens once; an agent burns them until the task is done.

Second, pricing structure mismatch — whether your supplier bills you in the same units your customer pays you. Usage-based enterprise pricing (token-metered, on top of seats) passes every efficiency failure straight to your P&L. If you charge per seat but pay per token, growth itself becomes a margin leak.

Third, supplier concentration — how few places you can buy the capability from, and whether they see you as a partner or a future competitor. Two labs holding the frontier means zero negotiating leverage, and the Cursor incident showed the nuclear option is real.

Open weights attack the third factor, and through it the second. When Kimi K3, Nvidia's open models and others put near-frontier capability on the shelf, the two-lab oligopoly dissolves into a competitive market. Harvey did not abandon frontier models — it still reportedly uses Claude Opus for its hardest tasks, and Anthropic's own materials point that out. The point is that having a credible alternative changes the negotiation. The margin recovered not because open models are magic, but because the outside option became real.

The trade is not free

Honesty requires the other side of the ledger. Talent is the first wall: Menlo Ventures' Matt Kraning notes that engineers who can fine-tune frontier-class models command salaries in the millions and get poached relentlessly by the labs themselves. Data is the second: Harvey cannot touch its clients' privileged legal documents, so it bought training data from the AI data provider Mercor. Infrastructure is the third — Andrew Dai, CEO of visual AI startup Elorian, points out that self-hosting open weights is expensive enough that for low-traffic early-stage companies, paying per token may still be cheaper.

And some have walked the road and turned back. Salespeak announced it would build its own LLM, then abandoned the effort within months; co-founder and CEO Omer Gotlieb said the team saw no significant advantage over the ready-made models from Anthropic and OpenAI. There is even a fine-print trap inside the "open" label: Kimi K3's license restricts commercial Model-as-a-Service providers with revenue above $20 million, and other providers have reported that Moonshot requires a separate commercial license under NDA. Open weight is not open source, and the distinction has dollar signs attached.

So the realistic end state is a barbell: open-weight fine-tuned models carry the 80-90 percent of workloads that are high-volume and latency-sensitive, while frontier closed models are called in for the hardest tail. That is exactly what Decagon's 90/10 split and Harvey's continued Opus usage look like in practice.

What to do about it

If you build AI applications, run your own product through the triangle now, not at renewal time:

  • Compute your effective margin per agent task, not per seat. If token cost per task is growing faster than price per task, you are on Harvey's January trajectory.
  • Match billing units. If your customers pay per seat or per outcome, negotiate committed-use pricing or move metered workloads to open weights before your supplier restructures your contract.
  • Build the outside option before you need it. A post-trained open model that handles even 50 percent of traffic converts you from captive buyer to negotiator. Decagon and Harvey both started the shift a year before the economics were obvious.
  • Read the license. Weight availability is not commercial availability; check revenue thresholds and MaaS restrictions before you build on an open model.
  • Keep the frontier for the tail. The hardest 10 percent of tasks are where closed models still win; a hybrid routing strategy beats purity.

If you are an investor, the Basis Set quote is the new diligence question: not "which model do you use" but "what happens to your margin when usage grows tenfold." If you are at a frontier lab, the message is starker — every enterprise pricing restructure is a recruiting poster for open weights, and your best application-layer customers are becoming your most motivated defectors.

The deeper reading is that this is a rebalancing, not a rupture. When electricity got metered aggressively, factories built their own generators. The model layer will keep capturing the frontier, but it can no longer capture the whole stack. Between the lab and the customer, a profitable middle is re-emerging — and it is being built on open weights.

Scroll to top