The Endgame Of Vertical Integration
TLDR
Model labs and agent labs are starting to do each other's jobs — Anthropic builds apps, Harvey trains models. The line between 'wrapper' and 'foundation model' is basically gone.
For the past couple of years there's been a tidy division of labor in AI: model labs build the raw intelligence, agent labs wrap it in the plumbing that makes it useful for a specific job — legal research, coding, customer support, whatever. That split is dissolving fast. Anthropic is now building first-party apps that compete directly with its own customers, while Harvey, the legal AI startup that made its name building harnesses on top of other people's models, is now training its own. Everyone's converging on the same territory.
The reason comes down to something a Moonshot AI cofounder, Yang Zhilin, put well: when you build a harness on top of someone else's model, you're reverse-engineering their training process, guessing at what tools, prompts and context structures will fit the model's internal distribution. When you train the model yourself inside your own harness, you flip that around entirely. You design the tools and environment first, then train the model to be great inside them. Zhilin thinks that second path has a much higher ceiling, because you can fix weaknesses on either side of the loop — adjust the tool, retrain the model, repeat.
Anthropic's own answer is blunter. Co-designing harness and model is basically required if you want max performance, because you're always testing the model against some harness — and if you're not building that harness yourself, you're stuck guessing. Poolside's Eiso Kant takes this even further on the coding side, arguing that stuffing a system prompt with 30 or 40 tools is already outdated; give a model a container, a codebase, some credentials, and let it work, and within a year nobody will be shipping bloated tool registries anymore.
The economic pressure behind all this is intelligence-per-dollar, which has become the metric enterprise buyers actually trust, since benchmark scores and token pricing don't reflect the messy, unbenchmarked work companies actually need done. Anthropic's API gross margins reportedly jumped from roughly 38-40% in 2025 to over 70% now, and that kind of margin expansion is exactly what agent labs are racing to match by building their own models tuned to their own harnesses. Whether legal AI specifically needed this — given its modest volumes and middling verifiability — is debatable. But the broader logic, that whoever owns both ends of the stack controls the frontier, is now undeniable.
So the two camps are meeting in the middle, not because anyone planned it that way, but because neither side can afford to leave the other's territory uncontested. Model labs have the compute and the training know-how; agent labs have the domain depth and the customer relationships. The winners will likely be whoever manages to do both cheaply enough to keep pushing their corner of the Pareto frontier further than anyone else can follow.
My take
This is the natural end state of an industry that spent two years pretending 'wrapper' was an insult — turns out owning the weights and the workflow beats owning either alone, and margin data proves it. Vertical-focused startups that don't start training their own models now are betting their survival on model labs staying too distracted to notice their niche, which is a terrible bet given how fast frontier labs are shipping first-party apps. Watch legal, finance and coding closely: whoever blinks first on training gets steamrolled by whoever didn't.
Read more about this at: TLDR
Related stories
The Rise of Intelligence Ownership: a task-trained open source model vs the frontier
TLDR Dev · 6 days ago ·
22
Hidden Technical Debt of AI Systems: Agent Harness
TLDR Dev · 1 month ago ·
31
Up the Stack: How AI’s Escape From the Commodity Trap Risks Enterprise Lock-in
AI Snake Oil · 3 weeks ago ·
14