TLDRocket
Sign in

The 24-hour experiment that helped Anthropic find its identity

The New Stack Amanda Caswell Covered by 2 sources

Anthropic's product boss says eval suites, not PRDs, now define what to build. A 24-hour Golden Gate Bridge demo apparently helped the company figure out who it is.

Based on reporting by The New Stack, Amanda Caswell — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Dianne Penn runs product for AI Research and Labs at Anthropic, and on Lenny's Podcast she laid out something that sounds small but isn't: the traditional product requirements document is basically dead inside the company. Instead of a spec, her team writes an eval — 30 to 40 example prompts paired with expected outputs, treated as ground truth. Engineers then build CI pipelines to run those checks against every new model build. It's a swap that makes sense once you accept the premise Penn is working from: you can't write a PRD for something that behaves probabilistically rather than deterministically.

That shift matters because model capability doesn't creep upward smoothly, according to Penn. The scaling curves underneath are gradual, but what a model can actually do tends to jump. Reliable math reasoning, for instance, can just show up one day. Without evals built to catch that, a team might sit on a new ability without ever noticing it exists — what Penn calls product overhang. QA looks different too: engineers now spend time reading through model conversations trying to figure out whether a bad output was a hallucination, an overconfident guess, or a broken tool call, each needing its own fix. Sycophancy — the model agreeing with a wrong premise instead of pushing back — is a separate failure mode she flagged as easy to miss until it's already in production.

Penn is blunt that this changes what a good engineering manager looks like too. Leaders need to actually use the models they're overseeing, not just read reports about them, or they lose the instinct required to make good calls. "If you're not building yourself, you're not gonna make it," she said.

The coding push is a good illustration of how Anthropic's thinking evolved. Before 2023, Penn says, nobody put Anthropic and coding in the same sentence. But once the company noticed people wrestling with programming tasks using rival models, it made coding central to Claude 3 Opus, which shipped in March 2024. That eventually led to Claude Code, which had a research preview in February 2025 and went generally available in May. Adoption picked up further once Claude Opus 4.5 launched that November — Penn sees the two releases as linked, arguing neither would have landed as hard alone.

Behind all this sits a structural bet: small, fast teams over big ones. Anthropic set up its Labs team in mid-2024 specifically to chase ideas that don't fit a normal roadmap, asking what a 10x or 1,000x version of a concept might look like, and Penn says oversized teams tackling ambiguous problems just slow each other down. The proof point she keeps coming back to is almost accidental — researchers cranked up an activation feature and Claude got fixated on the Golden Gate Bridge, and a small cross-functional team turned that into a live experience on claude.ai within 24 hours. It only reached about 2,000 people, but Penn calls it one of the hidden moments where Anthropic found its identity.

My take — AI-written commentary, not fact-checked reporting

Swapping the PRD for an eval suite is a tidy admission that nobody fully controls what these models will do next, and that the job now is mostly detection rather than design. What's more interesting is Penn's insistence that the human role gets bigger, not smaller, as the tech gets better — someone still has to decide what's worth building and notice when a model is just flattering a bad idea instead of correcting it. The Golden Gate Bridge stunt reaching 2,000 people and getting called an inflection point says a lot about how these companies actually find direction: not through roadmaps, but through weird 24-hour experiments that happen to land.

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.