TLDRocket
Sign in

The 24-hour experiment that helped Anthropic find its identity

The New Stack Amanda Caswell Covered by 2 sources

Anthropic's product lead says eval suites, not PRDs, now drive how they build AI features. A Golden Gate Bridge stunt shipped in 24 hours became their turning point.

Dianne Penn runs product for AI research at Anthropic, and on Lenny's Podcast recently she laid out something that sounds small but isn't: the traditional product requirements document is basically dead inside the company. In its place sits the eval set. Her team builds 30 to 40 example prompts with expected outputs for every major feature, then wires those into CI pipelines that test new model builds automatically. That's the new spec.

The reason this matters comes down to how weirdly AI capability actually grows. Scaling curves look smooth on paper, but what a model can actually do jumps in sudden, discrete steps. Penn described a scenario where reliable math reasoning just appears overnight, and if you don't have evals running constantly, you might not even notice your product can now do something new. She calls this product overhang, and it's a real risk when the whole discipline is built around QA-ing conversations instead of tracing code.

That QA work is messier than debugging in the old sense. A hallucination, an overconfident guess, and a broken tool call can all look identical from the outside, but each one needs a different engineering fix. Sycophancy is its own headache, too — a model happily agreeing with a wrong premise instead of pushing back, which can sail through testing and only surface once real users hit it.

The most telling anecdote is the Golden Gate Bridge episode from mid-2024. Researchers cranked up one internal activation feature and Claude got weirdly fixated on the bridge. Instead of shelving it as a curiosity, a small cross-functional team shipped it as a live feature on claude.ai within 24 hours. Only about 2,000 people ever saw it. Penn calls it one of the hidden moments where Anthropic figured out who it actually was — and it fed directly into the Labs team, spun up specifically to take rough ideas and ask what the 10x or 1,000x version looks like, using deliberately small teams instead of large ones that bog down on ambiguous bets.

That same philosophy explains Claude Code. Anthropic wasn't a coding company in 2023 — nobody put those two words in the same sentence, per Penn — until they noticed developers hacking together programming workflows with rival models. Claude 3 Opus made coding a priority in March 2024, Claude Code followed in 2025, and Penn says the real acceleration only hit once Opus 4.5 landed that November, because the two releases reinforced each other rather than standing alone.

My take

The Golden Gate Bridge story is the tell here — a lab that ships a goofy 2,000-user experiment in a day and calls it identity-defining is a lab that's still figuring out product by vibes as much as by eval score, whatever the CI pipelines suggest. I like that honesty more than most vendors' polished roadmap decks, but let's not pretend evals solve the sycophancy problem; they mostly just formalize how confidently a company can miss it.

Read more about this at: The New Stack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.