TLDRocket
Sign in

What We've Learned From A Year of Building with LLMs

Eugene Yan

Five AI practitioners wrote up everything they learned building real LLM products for a year. It's the opposite of a hype thread — just what breaks and how to fix it.

Based on reporting by Eugene Yan — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Eugene Yan and a handful of other engineers who've actually shipped LLM-powered products spent the last year comparing notes, and the result reads less like a blog post and more like a field manual. No launch announcement, no benchmark chart claiming state-of-the-art anything. Just a long, unglamorous accounting of what it takes to get a language model from a cool demo into something people rely on.

The piece splits into three layers, which is honestly the most useful part of it. There's the tactical layer — prompting tricks, retrieval setups, when fine-tuning actually pays off versus when it's a waste of a GPU budget. Then there's the operational layer, the stuff nobody puts in a keynote: how you build evals that catch regressions before users do, how you structure a team so a single prompt engineer isn't the bottleneck for every feature, how you handle the fact that the same input can produce three different outputs on three different days. And finally the strategic layer, which is really a business question dressed up as a technical one — build versus buy, how fast to move, when an LLM feature is a nice-to-have versus a liability.

What comes through in all three sections is a kind of hard-won humility. These are people who've watched confident demos fall apart the moment real users show up with real, weird inputs. Guardrails matter more than clever prompts. Evaluation is the unglamorous work that actually determines whether a product survives contact with customers, and most teams underinvest in it because it doesn't produce a flashy screenshot.

There's also a quiet argument buried in here against the industry's obsession with bigger models and longer context windows as the answer to everything. The authors' experience says the opposite: most failures aren't solved by a bigger model, they're solved by better retrieval, tighter evals, and teams that actually talk to their users. It's not a sexy conclusion. But it's the one that tends to be true a year into actually running this stuff in production, rather than a week into playing with it in a notebook.

My take — AI-written commentary, not fact-checked reporting

This is the kind of piece that never trends the way a new model release does, and that's exactly why it's worth more. Anyone can post a benchmark; almost nobody writes down what it actually costs to keep an LLM feature from embarrassing you in production. I'd trade ten 'GPT-5 is here' threads for one more essay like this.

Read more about this at: Eugene Yan

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.