The Data Scientist Show - Building end-to-end ML systems
Eugene Yan
Eugene Yan sat down for a near 2-hour podcast on building real ML systems. It's a rare deep dive into what actually breaks in production ML, not textbook theory.
Based on reporting by Eugene Yan — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Eugene Yan, the engineer-writer known for his practical takes on machine learning, recently joined Daliana Liu's show, The Data Scientist Show, for a conversation that ran nearly two hours. That's a long time for a podcast about data science, and it shows how much ground the two covered rather than skimming the usual highlights reel.
The conversation centered on three threads: the mistakes people keep making when they build ML systems, how to actually measure whether a model did anything useful, and what it takes to grow as a machine learning engineer over time. Yan has spent years working on recommendation systems and production ML at places like Amazon and Alibaba, so his answers lean on specifics rather than platitudes — the kind of war stories that don't show up in a course syllabus.
What stands out is the framing around impact. A lot of ML practitioners can build a model that beats a baseline on a held-out set, but far fewer can explain what that improvement meant in dollars, retention, or some other number a business actually cares about. Yan has built his reputation partly on bridging that gap, and the podcast digs into how he's quantified the value of his own past projects, not just described them.
The episode is available as video on YouTube and as audio on Spotify and Apple Podcasts, so there's no excuse based on format preference. For anyone building ML systems and tired of shallow takes on the field, this is less inspirational talk and more a working engineer explaining, at length, how the job really goes.
My take — AI-written commentary, not fact-checked reporting
I'll take a two-hour unscripted conversation between two people who've actually shipped ML systems over another glossy conference keynote any day — that's where the real lessons about impact and career growth hide, not in slide decks. More practitioners should be doing long-form, unedited chats like this instead of chasing viral takes.
Read more about this at: Eugene Yan