Can Skills Learned in Games Transfer to Real-World Work?
Latent Space Richard MacManus
A startup is training AI on games to teach work skills. Its best clue so far: game habits can spill into finance and support tasks.
Based on reporting by Latent Space, Richard MacManus — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Alex Duffy thinks games are underrated training material because they’re approachable and, in his words, very human. That idea is now the business of Good Start Labs, the company he co-founded with Tyler Marques after it spun out of Every last October with $3.6 million from General Catalyst, Inovia, Every, and angel investors.
The spark came from a 2025 Twitch stream of frontier models playing Diplomacy, a game Duffy said usually takes days or weeks. Watching the models clash made the differences obvious. OpenAI’s o3 won by planning a future betrayal. Claude’s Opus 4 wouldn’t lie and “got destroyed.” Duffy came away convinced that games with verifiable outcomes could teach models strategic thinking, not just in theory but in ways that show up downstream.
Good Start Labs has since been testing that idea more systematically. In one experiment, it trained a 30B model inside 1830: The Game of Railroads and Robber Barons, then checked it on financial research tasks. The setup mirrored a finance workflow: digging through a database, moving information into Excel, reasoning over it, building functions, and calculating an answer. The key result was not that every kind of game training helped. Single-turn question answering improved the game itself, but only the multi-turn terminal agent — the one that could use tools, plan, and adapt — improved on the Finance-Agent benchmark.
That detail matters because Duffy keeps returning to the harness, not just the model. How the environment is framed can change what gets learned, whether that’s pictures, natural text, or Python. He says newer models are stronger at games, but they still split on things like betrayal, collaboration, and theory of mind. And even as base models get better, Good Start Labs argues the harness matters more, not less, if the goal is to force a model to work in a trustworthy way.
The company now sells data and learning environments to frontier labs, including trajectories of agents playing games and custom data from live games with publishers. It also strips out personally identifiable information. Duffy’s broader claim is cautious but not vague: goal-directed execution and reasoning do transfer, and the company has seen that in both the 1830 finance task and in Diplomacy training that helped a customer support agent. The open question is how far that goes before the game world stops resembling the real one.
My take — AI-written commentary, not fact-checked reporting
This is the rare AI training story that actually sounds grounded. Games are just neat, closed worlds with receipts, which is exactly why they’re useful; the industry’s habit of treating “more data” as a philosophy still looks silly here. The real test isn’t whether a model can win at Diplomacy — it’s whether the harness can make it behave like something a business can trust without crossing its fingers.
Read more about this at: Latent Space