LLMs are stuck in a groupthink groove. This startup is trying to get them out.
MIT Technology Review Will Douglas Heaven
Ask ChatGPT, Claude, or Gemini for a random number and you'll get 7, almost every time. A startup called Springboards built an LLM that actually breaks the pattern.
Based on reporting by MIT Technology Review, Will Douglas Heaven — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Try this: ask any major chatbot for a random number between 1 and 10. You'll get 7. Ask again, you'll get 3 or 4. It's a party trick that works because today's LLMs aren't nearly as creative as their hype suggests — they converge on the same safe, high-probability answers, whether you're asking about numbers, cars, or band names.
A NeurIPS best-paper winner called "Artificial Hivemind" put numbers behind the vibe. Researchers asked 25 different LLMs, including top US models and open-source ones from China, to write a metaphor about time — 50 times each. Out of 1,250 responses, nearly all landed on some version of "time is a river" or "time is a weaver." Ask for a band name and you'll drown in "Glass Harbor," "Neon Hearts," and "Velvet Echo." OpenAI's explanation is that training for reliability naturally pushes models toward familiar answers, and pushing for novelty risks incoherence.
Australian startup Springboards is betting there's a market for the opposite. Its model, Flint, is built on top of Alibaba's open-source Qwen 3 — training a foundation model from scratch was never realistic for a small team, says cofounder Kieran Browne. Rather than just cranking up the "temperature" setting, which Browne found just makes models babble incoherently or randomly switch into code mid-sentence, Springboards trained Flint to spot the specific moments in a response where variety actually matters — the destination in a travel question, not every word around it — and inject randomness there.
In demos, the difference is real if modest. Asked for a car, ChatGPT and Claude reliably say Toyota or Honda; Flint said Ford F-150. Asked for a New Balance tagline, both mainstream models produced "Run your way"; Flint gave "Built to last, run to win." Business strategist Zoe Scaman ran a classic MBA case study — reinventing a finance company for younger customers — through all four models. The big three all reached for "teach financial literacy in a fun way." Flint suggested rebranding the entire concept of wealth accumulation instead.
Flint is still rough. It's aimed at ad agencies and marketers for now, and even its boosters admit it breaks down under pressure. Marketing exec Maximilian Weigl likes using it alongside the mainstream tools but notes that most of the time, "good enough" — meaning familiar and average — is exactly what clients want. Which is really the tension here: chatbots got optimized to be reliable and inoffensive, and that same optimization is what makes them boring at the exact moments you want them not to be.
My take — AI-written commentary, not fact-checked reporting
This is a niche fix for a real problem, and I like that Springboards didn't try to solve it by cranking a temperature dial — that's the lazy move everyone tries first. But the bigger story is what it reveals about the entire industry: every major lab trained its model to converge on the same safe middle, and calls that reliability. Homogenized creativity is a business risk nobody's pricing in yet, and a small Qwen-based startup shouldn't be the only one noticing.
Read more about this at: MIT Technology Review