TLDRocket
Sign in

How a Georgia Tech team used the open Olmo stack to trace social reasoning

Allen Institute (AI2)

Georgia Tech researchers traced where a model’s social reasoning came from using Ai2’s open Olmo stack. The twist: social smarts leaned on narrative, dialogue-heavy data more than on science text.

Based on reporting by Allen Institute (AI2) — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Glenn Matlin and Chandreyi Chakraborty wanted to answer a question that sounds simple and turns nasty fast: where does a model’s social reasoning actually come from? Not from theory alone, but from the training data itself — the bits of text that push a model toward reading beliefs, emotions, intentions, and moral choices the way it does.

To get there, they used influence functions, a method that estimates how much a single training document helped shape an answer on a benchmark. That only works if you know the model really trained on those documents, which is why they needed something fully open. They chose Olmo 3, one of the few LLMs released with its training corpora, checkpoints, and evaluation tools public alongside the model weights.

Their study pulled together five pieces of the Olmo ecosystem: Olmo 3, the Dolma 3 training data, WebOrganizer for categorizing documents, OlmoEval for benchmarks, and OLMES, a shared scoring standard. Dolma 3 is huge, with about 1.26 billion documents, so the team sampled about 5.68 million of them across 576 categories. They then measured how much each category influenced Olmo 3’s answers on four benchmarks: SocialIQA, ARC-Challenge, and MMLU’s social-science and STEM sets.

The result wasn’t the neat split people might expect. Social-science knowledge looked closer to STEM and reasoning than to social reasoning. SocialIQA stood apart, with Olmo 3 leaning heavily on narrative and interpersonal material such as literature, social life, customer support, and Q&A threads. Technical documentation and science writing mattered more for the other benchmarks.

And the team didn’t stop at explanation. They tested causality by having Olmo 3 unlearn the literature documents most tied to social reasoning. Its SocialIQA score fell more than it did when random documents from that category were removed. That doesn’t mean “more literature” is a magic fix, but it does suggest that training-data choices can be tested more directly than before.

The deeper point is about access. Tracing a capability through a model’s full training set has usually been a private-lab sport, which leaves outside researchers guessing. Ai2’s open stack gave an external team enough of the scientific plumbing to reproduce the work from scratch, which is exactly the kind of thing serious model auditing needs.

My take — AI-written commentary, not fact-checked reporting

This is what open models are for: not leaderboard bragging, but letting outsiders poke the thing until the story stops being convenient. The industry loves talking about transparency right up until somebody asks which documents taught the model its manners. Open weights without open data are just a nicer curtain.

Read more about this at: Allen Institute (AI2)

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.