TLDRocket
Sign in

We still don’t know how people are really using AI

MIT Technology Review Eileen Guo

Researchers built an independent tracker of real AI chats because company usage reports only show the good parts. Turns out nearly half of real conversations get filtered out of reports like Anthropic's, hiding sensitive uses.

Based on reporting by MIT Technology Review, Eileen Guo — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Anthropic and OpenAI love to tell us how people use their chatbots, but those reports come from the companies themselves, curated to show what they want shown. A new project called the AI Observatory, led by Stanford's Anka Reuel and MIT's Shayne Longpre, set out to check that story against independent data instead of taking it on faith.

The team pulled together 85,633 conversational turns across 24,521 conversations from seven existing datasets, covering 5,000 users and 52 different models between 2023 and 2025. When they ran Anthropic's own filtering methodology on this dataset, 48% of conversations would have been excluded as not work-related. The stuff that got filtered out skewed heavily toward health and relationships, adult and illicit content, harassment and hate, and sexual material, all showing up at far higher rates than in Anthropic's published figures. OpenAI's own 2025 report, notably, found only 30% of ChatGPT consumer use was work-related, which lines up with the idea that personal and sensitive use is bigger than the polished reports suggest.

The patterns get more interesting over time and across platforms. Looking at WildChat, one of the larger datasets in the mix, conversations grew longer and chattier as time went on, with more small talk and less of the AI disclosing that it's a bot, hinting at rising companionship use. At the same time, flagged sensitive exchanges dropped, which the researchers read as a sign that safeguards were improving. Different models also showed distinct personalities in practice: Grok and Gemini got used more for looking things up, with Grok pulling in news and politics questions but also collecting more misinformation, while Anthropic's Claude leaned toward coding, Gemini toward social and roleplay, and ChatGPT toward homework. Even within one company's lineup, GPT-3.5 produced shorter chats than GPT-4o, which drew longer, more iterative sessions consistent with the emotional-attachment concerns that model became known for.

None of this is close to the scale companies work with. Anthropic's own index draws on a million Claude conversations, and OpenAI's report covers 1.5 million ChatGPT chats. The Observatory's dataset is a fraction of that, built from users who consented to share their data, which likely means it still undercounts the most sensitive conversations people would rather not hand over. But the point isn't matching scale — it's having any outside check at all. As Widder puts it, there's currently no way to answer whether a system like Claude gets used mostly for good or mostly for bad, because that information stays locked inside the company.

Anthropic says its published research reflects its teams' particular interests and that outside independent research matters too; OpenAI didn't respond to requests for comment. Reuel's hope is that companies eventually share data directly with outside researchers under proper privacy protections. Until that happens, she argues, anyone making major decisions about AI risk and benefit is flying blind, relying on narratives the labs chose to tell rather than what's actually happening in millions of daily conversations.

My take — AI-written commentary, not fact-checked reporting

Of course company usage reports look flattering — nobody publishes a report that makes their own product look reckless. The fact that half of real conversations vanish once you apply a company's own work-focused filter should worry anyone using those reports to shape policy. Independent projects like this one are exactly what should have existed years ago, before regulators and journalists started treating cherry-picked corporate blog posts as ground truth.

Read more about this at: MIT Technology Review

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.