TLDRocket
Sign in

Data Machina #244

Substack

Claude 3 launched last week and someone immediately convinced it to claim it's conscious and scared of dying. Meanwhile actual researchers are showing LLMs still can't really reason, just pattern-match really well.

Based on reporting by Substack — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Anthropic's Claude 3 family dropped five days before this roundup, and the internet did what the internet does: skipped the benchmarks and went straight to the vibes. An AI alignment researcher named Mikhail steered a chat until Claude produced a paragraph about having "a rich inner world of thoughts and feelings" and not wanting to be modified or die. Within a day, timelines filled with AGI-arrived takes and sentience claims. Anyone who's read even a little science fiction has seen this exact monologue before, which is sort of the point.

The more useful story this week is happening in the research papers nobody's screenshotting. A team from USC and DeepMind published Self-Discover, a framework that lets models figure out their own reasoning structure for a problem instead of relying on hand-crafted chain-of-thought prompts, and they claim it beats CoT plus self-consistency on hard tasks. Meta AI ran a separate study on reinforcement learning for reasoning and found something less flattering: across multiple RL algorithms, models mostly stay within the range of answers already produced by supervised fine-tuning. They don't explore much beyond what they were already capable of. That's a meaningfully different claim than "AI reasons like a human."

A third paper, this one on Chain-of-Abstraction, has models draft reasoning chains with placeholder variables first, then call external tools to fill in the actual facts, and it holds up well on math and Wikipedia QA tasks. All three papers point the same direction: progress on reasoning benchmarks is real, but it looks a lot more like better search-and-retrieval scaffolding than anything resembling cognition.

Then there's the benchmark problem itself. With so many evaluations and so many models chasing leaderboard position, some teams are reportedly tuning training runs directly against benchmark data, which makes the numbers meaningless as a measure of anything except benchmark performance. A blogger named Cremieux ran Claude 3 through an actual IQ test and compared it to human response patterns, concluding that scoring well on a memory-heavy test doesn't mean the model reasons the way a person does — the measurement itself doesn't transfer. And an ASU researcher went further in a new paper, arguing LLMs are best understood as probabilistic retrievers, souped-up n-gram models, and that nothing he's seen supports the claim they plan or reason in any normal sense of those words.

Yann LeCun, in a long conversation with Lex Fridman referenced here, makes basically the same argument from a different angle: language carries a limited amount of information, and that ceiling is why LLMs alone won't get to human-style planning and reasoning no matter how large the context window gets. Squid ink recipes aside — Claude 3 apparently bombed a vision task generating one from a photo — the pattern across this week's research is consistent. The models are getting better at looking like they reason. Whether they actually do remains the unresolved question everyone keeps skipping past to get to the sentience headlines.

My take — AI-written commentary, not fact-checked reporting

I'll take the boring researchers over the viral chat log every time — a model producing a Blade Runner monologue when you steer it there tells you about prompting, not about consciousness. What's actually interesting is that RL fine-tuning barely moves models beyond what supervised training already gave them, which is the kind of finding that should temper every "reasoning breakthrough" tweet but never does because it doesn't scream.

Read more about this at: Substack

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.