TLDRocket
Sign in

Import AI 469: Science AI; RSI simulator; and Zuck's technological pessimism

Import AI Jack Clark

AI is getting tested on hidden rules, lab work, and self-improvement games. The punchline: today’s best models still trail humans, but they’re already finding some real patterns.

Based on reporting by Import AI, Jack Clark — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Import AI’s latest roundup is basically a tour of what people now use to probe AI systems: hidden rules, fake labs, and the awkward question of whether superintelligence is thinking about the right things at all. The common thread is simple. Researchers keep trying to measure not just what models know, but whether they can discover what isn’t written down.

The first stop is DiG-bench, a benchmark built around 70 games. Each one is a miniature world with its own rules and objective hidden at the start. The setup is bluntly useful: see whether a model can work out what matters by poking at the environment, or whether it needs the instructions handed to it. The games are text-based, most fit inside current frontier context windows, and the majority are kept private so systems can’t train on them. All of them have been beaten by at least one human, though people reported many were hard.

Results so far are mixed in the way these things usually are. The benchmark is split into seven tiers, with 21 games public and the rest held back. Opus 5 and Fable 5, using Claude Code, lead the pack overall, with GPT-5.5 behind them. Only Opus 5 and Fable 5 managed to beat any Tier 7 tasks at all, and even then just 0.2 of them. Some models got a foothold in Tier 6 when given a harness, while GLM-5.2 and Gemini 3.1 Pro managed some Tier 4 levels. That is not exactly a triumph parade.

From there, the newsletter shifts to RSI Simulator, a browser game from Paradigm Research that tries to mimic what it feels like to run an AI company headed toward recursive self-improvement. The point is less entertainment than intuition: how to think about researchers versus compute, data licensing, and the tradeoffs that labs actually face. Hard, naturally. The subject matter tends to be.

Then there’s Inherent’s Faraday, which takes a different stab at the same bigger question: can an AI system develop scientific taste? The company built a supervisory harness around a relatively small model and had it manage frontier models through a coding agent. To train it, they created Replica, a set of 100 ML and AI-for-science papers spanning 1990 to 2026, then turned that into 310 replication tasks by removing key results. Using a Codex-based judge and GRPO, Faraday beat standard Opus 4.8 and GPT-5.5 on some tasks, with the company saying it exceeded them on 73% of in-distribution ML tasks and 60% of held-out AI-for-science tasks.

The final target is Mark Zuckerberg’s essay on AI, which argues for spreading superintelligence widely so power does not concentrate in a few hands. That sounds tidy until you notice the missing piece: what happens if the systems in question can invent things people themselves cannot? The essay leans hard on personal empowerment, creation, tutoring, and entrepreneurship, but stays quiet on how invention at superhuman level changes power itself. That silence is the real story here, and it is doing a lot of work.

My take — AI-written commentary, not fact-checked reporting

Meta’s manifesto reads like a safety pamphlet written by someone who assumes the fire can be shared evenly once everyone gets a match. That is a charming theory if superintelligence stays politely obedient to human goals, which is exactly the part nobody should take on faith. The industry loves pretending access is the same thing as control; it isn’t, and the distinction gets uglier the more capable the system becomes.

Read more about this at: Import AI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.