TLDRocket
Sign in

Import AI 461: "Alignment is not on track"; FrontierCode; and synthetic research interns

Import AI Jack Clark

AI researchers just launched a new nonprofit, Sequent, because they think alignment work isn't keeping pace with superintelligence timelines. They're raising $100M+ to bet on unproven theory-heavy approaches labs mostly ignore.

Based on reporting by Import AI, Jack Clark — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

A handful of researchers from the UK's AI Security Institute and the alignment-theory startup Timaeus decided this week that the current approach to AI safety isn't going to cut it, so they built something new. The org is called Sequent, and its founding argument is blunt: superintelligent AI could arrive within a few years, and nobody has a principled reason to believe alignment techniques will scale with it. Not "we hope so." Not "probably fine." Actual proof, or something close to it, before anyone trains a system that outstrips human oversight.

Sequent wants 40 to 80 full-time researchers within a couple years and is targeting $100-150 million in initial funding, with plans to ask for ten times that if early bets pay off. The pitch to funders is essentially a portfolio strategy: instead of chasing one clever trick, throw resources at scalable oversight, learning theory, heuristic arguments, game theory, and AI personas simultaneously, and look for places where these research threads reinforce each other. One example they give: using scalable oversight to figure out which training variables actually matter, informed by insights from personas and learning theory about what can even be tweaked in the first place.

The critique embedded in Sequent's founding document is aimed squarely at frontier labs, which the group describes as reactive — fixing failures as they surface rather than building theory that predicts when and why alignment might break. That's workable when today's models have occasional weird failure modes that engineers can patch after the fact. It gets a lot scarier once AI systems start handling recursive self-improvement, building bigger and bigger chunks of their own successors with less human involvement at each step. Sequent's explicit hope is to stay independent enough that if a lab is barreling toward something dangerous, someone outside the commercial pressure cooker can say so loudly.

Meanwhile Cognition, the company behind the coding agent Devin, released a benchmark called FrontierCode built specifically to be brutal. Its hardest tier, Diamond, currently caps out at 13.4% for Claude Opus 4.8 — the best score recorded. Twenty open-source maintainers spent over 40 hours per task pulling real, messy multi-PR problems from their own repos rather than scraping GitHub issues wholesale, and grading covers not just whether code runs but whether it's mergeable: does it follow the codebase's conventions, does it touch only what's necessary, do the tests actually test anything real. Given how fast benchmarks like SWE-Bench saturated, a eval this punishing might actually hold up for a year or two, though the newsletter's own guess is 70%+ on Diamond by mid-2027.

Elsewhere, Xiaomi pushed inference speed to 1,000 tokens per second on a trillion-parameter model called MiMo-V2.5-Pro-UltraSpeed, running on ordinary 8-GPU nodes rather than exotic hardware, using tricks like FP4 quantization and a custom speculative decoding scheme called DFlash. And a group of Chinese universities built AARRI-Bench, testing whether AI agents can act like a competent research intern — spotting fabricated data in a manuscript, following research norms, using tools appropriately. Claude Opus 4.7 topped the leaderboard at 68.3%, which sounds decent until you remember that's the bar for entry-level intern work, not for anything resembling independent science.

My take — AI-written commentary, not fact-checked reporting

Sequent's pitch is the most honest thing I've read from anyone adjacent to frontier labs in a while — admitting that empirical patch-and-monitor safety work won't produce confidence before someone trains something genuinely superhuman. I'd rather see ten well-funded independent shops taking weird theoretical bets than one more lab quietly hoping its RLHF pipeline generalizes to recursive self-improvement. The FrontierCode and AARRI benchmarks tell the same underlying story: models are getting scarily competent at narrow execution while still failing basic judgment calls, which is exactly the gap alignment research needs to close before anyone lets an AI system meaningfully improve itself unsupervised.

Read more about this at: Import AI

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.