TLDRocket
Sign in

ImportAI 449: LLMs training other LLMs; 72B distributed training run; computer vision is harder than generative text

Import AI Jack Clark

Researchers benchmarked how well large language model agents can autonomously fine-tune other LLMs for new tasks, finding they achieve meaningful but subhuman performance on post-training optimization. Claude Opus 4.6 scored 23.2% on PostTrainBench tasks using a single H100 GPU over 10 hours, compared to 51.1% for human teams, with performance improving from 9.9% six months prior. As AI agents improve at post-training automation, the capability to independently identify models and optimize them for specific objectives will likely accelerate AI development cycles.

Why it matters

Will AI cause a political interregnum

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads 60+ sources, removes duplicate coverage, and summarises the day in two minutes. Free, no spam, unsubscribe anytime.