TLDRocket
Sign in

ImportAI 449: LLMs training other LLMs; 72B distributed training run; computer vision is harder than generative text

Import AI Jack Clark

Researchers benchmarked how well large language model agents can autonomously fine-tune other LLMs for new tasks, finding they achieve meaningful but subhuman performance on post-training optimization. Claude Opus 4.6 scored 23.2% on PostTrainBench tasks using a single H100 GPU over 10 hours, compared to 51.1% for human teams, with performance improving from 9.9% six months prior. As AI agents improve at post-training automation, the capability to independently identify models and optimize them for specific objectives will likely accelerate AI development cycles.

Why it matters

Will AI cause a political interregnum

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.