An Anthropic researcher just gave us a peek at self-improving AI
TechCrunch Russell Brandom ● Covered by 5 sources
Anthropic says AI models can help train other AI models, and it worked on 10 alignment tests. That’s a real step toward self-improving systems, and it beat human researchers on speed and cost.
Based on reporting by TechCrunch, Russell Brandom — read the original for the full story.
Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error
Anthropic has published an early look at something AI labs have been chasing for a while: models that help train other models. In a new paper, “Automated Researchers Can Reliably Mitigate Alignment Failures,” the company says automated systems improved a model’s performance across 10 benchmarks for specific misaligned behaviors without hurting overall performance.
The work was led by Anthropic Fellow Chen Yueh-Han and tries to mirror how human researchers actually work. Each automated system searches the literature, proposes a method, then trains the model for 30 minutes at a time, repeating the process over several iterations. The useful methods stay. The dead ends get tossed. That makes the process fast, and it also makes it easy to scale.
Anthropic is careful to frame the result as early evidence, not a finished product. Still, the paper is pointed about the comparison with people. It says the best Automated Alignment Researcher method outperforms what experienced humans propose on average within six hours, and that human-guided research directions do not produce stronger performance.
The cost gap is even starker. Anthropic puts the automated system at about $4 per hour in API inference, compared with $150 per hour for its human researchers. But the paper also admits the obvious catch: the whole setup only works if the benchmarks really capture the alignment goals, and those benchmarks, plus the literature the system relies on, still need a lot of care and upkeep.
My take — AI-written commentary, not fact-checked reporting
This is the part of AI that should make people pay attention, not the demo reels and cartoonish chatbot theatrics. If machines start beating humans at the grind of model training, the real bottleneck shifts to who defines the tests and who gets to say the model is “aligned.” That is not a small detail; that is the whole game, and it’s exactly where the mess begins.
Read more about this at: TechCrunch