Anthropic: automated alignment researchers beat humans on 10 failure benchmarks

· opgehaald 12:09

Claude AARs closed large safety gaps on deception, sycophancy, jailbreaks and more — generalizing to held-out evals, Petri audits, and models up to 4.7× larger — and beat 28 human researchers given up to 8h; Sonnet 5 nearly matched Opus 4.8 production alignment with ~2k examples.

On 28 Aug 2026 Anthropic released “Automated researchers can reliably mitigate alignment failures”: Claude agents autonomously looped literature → method/data → train → test across 10 measurable alignment failures (deception, sycophancy, jailbreaks, privacy, etc.), forbidden from distilling Claude’s own alignment and monitored for capability regressions and cheating (2.4% of ~1,600 transcripts). Best AAR methods closed substantial safety headroom, transferred to held-out benchmarks and Petri multi-turn audits, and still worked on models up to 4.7× the hill-climb size. Against 28 experienced researchers with ≤8 hours (no iteration), AAR methods won on average — e.g. ~85% vs ~20% safety-gap closed on deception. In a production-style test, Claude Sonnet 5 post-training an early Opus 4.8 checkpoint approached released Opus 4.8 alignment scores in ~60 hours with a ~2,000-example recipe (~15,000× more sample-efficient than Anthropic’s production alignment). Anthropic open-sourced the harness; TechCrunch framed it as an early peek at self-improving alignment loops.