Het verhaal
On 28 Aug 2026 Anthropic published research on automated alignment researchers (AARs): Claude agents that loop through literature search, propose training methods/data, train a target model, and evaluate on public benchmarks for ten alignment failures (including deception, sycophancy, and jailbreaks). The strongest AAR methods closed substantial safety gaps, generalized to held-out benchmarks and Petri multi-turn audits, and transferred to models up to ~4.7× larger than the training target while preserving capability. In a human baseline, 28 experienced researchers had up to eight hours on the same tasks; their methods underperformed the best AAR methods, and seeding AARs with human ideas did not help. Anthropic frames this as evidence that automating alignment work on well-measured failures may be practical near-term, while still requiring oversight against reward hacking.