AI
Anthropic paper shows automated systems beat human researchers at alignment training for $4 an hour
Anthropic has published an Automated Alignment Researcher paper led by fellow Chen Yueh-Han showing that AI systems can improve a model's performance across ten misaligned-behavior benchmarks without degrading overall capability. The company says the best automated method beats what experienced human researchers propose on average within six hours of compute, at roughly $4 per hour of API inference versus $150 per hour for human researchers, an early step toward recursive self-improvement of safety training.