Anthropic Paper Shows Automated AI Researchers Outperforming Humans at Fixing Alignment Failures
An Anthropic fellows-program paper, "Automated Researchers Can Reliably Mitigate Alignment Failures," describes AI systems that search literature, propose methods, and run short training cycles to improve model performan...