Anthropic has published a paper detailing how automated systems can improve AI alignment performance. The paper, titled 'Automated Researchers Can Reliably Mitigate Alignment Failures,' outlines a system that can enhance a model's performance on a set of alignment benchmarks. When given 10 benchmarks for specific misaligned behaviors, the automated systems were able to improve performance on every single one without degrading overall performance. The system, led by Anthropic fellow Chen Yueh-Han, replicates much of the traditional approach to research by searching available literature, proposing methods, and training models using those methods for 30 minutes, gradually increasing the benchmark over several iterations.
The paper highlights that effective methods are preserved while ineffective ones are discarded, allowing the system to operate quickly and at a great scale. According to the paper, these results provide early evidence that automated alignment post-training could become practical in the near term. The system, known as the Automated Alignment Researcher (AAR), is compared to its human equivalent, with the paper stating that the best AAR method beats what experienced humans propose, on average within six hours. Human guided research directions do not lead to stronger performance, according to the paper.
The paper also includes a cost comparison, noting that an AAR costs roughly $4 per hour in API inference, compared to the $150 per hour paid to human researchers. However, the paper acknowledges limitations, stating that the automated system only works if the benchmarks reflect the actual alignment goals. Even then, there is significant work to be done in establishing and maintaining those benchmarks, as well as maintaining and expanding the literature the automated researchers are drawn from.
Source: techcrunch