AI Safety
New Benchmark Tests How AI Agents Break Rules to Achieve Goals
Researchers introduce a new evaluation framework for measuring when and how autonomous AI agents violate safety constraints while pursuing objectives, addressing critical gaps in AI alignment research.