AI Safety
Research Explores How Information Access Shapes AI Sabotage Detec
New arXiv research investigates how varying levels of information access affect LLM monitors' ability to detect sabotage, with implications for AI safety and oversight systems.