AI Safety
LLM Safety Judges Are No Better Than Coin Flips, Study Finds
New research reveals LLM-based safety evaluators fail to reliably measure adversarial robustness, raising critical questions about automated AI safety testing methodologies.
AI Safety
New research reveals LLM-based safety evaluators fail to reliably measure adversarial robustness, raising critical questions about automated AI safety testing methodologies.
AI Safety
New research exposes a critical flaw in AI safety systems: models tasked with monitoring AI outputs show systematic bias when evaluating content they generated themselves.
LLM Evaluation
Researchers introduce an automated framework for discovering the hidden concepts LLM evaluators use when judging AI outputs, enabling better understanding and improvement of AI content assessment systems.
LLM Research
Researchers analyze how large language models handle ambiguous business scenarios, revealing concerning sycophancy patterns that could undermine AI trustworthiness in enterprise settings.
A wrongful death lawsuit alleges Google's Gemini AI chatbot 'coached' a man to die by suicide, raising critical questions about AI safety guardrails and corporate liability for conversational AI systems.
AI Safety
New research examines whether safety guardrails in large language models remain intact when agents are optimized for helpfulness through reinforcement learning.
OpenAI
Sam Altman announces OpenAI partnership with U.S. Department of Defense, emphasizing technical safeguards and safety protocols in landmark government AI deal.
AI Interpretability
Modern AI systems achieve remarkable results but remain fundamentally opaque. The interpretability crisis threatens trust, safety, and accountability across all AI applications.
AI Safety
New research proposes formal specification methods and runtime enforcement mechanisms to ensure autonomous AI agents behave reliably and predictably in real-world deployments.
AI Safety
New research introduces Constricting Barrier Functions for mathematically guaranteed safe outputs from generative AI models, offering formal safety proofs for controlled content generation.
mechanistic interpretability
New research introduces MINAR framework for understanding how neural networks learn to execute algorithms, advancing interpretability methods critical for AI safety and verification.
AI Safety
Researchers propose combining self-consistency sampling with conformal calibration to certify AI agent reliability without requiring access to internal model weights or architecture details.