AI Safety
Research: LLM Safety Training Survives RL Optimization
New research examines whether safety guardrails in large language models remain intact when agents are optimized for helpfulness through reinforcement learning.
AI Safety
New research examines whether safety guardrails in large language models remain intact when agents are optimized for helpfulness through reinforcement learning.
OpenAI
Sam Altman announces OpenAI partnership with U.S. Department of Defense, emphasizing technical safeguards and safety protocols in landmark government AI deal.
AI Interpretability
Modern AI systems achieve remarkable results but remain fundamentally opaque. The interpretability crisis threatens trust, safety, and accountability across all AI applications.
AI Safety
New research proposes formal specification methods and runtime enforcement mechanisms to ensure autonomous AI agents behave reliably and predictably in real-world deployments.
AI Safety
New research introduces Constricting Barrier Functions for mathematically guaranteed safe outputs from generative AI models, offering formal safety proofs for controlled content generation.
mechanistic interpretability
New research introduces MINAR framework for understanding how neural networks learn to execute algorithms, advancing interpretability methods critical for AI safety and verification.
AI Safety
Researchers propose combining self-consistency sampling with conformal calibration to certify AI agent reliability without requiring access to internal model weights or architecture details.
mechanistic interpretability
New research goes beyond behavioral analysis to trace the internal mechanisms LLMs use when weighing competing reward signals, offering insights into AI decision-making at the circuit level.
LLM Interpretability
New research introduces ADAPT, a hybrid optimization technique that combines discrete and continuous methods to visualize and understand internal features of large language models.
AI Agents
New research systematically documents technical and safety features across deployed agentic AI systems, creating a comprehensive index for understanding how autonomous AI operates in the wild.
Machine Unlearning
New research explores machine unlearning for LLM agents, addressing how autonomous AI systems can selectively forget data while maintaining tool-use and reasoning capabilities.
AI Agents
Technical guide to implementing traceable AI decision-making with comprehensive audit logging and human oversight checkpoints for accountable autonomous systems.