AI Safety
Formal Behavioral Contracts: Ensuring AI Agent Reliability
New research proposes formal specification methods and runtime enforcement mechanisms to ensure autonomous AI agents behave reliably and predictably in real-world deployments.
AI Safety
New research proposes formal specification methods and runtime enforcement mechanisms to ensure autonomous AI agents behave reliably and predictably in real-world deployments.
LLM
New research introduces AutoQRA, a framework that jointly optimizes mixed-precision quantization and low-rank adapters, enabling more efficient fine-tuning of large language models on limited hardware.
mechanistic interpretability
New research introduces MINAR framework for understanding how neural networks learn to execute algorithms, advancing interpretability methods critical for AI safety and verification.
Neural Networks
New framework converts opaque neural network decisions into interpretable mathematical expressions, enabling better model verification and understanding of AI behavior.
deepfake detection
New research reveals a surprising detection gap: while machines excel at spotting deepfake images, humans consistently outperform AI systems when identifying synthetic videos.
AI Security
New research introduces CREDIT, a certified framework for verifying deep neural network ownership and defending against model extraction attacks through provable security guarantees.
AI Agents
New research introduces MIRA, a framework that integrates memory architectures with reinforcement learning while minimizing expensive LLM calls, advancing efficient autonomous agent design.
AI Agents
New research systematically documents technical and safety features across deployed agentic AI systems, creating a comprehensive index for understanding how autonomous AI operates in the wild.
AI ethics
New research introduces Mirror, a multi-agent framework using AI to assist in ethics review processes, potentially transforming how AI systems evaluate content for safety and compliance.
AI Agents
New research introduces MAPLE, a sub-agent architecture enabling memory, learning, and personalization in agentic AI systems through modular design patterns.
AI Security
New research proposes a multi-agent AI reference architecture for securing enterprise AI deployments, addressing governance challenges in managing AI systems at scale.
Machine Unlearning
New research introduces a principled approach to removing harmful concepts from generative AI models using tempering and classifier guidance, with major implications for synthetic media safety.