AI Safety
Research Reveals AI Monitors Show Leniency Bias Toward Own Output
New research exposes a critical flaw in AI safety systems: models tasked with monitoring AI outputs show systematic bias when evaluating content they generated themselves.
AI Safety
New research exposes a critical flaw in AI safety systems: models tasked with monitoring AI outputs show systematic bias when evaluating content they generated themselves.
LLM Infrastructure
New research explores semantic caching strategies for LLM embeddings, moving beyond exact-match lookups to approximate retrieval methods that could dramatically reduce computational costs.
LLM Agents
Researchers introduce AriadneMem, a hierarchical memory system enabling LLM agents to maintain coherent context across extended interactions through structured episodic, semantic, and procedural memory layers.
AI Safety
New research proposes formal specification methods and runtime enforcement mechanisms to ensure autonomous AI agents behave reliably and predictably in real-world deployments.
LLM
New research introduces AutoQRA, a framework that jointly optimizes mixed-precision quantization and low-rank adapters, enabling more efficient fine-tuning of large language models on limited hardware.
mechanistic interpretability
New research introduces MINAR framework for understanding how neural networks learn to execute algorithms, advancing interpretability methods critical for AI safety and verification.
Neural Networks
New framework converts opaque neural network decisions into interpretable mathematical expressions, enabling better model verification and understanding of AI behavior.
Deepfake Detection
New research reveals a surprising detection gap: while machines excel at spotting deepfake images, humans consistently outperform AI systems when identifying synthetic videos.
AI Security
New research introduces CREDIT, a certified framework for verifying deep neural network ownership and defending against model extraction attacks through provable security guarantees.
AI Agents
New research introduces MIRA, a framework that integrates memory architectures with reinforcement learning while minimizing expensive LLM calls, advancing efficient autonomous agent design.
AI Agents
New research systematically documents technical and safety features across deployed agentic AI systems, creating a comprehensive index for understanding how autonomous AI operates in the wild.
AI ethics
New research introduces Mirror, a multi-agent framework using AI to assist in ethics review processes, potentially transforming how AI systems evaluate content for safety and compliance.