LLM Security
Vocabulary Trojans: A New Threat to LLM Security and Trust
Researchers reveal how malicious actors can embed hidden backdoors in LLMs through vocabulary manipulation, enabling stealthy sabotage that evades detection methods.
LLM Security
Researchers reveal how malicious actors can embed hidden backdoors in LLMs through vocabulary manipulation, enabling stealthy sabotage that evades detection methods.
AI Agents
Learn how to design production-grade agentic AI systems using LangGraph with two-phase commit protocols, human-in-the-loop interrupts, and safe rollback mechanisms for reliable automation.
AI Safety
New research proposes integrating actions, compositional structure, and episodic memory from neuroscience to build safer, more interpretable AI systems that could transform how we approach AI trustworthiness.
AI Safety
Researchers introduce DarkPatterns-LLM, a multi-layer benchmark designed to identify and evaluate manipulative behaviors in large language models, advancing AI safety and authenticity research.
AI Governance
Researchers propose comprehensive framework for governing agentic AI systems, mapping capabilities to risks and establishing safety protocols as autonomous agents become more prevalent.
AI Agents
New research proposes multi-agent deliberation framework where AI agents debate decisions before acting, generating human-readable rationales that improve transparency and reduce harmful behaviors.
OpenAI
OpenAI is hiring a new Head of Preparedness to lead efforts assessing and mitigating risks from frontier AI models, including potential misuse in synthetic media generation.
AI Agents
New research proposes combining blockchain monitoring with agentic AI to create verifiable perception-reasoning-action pipelines, addressing critical trust and authenticity challenges in autonomous AI systems.
AI Safety
Researchers introduce a new evaluation framework for measuring when and how autonomous AI agents violate safety constraints while pursuing objectives, addressing critical gaps in AI alignment research.
AI Safety
New research bridges efficiency and safety by developing formal verification methods for neural networks with early exits, enabling mathematically proven safety guarantees for adaptive AI systems.
LLM Interpretability
New research maps LLM internal representations to brain-derived axes, enabling interpretable reading and targeted steering of model behavior without fine-tuning.
LLM Security
New research reveals how adversarial control tokens can manipulate LLM-as-a-Judge systems into completely reversing their binary decisions, exposing critical vulnerabilities in AI evaluation pipelines.