LLM Safety
Global Subspace Projection: A New Approach to LLM Detoxification
Researchers propose a novel technique for removing toxic behaviors from large language models by projecting out malicious representations in the model's latent space.
LLM Safety
Researchers propose a novel technique for removing toxic behaviors from large language models by projecting out malicious representations in the model's latent space.
AI Safety
Researchers introduce GuardEval, a comprehensive benchmark evaluating LLM moderators across safety, fairness, and robustness dimensions—critical metrics for AI content authentication systems.
AI Safety
Comprehensive technical guide to implementing AI safety guardrails, from prompt-based filtering to advanced validation architectures. Covers practical methods for ensuring secure and relevant AI interactions with code examples.
AI Safety
New research exposes critical AI safety flaw: rhyming prompts bypass guardrails in 62% of language models tested, revealing how poetic formatting defeats content moderation systems through pattern recognition exploitation.
LLM Safety
New research demonstrates how multi-agent debate frameworks can evaluate LLM safety more efficiently than traditional methods, reducing costs while maintaining accuracy in identifying harmful model behaviors.
OpenAI
OpenAI unveils open-weight safety models designed to help developers build safer AI applications, marking a shift toward more accessible AI safety tooling and moderation infrastructure.