AI Funding
Humans& Raises Record $480M Seed from Anthropic, xAI Alumni
New AI startup Humans& secures one of the largest seed rounds ever at $480M, founded by veterans from Anthropic, xAI, and Google pursuing 'human-centric' AI development.
AI Funding
New AI startup Humans& secures one of the largest seed rounds ever at $480M, founded by veterans from Anthropic, xAI, and Google pursuing 'human-centric' AI development.
LLM research
New research reveals how LLMs develop 'directional attractors' during reasoning tasks, showing that similarity-based retrieval mechanisms systematically steer iterative summarization toward predictable patterns.
LLM research
New research introduces PrivacyReasoner, a framework enabling LLMs to emulate human privacy reasoning patterns for better protection of personal information in AI systems.
LLM Security
New research introduces State-Transition Amplification Ratio (STAR) to identify inference-time backdoor attacks in large language models by analyzing anomalous reasoning patterns.
LLM alignment
Researchers introduce ECLIPTICA, a framework using Contrastive Instruction-Tuned Alignment (CITA) to enable dynamic switching between aligned and unaligned LLM behaviors for safety research.
LLM unlearning
New research introduces domain-to-instance framework for generating synthetic data to help large language models selectively forget harmful knowledge while preserving useful capabilities.
AI Safety
Researchers introduce GuardEval, a comprehensive benchmark evaluating LLM moderators across safety, fairness, and robustness dimensions—critical metrics for AI content authentication systems.
AI Certification
New research proposes maturity-based certification for embodied AI systems, introducing quantifiable trustworthiness metrics that could reshape how we evaluate AI reliability and authenticity.
LLM Security
New research proposes ALERT, a training-free method to detect jailbreak attacks on LLMs by analyzing discrepancies between internal model representations and output behavior.
AI Safety
New research shows AI models frequently omit key reasoning steps in their explanations, raising critical questions about whether we can trust AI transparency and the reliability of chain-of-thought prompting.
LLM research
New research reveals a fundamental paradox in LLM self-correction: models that excel at fixing errors often produce fewer initial mistakes, while error-prone models struggle to correct themselves.
AI Safety
New research explores how LLM-powered agents may develop biases against humans based on belief systems, revealing critical vulnerabilities in autonomous AI decision-making.