Andrej Karpathy
Karpathy's Autoresearch: AI Agents Run ML Experiments Solo
Andrej Karpathy releases Autoresearch, a 630-line Python tool enabling AI agents to autonomously run machine learning experiments on single GPUs, democratizing ML research.
Andrej Karpathy
Andrej Karpathy releases Autoresearch, a 630-line Python tool enabling AI agents to autonomously run machine learning experiments on single GPUs, democratizing ML research.
Yann LeCun
Meta's Chief AI Scientist Yann LeCun argues AGI is fundamentally misdefined in new research paper, introducing Superhuman Adaptable Intelligence as alternative framework for measuring AI progress.
AI research
New research proposes treating AI models as clinical patients, introducing systematic diagnostic and treatment protocols for understanding model behavior, identifying failures, and applying targeted interventions.
Machine Learning
New research investigates representation collapse in continual learning, revealing why neural networks catastrophically forget previous tasks and proposing mechanisms to understand this fundamental limitation.
LLM evaluation
Researchers introduce an automated framework for discovering the hidden concepts LLM evaluators use when judging AI outputs, enabling better understanding and improvement of AI content assessment systems.
LLM Agents
New research introduces PlugMem, a task-agnostic plugin memory module enabling LLM agents to maintain context across sessions without task-specific training.
Deepfake Detection
New research reveals a surprising split in deepfake detection: machines outperform humans at identifying synthetic images, while humans maintain an edge in spotting fake videos.
AI research
New research introduces learned policies for context window management in AI agents, enabling more efficient handling of long-running tasks that exceed memory limits.
LLM evaluation
Researchers introduce Autorubric, a unified framework that brings systematic rubric-based evaluation to large language models, addressing inconsistent assessment methods across AI systems.
LLM Safety
New research introduces FlexGuard, a continuous risk scoring framework that enables adaptive content moderation strictness for LLMs, moving beyond binary safe/unsafe classifications.
LLM
New research combines reinforcement learning with knowledge distillation to improve how smaller language models learn complex reasoning from larger teacher models.
LLM Agents
New research introduces Tool-R0, a framework enabling LLM agents to autonomously learn tool usage through self-evolution, eliminating the need for curated training datasets while achieving state-of-the-art performance.