AI research
Neural Paging: AI Learns to Manage Its Own Memory Limits
New research introduces learned policies for context window management in AI agents, enabling more efficient handling of long-running tasks that exceed memory limits.
AI research
New research introduces learned policies for context window management in AI agents, enabling more efficient handling of long-running tasks that exceed memory limits.
LLM Evaluation
Researchers introduce Autorubric, a unified framework that brings systematic rubric-based evaluation to large language models, addressing inconsistent assessment methods across AI systems.
LLM Safety
New research introduces FlexGuard, a continuous risk scoring framework that enables adaptive content moderation strictness for LLMs, moving beyond binary safe/unsafe classifications.
LLM
New research combines reinforcement learning with knowledge distillation to improve how smaller language models learn complex reasoning from larger teacher models.
LLM Agents
New research introduces Tool-R0, a framework enabling LLM agents to autonomously learn tool usage through self-evolution, eliminating the need for curated training datasets while achieving state-of-the-art performance.
LLM Agents
Researchers introduce a new benchmark for evaluating how general LLM agents perform when given additional compute resources at inference time, addressing a critical gap in agent evaluation.
LLM Interpretability
New research introduces ADAPT, a hybrid optimization technique that combines discrete and continuous methods to visualize and understand internal features of large language models.
LLM fine-tuning
New research introduces proxy methods that preserve gradient influence signals while dramatically reducing computational costs for selecting optimal training data in large language model fine-tuning.
LLM
New research applies software product line variability modeling to systematically optimize LLM inference hyperparameters like temperature and sampling strategies.
LLM Safety
New research explores whether constraining specific parameter regions in large language models can ensure safety, examining the theoretical foundations of alignment through architectural constraints.
LLM
New research reveals that LLMs reason better using their own examples rather than human-provided ones, suggesting the process of generation matters more than example quality.
Agentic AI
New research proposes proxy state-based evaluation for multi-turn tool-calling LLM agents, addressing the challenge of scalable reward verification in complex agentic workflows.