LLM Agents
New Benchmark Tests How LLM Agents Scale at Inference Time
Researchers introduce a new benchmark for evaluating how general LLM agents perform when given additional compute resources at inference time, addressing a critical gap in agent evaluation.