LLM Agents New Benchmark Tests LLM Agents Against Messy Real-World APIs Researchers challenge the assumption that LLM agents work reliably with perfect APIs, revealing how real-world complexity degrades AI performance.