LLM LLM Inference Evolved: A Guide to Decoding Algorithms A technical look at how decoding algorithms — from greedy search to nucleus sampling and speculative decoding — shape LLM inference quality, latency, and cost in modern generative AI systems.