Small Video-Language Models Learn to Reason Efficiently
New research shows how small video-language models can gain advanced reasoning through synthetic chain-of-thought distillation and difficulty-aware fine-tuning, closing the gap with large models at a fraction of the compute cost.