DeepSeek-V4.1-Flash Brings 1M Context and FP4 KV Cache
DeepSeek's new V4.1-Flash pushes a 1M-token context window with FP4 KV cache compression and cross-layer attention reuse, slashing memory costs for long-context inference and reshaping the efficiency frontier for large models.