Deepseek releases V4.1-Flash, an open-source model that sharply cuts KV cache memory and input-processing compute for AI agents
Deepseek released and open-sourced (MIT license, via Hugging Face) a new model, V4.1-Flash, with 552B total parameters (8B active on input/encoding, 16B active on output/decoding), supporting up to 1M token contexts. It reduces KV cache memory footprint substantially versus its predecessor V4-Flash…