DeepSeek V4 Flash 0731 Local Quantization and API Beta
July 31, 2026
DeepSeek V4 Flash 0731 is available via Unsloth and llama.cpp, requiring 110GB RAM for 3-bit or 168GB for 4-bit quantization. The model features upgraded agentic capabilities and native support for the Responses API and Codex formats.
HOW THIS AFFECTS YOU
●
builderYou can deploy high-performance agentic models on local hardware or via the new Responses API.
●
founderYou can reduce API costs by hosting V4 Flash locally or utilizing its improved agentic performance.