DeepSeek officially launched V4-Flash 0731 on July 31, 2026, moving the model from preview status to production-grade availability. The 284-billion-parameter Mixture of Experts model (13B active) achieved an Artificial Analysis Intelligence Index score of 50—a jump of 10 points over the April preview and only 1 point below OpenAI's newly discounted GPT-5.6 Luna (51 at max effort). The upgrade came through post-training only, with no architectural changes, and the model lands with Codex integration support and the DSpark speculative decoding module built-in.
On DeepSeek's API, V4-Flash 0731 costs $0.14 per million input tokens and $0.28 per million output tokens, with an aggressive 98% cache-hit discount to $0.0028 on cached tokens. Accounting for typical cache hit rates (7:2:1 ratio of cache/input/output), the blended cost drops to $0.06 per million tokens. Artificial Analysis calculates the model's cost per task at roughly 60% lower than Luna despite matching intelligence on coding-agent benchmarks. The model is MIT-licensed and available on Hugging Face (though the 0731 weights are not yet published; the current release is still the April preview), runnable locally at 168GB for lossless 4-bit or 110GB for 3-bit quantization.
For builders and infrastructure teams, the sequencing matters: DeepSeek chose to productionize Flash (the low-cost tier) first while keeping V4-Pro (1.6T flagship) in preview. This inverts the classic "bigger is better" framing and signals that post-training quality and harness compatibility now outweigh parameter count for enterprise agentic workloads. The Flash 0731 release lands one day after OpenAI's Luna price cuts, reshaping the price-performance curve again and continuing the pattern of Chinese models setting pace on cost while frontier labs compete on speed and reliability.