DeepSeek V4-Flash 0731 exits preview with 50 intelligence score, 60% cheaper than GPT-5.6 Luna
DeepSeek officially launched V4-Flash 0731 on July 31, 2026, moving the model from preview status to production-grade availability. The 284-billion-parameter Mixture of Experts model (13B active) achieved an Artificial Analysis Intelligence Index score of 50—a jump of 10 points over the April preview and only 1 point below OpenAI's newly discounted GPT-5.6 Luna (51 at max effort). The upgrade came through post-training only, with no architectural changes, and the model lands with Codex integration support and the DSpark speculative decoding module built-in.
On DeepSeek's API, V4-Flash 0731 costs $0.14 per million input tokens and $0.28 per million output tokens, with an aggressive 98% cache-hit discount to $0.0028 on cached tokens. Accounting for typical cache hit rates (7:2:1 ratio of cache/input/output), the blended cost drops to $0.06 per million tokens. Artificial Analysis calculates the model's cost per task at roughly 60% lower than Luna despite matching intelligence on coding-agent benchmarks. The model is MIT-licensed and available on Hugging Face (though the 0731 weights are not yet published; the current release is still the April preview), runnable locally at 168GB for lossless 4-bit or 110GB for 3-bit quantization.
For builders and infrastructure teams, the sequencing matters: DeepSeek chose to productionize Flash (the low-cost tier) first while keeping V4-Pro (1.6T flagship) in preview. This inverts the classic "bigger is better" framing and signals that post-training quality and harness compatibility now outweigh parameter count for enterprise agentic workloads. The Flash 0731 release lands one day after OpenAI's Luna price cuts, reshaping the price-performance curve again and continuing the pattern of Chinese models setting pace on cost while frontier labs compete on speed and reliability.
Sources
- Primary source
- DeepSeek V4 Flash 0731: Official Release, Agent Benchmarks
“284B/13B MoE, same architecture re-post-trained, MIT-licensed, Artificial Analysis scores 50 on Intelligence Index”
- DeepSeek API Changelog
“DeepSeek-V4-Flash API now in public beta with Codex support and significantly enhanced agent capabilities”
- Artificial Analysis: DeepSeek V4 Flash 0731
“Cost per task 60% lower than Luna at matching intelligence; $0.06 blended rate on 7:2:1 cache/input/output ratio”