DeepSeek V4 Flash, released August 1, 2026, costs roughly 3 cents per Artificial Analysis Intelligence Index test—the cheapest well-known model to run by a factor of 30. Moonshot Kimi K3 costs 86 cents per test, OpenAI's GPT-5.6 Sol $1.86, and Anthropic's Claude Fable 5 $3.15. Published pricing sits at $0.14 per million input tokens and $0.28 per million output tokens. V4 Flash scores 50 on the Intelligence Index, matching Google Gemini 3.6 Flash and placing it one point behind Meta's Muse Spark 1.1, but well below frontier models Claude Opus 5 (61), Claude Fable 5 (60), and GPT-5.6 Sol (59).
The model is a Mixture-of-Experts with 284 billion total parameters and 13 billion activated, supporting a 1 million-token context window. A July 31 post-training update drove a massive jump: DeepSeek V4 Flash scores 54% on agentic coding benchmarks (DeepSWE), up from 7% in the preview version. On Terminal Bench 2.1, it scores 82.7%, beating the heavier V4-Pro-Preview at 72.1%—a reversal where DeepSeek's budget model outperforms its premium sibling. The model is available via the DeepSeek API with a new V4-Flash-0731 checkpoint; weights are open under MIT License on Hugging Face.
For practitioners, V4 Flash rewrites the unit economics of high-volume agentic deployments. At current token prices and the cost-per-test benchmark, it becomes the default for coding assistants, retrieval-augmented generation, and agent loops where throughput matters more than frontier reasoning. DeepSeek's recent $7 billion fundraise gives the company balance-sheet room to sustain aggressive pricing as other labs cut rates. The tradeoff is clear: frontier models stay ahead on reasoning and long-horizon agentic work, but gap to V4 Flash is now small enough that routing, caching, and batching—not model choice—become the cost lever.