Together AI ships DeepSeek V4 Pro with 512K context and cached input pricing
Together AI now hosts DeepSeek V4 Pro, the 1.6-trillion-parameter mixture-of-experts model, with 512K-token context and three controllable reasoning modes (Non-Think, Think High, Think Max). The model uses hybrid Compressed Sparse Attention and Heavily Compressed Attention, reducing single-token inference FLOPs to 27% and KV cache to 10% versus DeepSeek V3.2 at million-token context. Pricing is $2.10 per 1M input tokens, $0.20 per 1M cached input tokens, and $4.40 per 1M output tokens.
DeepSeek V4 Pro benchmarks at 93.5% on LiveCodeBench, 90.1% on GPQA Diamond, 80.6% on SWE-Bench Verified, and 83.5% on MRCR 1M comprehension. The 1.6T parameter count activates only 49B parameters per forward pass, giving the model frontier-level knowledge capacity while keeping inference costs comparable to a 49B dense model. Cached input pricing provides 90% cost reduction for repeated analysis over the same large context—critical for code agents, document intelligence, and long-horizon agentic workflows.
For teams building long-context reasoning systems, DeepSeek V4 Pro on Together AI removes the operational burden of running a trillion-parameter MoE locally while maintaining serverless flexibility or moving to dedicated, reserved capacity for production SLA guarantees. Workloads like repository analysis, policy comparison, and multi-step agentic decision-making can now leverage million-token context at hosted-inference pricing. Architects evaluating open-weight inference hosting should benchmark V4 Pro's reasoning modes and cached-input cost profile against Groq, Fireworks, and direct cloud provider inference endpoints.