Together AI and IBM signed a multi-year, $240 million infrastructure deal on August 11 to build a large-scale open-source AI inference cluster on IBM Cloud powered by NVIDIA HGX B300 systems and Spectrum-X Ethernet networking. The cluster is the first dedicated, large-scale inference facility on IBM Cloud using this hardware configuration; availability is targeted for Q1 2027. Together AI reported serving 400 trillion tokens monthly and recently closed an $800 million Series C at an $8.3 billion valuation.
Together AI claims #1 inference speed rankings across open-source models on NVIDIA Blackwell architecture, achieving up to 2x faster output speed than competing providers for Qwen, DeepSeek, and Kimi models. The company attributes performance gains to a fully modernized inference engine optimized for Blackwell hardware, advanced speculative decoding (ATLAS), and low-bit quantization (FP4/FP8). For DeepSeek-V4-Flash, Together AI delivers 13% faster performance than the next fastest provider. Together also secured commitments for 500+ megawatts of compute capacity independently capitalized by investors to support anticipated growth.
For engineers and architects evaluating open-source inference providers: Together AI's IBM partnership and Blackwell optimization represent a meaningful shift in commodity inference economics. The 2x speedup claims (where validated on your workload) translate directly to lower token costs; at scale, this compounds. However, the Q1 2027 launch date means existing deployments should evaluate performance on their own models and traffic patterns today. The implicit market message is that inference cost curves continue falling fastest for open-weight models with purpose-built kernels—and that hyperscalers and platform providers (IBM, NVIDIA, Together) are betting heavily on this trend to drive cloud compute adoption away from proprietary APIs.