IBM and Together AI signed a $240 million multi-year agreement to deploy a large-scale inference cluster on IBM Cloud using Nvidia HGX B300 systems (Blackwell architecture) paired with Spectrum-X Ethernet networking. The initial U.S.-based cluster will feature approximately 2,000 Blackwell 300 chips and become available in H1 2027. Together AI's CRO said the capacity will 'sell out at least two to three months ahead of time.'
The deal reflects a structural shift in AI economics from training to inference: as enterprises move AI from proof-of-concept to production, inference—the process of running trained models to generate responses—is becoming the primary driver of compute demand. Together AI currently processes 400 trillion tokens monthly and focuses on open-source models (DeepSeek, MiniMax, Kimi) that enterprises view as cheaper and less risky than closed systems from OpenAI, Anthropic, or Meta.
IBM is positioning itself as the infrastructure intermediary for enterprises seeking open-source AI and cost efficiency. The deal reflects growing cybersecurity skepticism toward closed models and data residency concerns, particularly among regulated industries (banks, hospitals). Together AI, valued at $8.3B in July, represents a bet that open inference at scale will capture enterprises unwilling to pay premium pricing for proprietary models.