IBM and Together AI Ink $240M Deal to Revolutionize Open-Source AI Inference
In a move that signals a massive shift in how enterprises approach artificial intelligence, IBM has announced a significant $240 million deal with neocloud provider Together AI. This partnership isn't just a simple infrastructure agreement; it’s a strategic play to bring large-scale, open-source AI inference workloads into the mainstream through IBM Cloud. By combining IBM’s enterprise-grade infrastructure with Together AI’s specialized platform, the two companies are positioning themselves as the go-to destination for developers who want to scale AI without being locked into proprietary silos.
The Powerhouse Infrastructure Behind the Deal
At the heart of this collaboration lies some of the most advanced hardware currently available on the market. The service will be powered by a massive cluster of Nvidia HGX B300 systems, all residing within the IBM Cloud ecosystem. To ensure that data moves as quickly as the processors can handle it, the setup utilizes Nvidia Spectrum-X Ethernet networking. This high-performance networking fabric is specifically designed to minimize latency and maximize throughput for AI workloads, which is a critical requirement when you're dealing with the massive parameter counts of modern large language models.
Together AI will leverage this high-octane cluster to offer its specialized inference services. By running their platform on IBM’s bare-metal and virtualized environments, they can provide a reliable, high-performance foundation for companies looking to move their AI projects from the experimental phase into full-scale production.
Why Open Source is Winning the Enterprise Debate
One of the most compelling aspects of the Together AI platform is its support for a diverse range of open-source models. The platform currently supports heavyweight names like DeepSeek, Nemotron, MiniMax, Kimi, and GLM. According to Together AI, this variety gives developers and corporate IT departments the freedom to customize and fine-tune models to their specific needs.
More importantly, it offers a path toward lower costs. Instead of being tethered to the pricing and limitations of a single proprietary model provider, enterprises can use these open-source alternatives to build tailored applications at a fraction of the cost. The goal here is clear: provide GPU infrastructure and AI services that make open-source the obvious, most cost-effective choice for the modern enterprise.
Scaling from Pilot to Production
IBM and Together AI are focusing heavily on the reliability factor. A hybrid environment on IBM Cloud, backed by Nvidia’s latest GPUs and Spectrum-X networking, provides the stability that large organizations demand. This agreement also marks another chapter in the deepening relationship between IBM and Nvidia. The two giants recently expanded their collaboration to focus on GPU-native data analytics, intelligent document processing, and regulated infrastructure deployments.
Vipul Ved Prakash, the CEO of Together AI, highlighted the importance of this foundation. He noted that the collaboration will accelerate their expansion into the enterprise sector, making production-grade inference more accessible than ever. "This cluster lets us bring production-grade inference to more companies, faster," Prakash stated, emphasizing that this is a major step in making open-source AI the primary choice for businesses worldwide.
The Shift from Training to Inference
This deal arrives at a pivotal moment in the AI market. While much of the initial AI hype focused on training massive models, the industry is now pivoting toward inference—the process of actually running those models to generate results. Gartner recently released a forecast predicting that worldwide spending on AI-optimized Infrastructure as a Service (IaaS) will grow by a staggering 96% through 2026, reaching a total of $42 billion. By 2027, that figure is expected to hit $66 billion.
Your brand deserves a better website.
We don't just use templates. We build custom web apps, landing pages, and company profiles designed specifically for what you need.
Perhaps the most telling statistic from Gartner is the shift in spending priorities. In 2026, global spending on inference is expected to reach $23.3 billion, officially surpassing the $19 billion projected for model training. This shift is driven by the rise of "agentic AI"—autonomous systems that perform multi-step tasks—which significantly increases the demand for continuous, real-time compute power.
A Competitive Landscape
Hardeep Singh, a senior principal research analyst at Gartner, points out that as organizations move toward production-scale deployment, they are increasingly relying on domain-specific models (DSMs) that require constant execution rather than periodic training. This creates a sustained demand for AI-optimized infrastructure, which is exactly what the IBM and Together AI partnership aims to provide.
As the AI IaaS market expands, IBM Cloud is finding itself in a fierce battle for dominance against incumbents like AWS, Google Cloud, and Microsoft Azure, as well as specialized players like Coreweave and Lambda. By doubling down on open-source inference and high-end Nvidia hardware, IBM is making a clear bet that the future of enterprise AI lies in flexibility, performance, and the ability to scale inference workloads efficiently.