TL;DR · 30-second read
The Short Version
IBM, the century-old computing company, has signed a $240 million, multi-year deal with Together AI, a younger firm that runs freely available artificial intelligence programs for other businesses.
IBM’s cloud, meaning its rentable computers in data centers, will host a large set of powerful Nvidia chips set aside for one job: answering people’s questions to those programs, the everyday work of a chatbot.
The surprising part is how it is bought. Instead of paying per use, the computing power is booked years ahead, like leasing a whole building instead of renting hotel rooms by the night.
IBM and Together AI have signed a $240 million multi-year agreement for an Nvidia-powered AI inference cluster that will run on IBM Cloud, Reuters reported on August 11, 2026. Inference is the stage where a trained AI model answers requests, as opposed to training, where the model is built.
Together AI operates a cloud platform that serves open-weight AI models, meaning models whose parameters are publicly released, to developers and enterprises. Under the deal, a dedicated cluster of Nvidia accelerators on IBM’s infrastructure is committed to that inference work for several years.
Executive Summary
The headline number is $240 million, but the structure matters more. This is not a customer swiping a card for GPU hours. It is a multi-year commitment to a specific, dedicated cluster, which is how data center space and power have long been contracted, and increasingly how AI compute is contracted too.
For IBM, the deal places IBM Cloud in the business of hosting third-party AI inference at scale, a market where the large hyperscalers and specialist GPU clouds have been most visible. For Together AI, it secures a block of capacity for serving open models at a time when access to Nvidia hardware still depends heavily on who has committed to it in advance.
The companies have not disclosed the hardware generation, the size of the cluster, where it sits, or how the $240 million is paid over time. Those details determine whether this is a modest capacity reservation or a meaningful new source of inference supply.
Inference Is Being Booked Like a Lease
Most people meet AI inference as a metered service: a developer sends a prompt, a provider returns an answer, and the bill is calculated per token, a token being a small chunk of text. That pricing looks elastic, but the hardware underneath it is not. Nvidia accelerator clusters are expensive to buy, take time to deliver and install, and lose value as each new chip generation arrives. Someone has to commit to them before the first token is served.
A $240 million, multi-year contract for a dedicated cluster is that commitment made explicit. It shifts inference from the on-demand model of cloud computing toward the reserved-capacity model long familiar from colocation and wholesale data center leasing, where a tenant contracts a defined block of infrastructure for a defined term. The provider gets revenue visibility that justifies deploying hardware; the tenant gets guaranteed supply it can resell to its own customers with predictable performance.
This is one deal, and it does not prove an industry-wide shift on its own. But it shows clearly that open-model inference, often described as a commodity service racing to the lowest per-token price, is being underwritten by the same kind of long-dated capacity contracts as training clusters. The per-token price a developer sees sits on top of a fixed, multi-year obligation somewhere in the stack.
What IBM Gains by Hosting Someone Else’s Models
IBM has spent recent years positioning around open and hybrid AI: its Granite models are openly released, and its Red Hat business is built on open-source software. Hosting a platform whose entire proposition is serving open-weight models fits that posture more naturally than competing head-on with the largest clouds for frontier-lab training contracts.
It also gives IBM Cloud a substantial, recurring AI workload with an anchor tenant. For a cloud provider, a committed customer reduces the risk of buying accelerators that might otherwise sit underused. For enterprise buyers who already run on IBM infrastructure, the arrangement could, in principle, place open-model inference closer to their existing systems, though neither company has described how enterprise customers would access the cluster.
The Risk Sits in the Utilisation Curve
Reserved capacity moves risk; it does not remove it. Inference demand can shift quickly as new models arrive, and per-token prices for open models have been falling as providers compete. A multi-year commitment to a fixed pool of hardware exposes whichever party carries it to two pressures: whether the cluster stays busy, and whether the hardware stays competitive as Nvidia releases newer, more efficient generations.
There is also a physical side. Dense Nvidia clusters draw significant power and increasingly require liquid cooling, which constrains which data halls can host them and how quickly. The value of a contract like this depends on the cluster actually being energised and in service on schedule, and none of those milestones have been published.
Background
IBM has operated a public cloud for more than a decade and in recent years has concentrated its AI strategy on enterprise and hybrid deployments, including its watsonx platform, its openly released Granite models and the open-source software business of Red Hat. IBM Cloud offers Nvidia GPU infrastructure alongside its traditional enterprise services.
Together AI emerged during the generative AI boom as a platform for running open-weight models, giving developers an alternative to closed models from the largest AI labs. Like other inference providers, it depends on securing large volumes of Nvidia accelerators, whose supply has been shaped by long-term commitments from cloud providers and AI companies. Source: IBM, Together AI ink $240 million deal for Nvidia-powered AI inference cluster (Reuters): report on a multi-year agreement for a dedicated Nvidia inference cluster on IBM Cloud.Sources

