Together AI and IBM Partner on $240 Million NVIDIA B300 Inference Infrastructure
Together AI signed a $240 million multi-year agreement to deploy a large-scale inference cluster on IBM Cloud, built on NVIDIA HGX B300 systems and Spectrum-X Ethernet networking.
Shane Barrett·updated August 23, 2026

According to reporting by Pulse 2.0, the cluster is scheduled to come online in Q1 2027 and represents the first dedicated large-scale inference deployment of its kind on IBM Cloud infrastructure. The contract signals a deliberate restructuring of open-source inference economics rather than a routine capacity expansion.
Hardware configuration and throughput claims
The infrastructure pairs HGX B300 GPU systems with Spectrum-X Ethernet fabric, a combination NVIDIA positions as capable of delivering up to 30x more AI factory output than prior generations. The figure originates from vendor materials bundled with the announcement; no independent benchmark has yet been published to validate it. For practitioners evaluating the deployment, the operationally relevant variables are tokens-per-watt, tokens-per-GPU-hour, and tail-latency under concurrent multi-tenant load, none of which are disclosed in current materials. The HGX B300 platform itself is expected to reach broader availability around the same Q1 2027 window, placing Together AI's cluster among the early production testbeds for this silicon generation. Memory bandwidth, FP8 throughput, and interconnect topology will determine whether the claimed gains hold in real serving scenarios.
Together AI scale and capital context
Together AI reports that its inference stack currently serves approximately 400 trillion tokens per month, establishing a baseline against which incremental capacity from the new cluster can be measured. The company's $800 million Series C round, closed at an $8.3 billion valuation, indicates that infrastructure expenditure of this magnitude aligns with existing operational scaling rather than speculative capacity buildout. Together AI operates an AI Native Cloud spanning inference, training, fine-tuning, and agentic workflows; the IBM-hosted cluster slots specifically into the inference tier. Selection criteria cited in the announcement emphasize GPU capacity at scale and roadmap alignment over peak throughput ceilings. The deal also extends IBM's broader NVIDIA partnership, which already spans GPU-native analytics, unstructured data extraction, and hybrid on-premises deployments. The agreement lands within a wider pattern of NVIDIA capital commitments directed at AI infrastructure buildout, as documented in recent financing arrangements aimed at accelerating global AI infrastructure.
Variables to monitor
Three data points warrant tracking as the Q1 2027 deployment approaches. First, public benchmarks comparing B300-based inference against incumbent H100 and H200 clusters on standard open-source models at matched precision and batch sizes. Second, token pricing disclosures from Together AI, which will indicate whether the claimed throughput improvements translate into measurable reductions in per-token cost. Third, integration depth between Together AI's inference stack and IBM Cloud's enterprise tooling, particularly watsonx-adjacent orchestration and consulting services. Until benchmark numbers are published, the headline 30x throughput figure remains a hypothesis awaiting empirical validation rather than an established result.