Expanding our enterprise inference capacity with IBM Cloud and NVIDIA

Together AI said it is working with IBM and NVIDIA to scale enterprise-grade AI inference, beginning with a large cluster of NVIDIA B300 GPUs on IBM Cloud, backed by NVIDIA Spectrum-X Ethernet networking. According to the company, it is the first dedicated, large-scale inference cluster of its kind on IBM Cloud, and Together AI is the first customer running on it.
The infrastructure is a dedicated NVIDIA B300 GPU cluster purpose-built for inference on IBM Cloud. In the arrangement described, Together AI operates the inference layer, IBM provides the cloud, and NVIDIA supplies the silicon and networking. Together AI said it is planning for the future as token demand increases.
NVIDIA contributes B300 GPUs and Spectrum-X Ethernet networking engineered for high-throughput inference. IBM contributes its experience running mission-critical infrastructure for large enterprises. Together AI contributes its inference platform for running open models in production.
Together AI said it serves hundreds of trillions of tokens per month to over a million developers. The company said enterprises and AI-native companies choose open models because their sovereign data stays theirs and they get frontier-level performance at a fraction of closed-model cost.
Together AI described the result as enterprise-grade inference at massive scale, with the reliability, security and guardrails enterprises expect. The company said open-source AI must run everywhere, at scale, as fast and reliably as closed systems, and called the collaboration a step toward that goal.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Enterprises can now run open models at production scale on a dedicated B300 inference cluster, built by Together AI, IBM Cloud, and NVIDIA