AivexaNewsSearch
AI news for builders and product teamsChecked every hour

What does 99.9% uptime mean for inference?

Collected Oct 1, 2026

Together AI published an explainer on what different uptime tiers mean for inference, arguing that each reliability level maps to a specific failure domain and requires its own architecture to survive it.

In Together AI's account, 99% uptime means an architecture can survive node-level failures such as GPU hardware faults, driver crashes and thermal events, generally through automated health checking, node draining and fast replica replacement within a single data center. The company says the ceiling at that tier is the building itself, citing thermal issues, substation events or edge router failures.

99.9% means surviving a full data center failure, which Together AI says usually requires model weights deployed across two facilities, enough capacity on each side to absorb full load, and live traffic routing to both rather than a cold standby. The company says it chose continuous routing to both facilities and argues that providers renting capacity from a hyperscaler or neocloud do not own their failure domains.

99.99% means surviving a regional outage, which Together AI says typically calls for multi-region deployment with availability zone redundancy and reserved failover capacity sized to absorb a full regional outage.

The post attributes differences in GPU and CPU failure modes and says GPU inference failure modes span compute, network, storage and software layers. It lists ECC errors in VRAM as the most common issue Together AI sees, saying they corrupt weights silently. It also says availability and performance are different contracts, and that under Provisioned Throughput a service delivering 30% of contracted throughput is not meeting the deal.

Together AI says it runs inference for teams including Cursor, Decagon, Cartesia and Yutori, and that it measures uptime at inference completion rather than at the gateway, counting a request that reaches the load balancer but fails at the GPU as downtime. The post includes questions it suggests asking providers about infrastructure ownership, full-stack observability, failover testing and how each SLA tier is measured.

Read at Together AI

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Reliability numbers are easy to publish. We break down what 99%, 99.9%, and 99.99% uptime actually require, the failure domains each tier has to survive, and the questions to ask any inference provider before you commit.