AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Canary rollouts: upgrade models in production without downtime

Collected Oct 1, 2026

Together AI published a description of canary rollouts, a mechanism for moving live traffic between two deployments on the same endpoint while it is serving requests: a source deployment and a target deployment. A rollout is created in a PENDING state and does nothing until explicitly started; the CLI's rollout command creates and starts in one step.

Three strategies are offered. Canary steps traffic through shares defined by the operator (default ladder 5% to 25% to 50% to 100%), each held for a wait window with optional metric checks. Blue-green is a single gated 0% to 100% cutover. Rolling swaps replica by replica while preserving total capacity and is described as best for same-model config changes.

Every step runs in a stated order: scale the target, run health checks, shift traffic, wait 30 seconds for routing to converge, drain the source, wait, evaluate the gate, then record the step. Health checks run before traffic moves, and each side is kept at the replica count its traffic share needs. The platform states capacity is never rounded down.

Metric gates apply to canary only and evaluate three router-side metrics: router_error_rate, router_latency and inflight_requests. Each rule uses a relative regressionCheck or an absolute thresholdCheck. The documented run was a Qwen2.5-7B to Qwen3.5-9B canary on one H100 per deployment whose gate caught a 137% p95 regression at 10% of traffic; it was canceled and reversed with live requests served and none failed.

A tripped gate routes to SYSTEM_PAUSED by default, holding the current split for human review. Together AI states there is no automatic abort: a confirmed regression always parks the rollout for a person. Before pausing on a regression the system re-queries for about 90 seconds, and a gate without trustworthy data pauses with METRICS_UNAVAILABLE rather than counting as a regression. Recoverable causes such as capacity shortfall are retried every 15 minutes for up to three hours. PAUSED, set by an operator, is never auto-resumed.

Operators can pause, resume, promote, or cancel. Cancel freezes the current traffic split and leaves both deployments serving; there is no rollback verb, so going back means creating a new rollout with source and target swapped. Rollouts end COMPLETED or CANCELED. The CLI ships in the together Python package version 2.34.0 or newer, and a REST API and Python SDK are also documented.

Read at Together AI

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

A hard model swap exposes every user at once, and rolling back means cold-starting the old deployment under pressure. Here's how staged traffic ramps, metric gates, and automatic rollback work on dedicated inference.