AivexaNewsSearch
AI news for builders and product teamsChecked every hour

A/B test models in production

Collected Oct 1, 2026

Together AI published a description of endpoint-level A/B testing for production models, positioning it as a way to measure whether a candidate model performs better with real users rather than only on benchmarks. The company states that shadow traffic can verify latency, errors and throughput but discards responses, so no user acts on them, meaning quality questions require real exposure.

According to the company, an A/B experiment attaches to an endpoint and declares members with exactly one control and one or more variants, each pointing at a deployment and carrying a percent setting that must sum to 100. Variant deployments must carry zero weight in the endpoint's base traffic split; only the control lives in that split. Together AI says experiment percents are true fixed traffic shares independent of replica counts, unlike traffic-split weights, which follow capacity.

Ramping is performed by resending the full member list. Updates are guarded by an etag, so Together AI says a concurrent teammate ramp is rejected rather than overwritten. Up to 20 variant members are supported, allowing multi-way tests such as a full-precision endpoint plus quantized variants. The company says platform metrics are available per deployment, while product quality signals must be logged by the user and joined by deployment ID, which appears in response metadata.

Ending an experiment is described as two steps: promote via a blue-green rollout from control to variant, then delete the experiment. If the variant loses, Together AI says deleting the experiment returns 100% of traffic to the control.

On cohorts, the company states assignment uses the request's sampling key, such as a top-level prompt_cache_key or user field, so requests with the same key route consistently; requests without a key are assigned effectively at random per request. Each member deployment autoscales on its own policy. Together AI reports running a full lifecycle on a live endpoint at 95/5, 80/20 and 50/50 before deletion, with roughly 75 seconds of propagation wait after each update, and that 360 consecutive requests after deletion all landed on the control.

Read at Together AI

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Shadow traffic proves a candidate is operationally sound. It can't tell you if users like it better. Run the split at the endpoint instead of in your app code.