AivexaNewsSearch
AI news for builders and product teamsChecked every hour

The production platform for open-weight AI inference

Collected Oct 1, 2026

Together AI announced a significant update to its inference platform for running open-weight models in production, along with a closed beta for custom training. The company said the platform is designed to give users control over performance, cost, and quality without building their own inference stack.

The update supports deploying models from Together's model platform or bringing your own open-weight or fine-tuned model, with full weights or adapters uploaded from Hugging Face, S3, or a local machine. Deployments can be pinned to a specific region or left flexible, and autoscaling can be driven by inflight requests, GPU utilization, time to first token, latency, decoding speed, or throughput. Multiple deployments can sit behind one stable endpoint, with canary, blue-green, and rolling updates that auto-roll-back on user-defined thresholds, plus A/B and shadow testing against real traffic.

Together said it rebuilt its model caching and distribution layer so weights can be shared across the fleet and proactively warmed, delivering roughly 4× faster warm starts across several frontier models in internal testing. The platform also provides an organization-level, scrapeable Prometheus endpoint for observability.

The custom training closed beta includes full-weight and LoRA reinforcement learning and supervised fine-tuning, configurable through the Python SDK, with checkpoints deployable natively to inference deployments. Access is limited during the beta.

Max Lu of Decagon said Together had already met its latency needs for voice, and that the new platform lets the company canary a new fine-tuned model on a small percentage of live traffic and automatically roll back if key metrics regress. Together said it serves more than 400 trillion tokens per month and will host a webinar on August 6.

Read at Together AI

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Run open models in production with full control over performance, cost, and quality. Deploy in minutes, roll out safely, and scale to your SLOs.