AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Together AI announces strategic partnership with Moonshot AI to natively serve Kimi models

Collected Oct 1, 2026

Together AI announced a strategic partnership with Moonshot AI under which Together AI becomes a launch platform for Moonshot's model releases, starting with Kimi K3 and extending to every open weights model Moonshot ships going forward.

Kimi K3 is described as the largest open model released to date: a 2.8T parameter sparse Mixture-of-Experts model with native vision support and a 1M token context window. Moonshot built it around two new architectural components: Kimi Delta Attention (KDA), which changes how information flows across sequence length and delivers significantly faster decoding at long context lengths, and Attention Residuals (AttnRes), which improves how representations are retrieved across model depth at minimal extra compute cost. Layered on top are optimizer refinements including per-head Muon and quantile-based expert load balancing, which Moonshot says deliver roughly 2.5x better scaling efficiency compared to Kimi K2. Together AI says benchmark results place it in direct competition with leading proprietary systems.

Kimi models are available across Together AI inference products including Serverless, Provisioned Throughput and Dedicated Inference. Provisioned Throughput offers reserved inference capacity with token-based pricing and a 99% uptime SLA; Dedicated Model Inference is offered for teams wanting more control.

Developers can post train Kimi models with custom training including full-weight and LoRA reinforcement learning and supervised fine-tuning, configured through the Python SDK, with multiple LoRA experiments run concurrently on dedicated capacity. Checkpoints can be deployed natively to inference without a separate handoff between training and serving.

Together AI says day zero availability applies to Kimi K3, validated by Moonshot, and that future Moonshot releases will also ship on Together at launch. Together customers can use Kimi models as a base model for fine-tuning, with flexibility to remove the attribution required in the MIT license. Together AI states its research-optimized inference stack is tuned for large sparse MoE models like K3, and that it already serves production traffic for companies including Cursor, Y Combinator and Decagon.

Read at Together AI

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Together AI partners with Moonshot AI to natively serve Kimi models, starting with the 2.8T parameter Kimi K3, with day zero access and post-training.