AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Kimi K3: the complete developer guide

Collected Oct 1, 2026

Moonshot AI has released Kimi K3, which Together AI describes as a 2.8-trillion-parameter model, the largest open-weight model released to date, and the first open-source model in the 3-trillion-parameter class. According to Together AI, Kimi K3 is positioned for long-horizon coding, end-to-end knowledge work, and deep reasoning, and Together AI says it is working directly with the Moonshot team to serve the model.

Two architectural changes are described as K3's backbone: Kimi Delta Attention (KDA), a hybrid linear attention mechanism, and Attention Residuals (AttnRes), which selectively retrieves representations across model depth. The first Kimi model to support a 1M context length, K3 also uses the Stable LatentMoE framework, activating 16 of 896 experts, roughly 2% per token. Supporting training techniques listed include Quantile Balancing, Per-Head Muon, Sigmoid Tanh Unit (SiTU), and Gated MLA.

On Together AI the API is OpenAI-compatible and uses the official Together Python SDK. Reasoning is configurable through a top-level reasoning_effort field with low, high, and max levels, max being the default; thinking can be switched off via reasoning={"enabled": False}. Streaming returns separate reasoning_content thinking traces and final-answer content. Multiple images are accepted as input, with no limit on image count but a request body under 100 MB and recommended images up to 4096x2160. Structured output uses response_format with json_schema and strict: true, or the looser json_object mode. Standard tool calling is supported, including dynamic tool loading via a system message carrying a tools field, and automatic context caching.

Sampling parameters are fixed (temperature 1.0, top_p 0.95, n 1, presence_penalty 0, frequency_penalty 0), and K3 was trained in preserved thinking history mode. Together AI states that across its evaluation suite K3 leads several coding and agentic benchmarks including SWE Marathon, BrowseComp, DeepSearchQA, AutomationBench, and OmniDocBench, remains competitive with the strongest proprietary models elsewhere, and outperforms GLM-5.2, while trailing Claude Fable 5 and GPT 5.6 Sol on a handful of benchmarks, consistent with Moonshot's positioning. All listed K3 results used reasoning effort set to max.

Read at Together AI

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Kimi K3 is the first open 3T-class model. See how it benchmarks, what it costs, and how to call it on the Together AI API, with copy-paste code examples.