AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Together AI brings Thinking Machines Lab’s new model Inkling on day 0

Collected Oct 1, 2026

Thinking Machines Lab released Inkling, a new multimodal mixture-of-experts model built for token-efficient reasoning, native multimodal understanding, and broad task versatility. Together AI said it is collaborating with the Thinking Machines Lab team to make Inkling available to developers on its inference platform on day zero.

Inkling accepts text, image, and audio inputs and produces text outputs through a unified decoder architecture, according to Together AI. It supports controllable inference effort, letting developers adjust how much reasoning the model applies per task. Together AI said the model's post-training spans scientific reasoning, coding, agentic workflows, forecasting, and calibrated prediction.

Together AI described Inkling as a decoder-only mixture-of-experts model with 975B total parameters, 40B active parameters per token, and a 1M-token context window. It incorporates token position into attention through a learned, query-conditioned relative bias rather than RoPE or absolute position embeddings, and mixes sliding-window and full causal attention, with five local-attention layers followed by one full-attention layer.

The architecture also introduces sconv, a channelwise causal convolution with a four-token receptive field applied to key and value streams and to attention and feed-forward sublayer outputs. Its feed-forward layers use a mixture-of-experts design with a shared expert sink that normalizes shared and selected routed experts together. Lightweight embedding towers convert image patches and quantized audio features into embeddings of the same width as text tokens.

On Together AI, Inkling runs with an optimized FlashAttention-4-based attention kernel designed to support its query-conditioned relative attention mechanism in production. Together AI said preliminary evaluations of the current Inkling checkpoint show strong results across scientific reasoning, mathematics, coding, agentic, vision, and audio benchmarks at its highest evaluated effort setting.

Inkling is available through Together AI Serverless with a 1M context window and OpenAI-compatible APIs, which Together AI said handles text, image, and audio inputs through a single API call.

Read at Together AI

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Together AI offers day zero access to Inkling, Thinking Machines Lab's multimodal mixture-of-experts model for text, image, and audio reasoning.