AI news for builders and product teamsUpdated Oct 10, 2026, 20:01 UTC
Together AI
First-party releases and research from Together AI. Headlines and excerpts link to the original articles.
Latest stories
Newest first
Thinking Machines Lab released Inkling, a multimodal mixture-of-experts model for text, image, and audio reasoning. Together AI is offering day zero access to Inkling on its inference platform, including serverless availability with a 1M context window and OpenAI-compatible APIs.

Together AI announced updates to its GPU Clusters covering passive health checks, auto node repair, a rebuilt Slurm-on-Kubernetes stack (Slinky 1.0), external OIDC for Kubernetes RBAC, startup scripts, a new cluster details view, and an acceptance-test opt-out.

Together AI announced an $800 million Series C round from investors including Aramco Ventures, NVIDIA, Vista Equity, and General Catalyst. The company also secured commitments for over 500 MW of compute capacity, to be capitalized independently by its new investors.

Together AI says it will present nine papers at ICML 2026, spanning agent evaluation, model training, inference optimization, and GPU kernels, with some research shipping in its production platform. The company will be at booth B714 in Seoul from July 6 to 11.

Together AI has released ParallelKernelBench, a benchmark testing whether LLMs can write fast multi-GPU CUDA kernels across 87 real workloads. The best model solves fewer than a third of problems, but a few generated kernels beat any public implementation.

Together AI compared Kimi K2.7 Code and Claude Fable 5 on 12 generated landing pages, reporting Kimi cost about 94% less on average while scoring within a few points on nearly every page. A custom design-inspiration MCP server improved Kimi's output quality.