AI news for builders and product teamsUpdated Oct 10, 2026, 19:01 UTC
Together AI
First-party releases and research from Together AI. Headlines and excerpts link to the original articles.
Latest stories
Newest first
Together AI is collaborating with IBM and NVIDIA to launch a dedicated NVIDIA B300 GPU inference cluster on IBM Cloud, backed by Spectrum-X Ethernet networking. Together AI says it is the first customer on the cluster, which it describes as the first dedicated large-scale inference cluster of its kind on IBM Cloud.

Together AI announced Together Link, which connects existing coding agent harnesses such as Claude Code, Claude Desktop, Codex, OpenCode, and Pi to open models including GLM 5.3 and Kimi K3. Together AI says it cuts model spend by over 50%.

Together AI launched together/Tev1-4B-experimental, a classifier built on Qwen3.5 4B, and published a guide to fine-tuning a similar model. The post says training on 38,340 examples costs about $17 and takes roughly 25 minutes.

Together AI described canary rollouts for moving live traffic between model deployments on dedicated inference without downtime, using staged traffic ramps, health checks, metric gates and human-approved pausing or reversal. It reported a Qwen2.5-7B to Qwen3.5-9B canary whose gate caught a 137% p95 regression at 10% of traffic, then was canceled and reversed with live requests served and none failed.

Together AI described how a global fintech ran its coding assistant on GLM 5.2 through Together's Dedicated Model Inference, giving the customer's engineers direct control over scaling, model rollouts, and testing without filing tickets with Together.

Together AI published a five-stage playbook for migrating from closed source to open source models: discover, evaluate, adapt, decide, and production. The company says such migrations can take weeks to months rather than months to years, and reports up to 70% cost reduction in some customer transitions.

Together AI has expanded its Together Fine-Tuning service with support for new open-weight models, live experiment tracking, Expert LoRA adapters, early stopping, dataset previews, pre-flight validation, and training price cuts of 30% to 70% on selected models.

Together AI has launched a public preview of preemptible compute for Together GPU Clusters, offering the same NVIDIA GPU capacity at a flat 50% of the on-demand rate. Preemptible nodes can be reclaimed with a five-minute drain window and are billed sub-hourly.

Together AI ported its ThunderKittens kernels to NVIDIA's Vera Rubin NVL72 platform and rebuilt its NVFP4 GEMM, reporting over 22 PFLOPS, up from 42% of roofline on Blackwell. The work adds support for NVFP4 and FP8 GEMMs on the new hardware.

Together AI published a deep dive into the open model AI stack, describing five independent layers — model, inference, gateways and routers, harness, and tools — and arguing that keeping them separate lets developers swap in a new open model in minutes instead of rebuilding their workflow.

Together AI reports that GLM-5.3 and Claude Fable 5 tied on DeepSWE pass@1 across 904 rollouts, with Fable 5 at 69.7% and GLM-5.3 at 69.0%. GLM-5.3 led pass@2 and pass@4 and cost 5.4x less per rollout, $3.99 versus $21.63.

Together AI compared GLM-5.3 and GPT-5.6 Sol across 904 DeepSWE rollouts on 113 tasks. Sol led pass@1 at 72.7% versus 69.0%, while GLM-5.3 won pass@4 and cost about 2.1x less per rollout; a GLM-first cascade with Sol escalation solved 85.9% of tasks at $6.61 each.

Together AI describes running A/B experiments at the endpoint level, splitting live traffic among one control and up to 20 variants with fixed percentages. It covers ramping via member updates, measuring platform and product metrics per deployment, and ending tests by promoting a rollout or deleting the experiment.

Together AI ran 904 DeepSWE rollouts comparing DeepSeek V4 Pro 0813 and Claude Fable 5. Fable leads pass@1 at 69.7% versus 62.8% but costs 90x more per rollout; Pro wins pass@4, and a Pro-first cascade with Fable escalation solves 82.7% of tasks at $8.28 each.

Moonshot AI's Kimi K3, described as a 2.8-trillion-parameter open-weights model and the first open-source model in the 3-trillion-parameter class, is now served on Together AI via an OpenAI-compatible API. The guide covers KDA and Attention Residuals architecture, Stable LatentMoE sparsity, reasoning effort levels, 1M context, tools, vision, benchmarks, and per-token pricing.

Together AI and academic collaborators introduced ThunderAgent, a program-aware scheduler for agentic inference that treats each agent workflow as a schedulable program. It reports over 2x single-node throughput, 2.4x speedup on an 8-node cluster, and acceptance as an ICML 2026 Spotlight paper.

Together AI announced a strategic partnership with Moonshot AI to natively serve Kimi models, beginning with Kimi K3, a 2.8T parameter sparse Mixture-of-Experts model with native vision and a 1M token context window. Together AI becomes a launch platform for Moonshot's open weights releases, offering day zero access and post-training.

Together AI released an update to its inference platform for running open-weight models in production, with rollout controls, autoscaling, and observability. It also opened a closed beta for custom training, including full-weight and LoRA reinforcement learning and supervised fine-tuning.

Together AI and Y Combinator announced a partnership to launch the first dedicated YC GPU cluster, giving YC portfolio startups access to compute for inference and training with short-term sprints at long-term rates instead of long-term commitments. The cluster is running at full utilization today.

Together AI published an explainer describing what 99%, 99.9% and 99.99% inference uptime tiers require architecturally, including the failure domains each must survive, and listed questions to ask inference providers before committing.