AivexaNewsSearch
AI news for builders and product teamsChecked every hour

The Open Source AI Stack

Collected Oct 1, 2026

Together AI has published a deep dive into the open model AI stack, aimed at developers and organizations moving from closed to open source models for what the post describes as more ownership, control and economics. The post states that using open models for agentic software development does not require training models, buying GPUs, or becoming a machine learning expert.

The stack is broken into five parts: the model, which interprets requests and decides what to do; inference, the infrastructure and provider where the model runs; gateways and routers, which decide which model or provider handles each request based on cost, speed and capability; the harness, the application that manages the conversation, gives the model tools and connects it to the codebase; and tools, described as skills and MCP, which supply knowledge for specific tasks. The post states these layers are independent, allowing decisions at each layer and experimentation with new models as they are released.

On models, the post distinguishes large from small by how much ambiguity they can handle rather than quality. It cites Kimi K3 as a large open model with 1.8T total parameters and 104B active parameters, and GLM 5.3 Flash as a small open model with 320B total and 18B active parameters, described as roughly 6 times smaller and 20 times cheaper. It notes most leading open models are Mixture-of-Expert models that activate a subset of experts per token. Other popular open models named are DeepSeek V4 Flash and MiniMax M3, with the post noting these change rapidly. Leaderboards cited for discovery include The Open Frontier and Artificial Analysis.

On inference, the post says cloud providers run models on their GPUs and charge per token, and that two providers running the same model should generally produce similar results with the same version and sampling settings, though performance, price, latency and API features may differ. Gateways named include OpenRouter and Vercel AI Gateway, with LiteLLM for self-hosted routing. Harnesses named with open model support include PI, OpenCode and Amp; the post says closed-source harnesses such as Claude Code and Codex can be connected to open-weight models via TogetherLink.

For tools, the post describes skills as reusable instructions loaded by the harness and MCP as a standard for connecting agents to external tools and data sources, pointing to skills.sh and mcp.so. It also covers context management, recommending starting new sessions when switching tasks or models, and suggests a plan, implement, review workflow split across three models. The post concludes that swapping a model is often a configuration change, and switching providers is pointing the same harness at a different API endpoint.

{"category": "Developer tools"}

Read at Together AI

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

A deep dive into the open model AI stack — model, inference, gateways and routers, harness, and tools — and how keeping each layer independent lets you swap in a new open model in minutes instead of rebuilding your workflow.