AivexaNewsSearch
AI news for builders and product teamsChecked every hour

How to Build a Model Router in the Harness

Collected Oct 1, 2026

LangChain described building a model router into the harness of Open SWE, its open source coding agent, in a post explaining how to build your own. Compared with a baseline of always using a top-tier frontier model, the router cut median cost per thread by 64% with no measurable change in quality.

The router picks one of three models at the start of a thread and uses it for the whole thread: GLM-5.3-Flash for fast tasks, GPT-5.6 Sol for balanced tasks, and GPT-6 Astra for performance tasks. It runs on the thread's first human message and has three parts: a base prompt, criteria per tier, and a classifier model. LangChain said the first version used an LLM with structured output; the classifier now runs on Jev, a newly released decision model, making classification almost 50 times faster.

In a first A/B test over 973 threads, half routed and half always using GPT-6 Astra, 29.2% of routed threads ended in a merged PR versus 27.3% of control (p = 0.49), and PR open rates were 38.9% versus 39.6% (p = 0.82). Median routed thread cost was $0.94 versus $2.61 on control, a 64% reduction; the mean dropped 42% and the p90 dropped 37%. Of routed threads, 56% went to balanced, 34% to fast, and 10% to performance. Median thread cost was $0.097 on fast, $1.50 on balanced, and $2.88 on performance. LangChain said a second A/B test against a fast-only control was ended within a day before producing statistically meaningful results because engineers flagged problems with the fast-only arm.

LangChain said the router belongs in the harness rather than a generic gateway because choosing a model requires the domain and task context the harness already assembles. It also released model routing middleware so developers can supply a base prompt, model tiers, and criteria per tier.

Read at LangChain

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

How we built a model router into Open SWE's harness that cut median cost per coding task by 64% with no measurable drop in quality, and how to build your own.