How to Build a Model Router in the Harness

LangChain described building a model router into the harness of Open SWE, its open source coding agent, in a post explaining how to build your own. Compared with a baseline of always using a top-tier frontier model, the router cut median cost per thread by 64% with no measurable change in quality.
The router picks one of three models at the start of a thread and uses it for the whole thread: GLM-5.3-Flash for fast tasks, GPT-5.6 Sol for balanced tasks, and GPT-6 Astra for performance tasks. It runs on the thread's first human message and has three parts: a base prompt, criteria per tier, and a classifier model. LangChain said the first version used an LLM with structured output; the classifier now runs on Jev, a newly released decision model, making classification almost 50 times faster.
In a first A/B test over 973 threads, half routed and half always using GPT-6 Astra, 29.2% of routed threads ended in a merged PR versus 27.3% of control (p = 0.49), and PR open rates were 38.9% versus 39.6% (p = 0.82). Median routed thread cost was $0.94 versus $2.61 on control, a 64% reduction; the mean dropped 42% and the p90 dropped 37%. Of routed threads, 56% went to balanced, 34% to fast, and 10% to performance. Median thread cost was $0.097 on fast, $1.50 on balanced, and $2.88 on performance. LangChain said a second A/B test against a fast-only control was ended within a day before producing statistically meaningful results because engineers flagged problems with the fast-only arm.
LangChain said the router belongs in the harness rather than a generic gateway because choosing a model requires the domain and task context the harness already assembles. It also released model routing middleware so developers can supply a base prompt, model tiers, and criteria per tier.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
How we built a model router into Open SWE's harness that cut median cost per coding task by 64% with no measurable drop in quality, and how to build your own.