GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

Together AI reported results from a head-to-head run of GLM-5.3 (max) against GPT-5.6 Sol (max) on all 113 DeepSWE tasks, four trials per configuration, totaling 904 rollouts, 452 per model. DeepSWE is described as a benchmark testing software engineering ability across task types and programming languages.
Under DeepSWE's official scoring, GPT-5.6 Sol led single-shot pass@1 at 72.7% against GLM-5.3's 69.0%, a 3.7-point gap. With retries, GLM-5.3 tied Sol at pass@2 (81.1% versus 81.0%) and led pass@4 (87.6% versus 85.8%).
Cost was reported at $3.99 per rollout for GLM-5.3 against $8.37 for Sol, about 2.1x lower. Per $100 spent, the report said GLM-5.3 solved 17 tasks and Sol solved 9. Sol averaged 19 minutes and 61 steps per rollout, against GLM-5.3's 35 minutes and 124 steps.
On reliability, Sol solved 61 tasks four for four, versus 48 for GLM-5.3, with reliability at 84.5% and 78.8% respectively. Sol broke existing tests in 20% of its failures, against 11% for GLM-5.3.
Task-type results split four apiece. Sol led data modeling and serialization (92%), build and ops tooling (73%), concurrency and durability (72%), and protocol conformance (59%). GLM-5.3 led query and config languages (88%), language and runtime internals (83%), and stateful reactivity (73%), and tied program analysis at 64. By language, GLM-5.3 led JavaScript 90 to 75 and Rust 70 to 60; Sol led Python 74 to 66, Go 79 to 76, and TypeScript 66 to 61.
Per-task correlation was 0.43. Each model solved 90 tasks; GLM-5.3 alone solved 9, Sol alone 7, and 7 defeated both. Their union covered 106 of 113 tasks (93.8%). A GLM-first cascade that escalates to Sol when tests fail solved 85.9% of tasks at $6.61 each; Sol alone solved 72.7% at $8.37. The report stated a perfect one-shot oracle router reached 83.8%.
Method notes: data came from a DeepSWE v1.1 export of published per-trial records. Costs are the published per-trial cost_usd. Per-turn trajectory JSONs for the GLM-5.3 batch were not on the public CDN at analysis time, so the analysis is index-level. Task types were classified by an LLM from benchmark prompts. The cascade assumes a verifier decides escalation.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
We ran 904 DeepSWE rollouts on GLM-5.3 and GPT-5.6 Sol. Sol leads pass@1 by 3.7 points; GLM-5.3 wins pass@4 at half the cost, and a GLM-first cascade hits 85.9%.