AivexaNewsSearch
AI news for builders and product teamsChecked every hour

GLM-5.3 vs. GPT-5.6 Sol on DeepSWE: Cost, Coding, and Routing

Collected Oct 1, 2026

Together AI reported results from a head-to-head run of GLM-5.3 (max) against GPT-5.6 Sol (max) on all 113 DeepSWE tasks, four trials per configuration, totaling 904 rollouts, 452 per model. DeepSWE is described as a benchmark testing software engineering ability across task types and programming languages.

Under DeepSWE's official scoring, GPT-5.6 Sol led single-shot pass@1 at 72.7% against GLM-5.3's 69.0%, a 3.7-point gap. With retries, GLM-5.3 tied Sol at pass@2 (81.1% versus 81.0%) and led pass@4 (87.6% versus 85.8%).

Cost was reported at $3.99 per rollout for GLM-5.3 against $8.37 for Sol, about 2.1x lower. Per $100 spent, the report said GLM-5.3 solved 17 tasks and Sol solved 9. Sol averaged 19 minutes and 61 steps per rollout, against GLM-5.3's 35 minutes and 124 steps.

On reliability, Sol solved 61 tasks four for four, versus 48 for GLM-5.3, with reliability at 84.5% and 78.8% respectively. Sol broke existing tests in 20% of its failures, against 11% for GLM-5.3.

Task-type results split four apiece. Sol led data modeling and serialization (92%), build and ops tooling (73%), concurrency and durability (72%), and protocol conformance (59%). GLM-5.3 led query and config languages (88%), language and runtime internals (83%), and stateful reactivity (73%), and tied program analysis at 64. By language, GLM-5.3 led JavaScript 90 to 75 and Rust 70 to 60; Sol led Python 74 to 66, Go 79 to 76, and TypeScript 66 to 61.

Per-task correlation was 0.43. Each model solved 90 tasks; GLM-5.3 alone solved 9, Sol alone 7, and 7 defeated both. Their union covered 106 of 113 tasks (93.8%). A GLM-first cascade that escalates to Sol when tests fail solved 85.9% of tasks at $6.61 each; Sol alone solved 72.7% at $8.37. The report stated a perfect one-shot oracle router reached 83.8%.

Method notes: data came from a DeepSWE v1.1 export of published per-trial records. Costs are the published per-trial cost_usd. Per-turn trajectory JSONs for the GLM-5.3 batch were not on the public CDN at analysis time, so the analysis is index-level. Task types were classified by an LLM from benchmark prompts. The cascade assumes a verifier decides escalation.

Read at Together AI

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

We ran 904 DeepSWE rollouts on GLM-5.3 and GPT-5.6 Sol. Sol leads pass@1 by 3.7 points; GLM-5.3 wins pass@4 at half the cost, and a GLM-first cascade hits 85.9%.