AivexaNewsSearch
AI news for builders and product teamsChecked every hour

GLM-5.3 vs. Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

Collected Oct 1, 2026

Together AI published a comparison of GLM-5.3 (max) and Claude Fable 5 (max) on DeepSWE, a software engineering benchmark covering many task types and programming languages. The run used all 113 DeepSWE tasks with four trials per configuration, totaling 904 rollouts, 452 per model. Figures come from that run and may differ from other public scorecards.

On first-attempt accuracy the two are within noise. Fable 5 posted 69.7% pass@1 and GLM-5.3 posted 69.0%, a 0.7-point gap described as inside the error bars on both sides. GLM-5.3 led at higher attempt counts: pass@2 at 81.1% versus 77.1%, and pass@4 at 87.6% versus 84.1%.

Cost differed by 5.4x, at $3.99 per rollout for GLM-5.3 against $21.63 for Fable 5. Per $100 spent, GLM-5.3 solved 17 tasks and Fable solved 3. GLM-5.3 averaged 35 minutes per rollout against Fable's 34. Fable used 114k output tokens over 85 steps; GLM-5.3 used 80k tokens over 124 steps.

Reported per-task correlation was 0.65, the highest agreement in the set. Both models solved 88 tasks; GLM alone solved 11, Fable alone 7, and 7 defeated both. Their union covered 106 of 113 tasks (93.8%). The report recommends GLM-5.3 as default and escalating to Fable for Rust and serialization work. It states GLM-5.3 is open weight and Fable 5 is closed.

Method notes: scoring used DeepSWE's official included_in_score with infra errors excluded; Fable 5 had 16 infra errors from model routing (strict errors-as-failures score 67.3) and GLM-5.3 had 1. Costs are published per-trial cost_usd. A GLM-first cascade was reported at 81.1% solved at $10.74 per task.

Read at Together AI

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

We ran 904 DeepSWE rollouts on GLM-5.3 and Claude Fable 5. A tie on pass@1, but GLM-5.3 wins pass@4 and costs 5.4x less: \$3.99 per rollout vs. \$21.63.