DeepSeek V4 Pro 0813 vs Claude Fable 5 on DeepSWE: Cost, Coding, and Routing

Together AI published results from 904 DeepSWE rollouts comparing DeepSeek V4 Pro 0813 (max) and Claude Fable 5 (max) across all 113 DeepSWE tasks, four trials each, using published per-trial records. DeepSWE is described as a benchmark testing software engineering ability across task types and programming languages.
On first-attempt accuracy, Claude Fable 5 leads at 69.7% pass@1 versus DeepSeek V4 Pro 0813's 62.8%, a 7-point gap. The lead narrows under retries: the two are level at pass@2 (78.5% versus 77.1%), and Pro leads at pass@4 (88.5% versus 84.1%).
Cost differs sharply. Together AI reports $0.24 per rollout for Pro against $21.63 for Fable, a 90x gap, or 260 solves per $100 for Pro versus 3 for Fable. Median rollout time is roughly even: 31 minutes for Fable versus 35 for Pro. Fable produced 115k output tokens over 79 steps; Pro used 146 steps.
Failure profiles also differ. Each model regressed the existing test suite in 11% of failures. Fable had the larger big-miss share (18% versus Pro's 10%); Pro's failures were near misses more often (66% versus 57%).
By domain, Fable won 6 of 8 task types, including data modeling and serialization at 88%, and language internals at 78%. Pro won stateful reactivity (66 versus 64) and concurrency and durability (58 versus 45). By language, Fable led Rust (85% versus 65%), Python, Go and JavaScript; Pro led TypeScript (61% versus 57%).
Per-task correlation was 0.39. Both solved 88 tasks; Pro alone solved 12, Fable alone solved 7, and 6 defeated both, covering 107 of 113 tasks. Together AI reports a Pro-first cascade that escalates to Fable on test-suite rejection solves 82.7% of tasks at $8.28 each; the reverse order costs $21.71 for the same accuracy.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
We ran 904 DeepSWE rollouts on DeepSeek V4 Pro 0813 and Claude Fable 5. Fable leads pass@1 at 90x the cost; Pro wins pass@4, and a Pro-first cascade hits 82.7%.