Project HydraFusion: Frontier quality via multi-model orchestration
GitHub announced Project HydraFusion, a research preview that delivers frontier intelligence through runtime orchestration, available to users on all GitHub Copilot plans through /experimental in GitHub Copilot CLI.
HydraFusion creates a full execution plan, choosing from models across multiple providers to draft, critique and revise, or cascade to more powerful models. GitHub said it fills a key role in its strategy to deliver automated semantic routing between local, cloud, and compound models. Developers select HydraFusion like any other model, and it chooses a workflow that balances performance, cost, and latency for each task. Usage is based on tokens consumed by the models HydraFusion uses, priced at each model's standard rate.
For each request, HydraFusion currently chooses one of three execution patterns. Single has one selected model solve the task directly. Cascade has an efficient model draft a solution, with a quality gate deciding whether to accept it or escalate to a stronger model. Critique has one model draft a result, an independent read-only critic from a different model family review it following the same review pattern as Rubber Duck, and the drafting model revise once. The company said HydraFusion uses capability signals for reasoning, code generation, debugging, and tool use to select the most efficient execution pattern to meet the quality bar.
GitHub reported offline evaluations across three agentic coding benchmarks using Claude Opus 5 and GPT-5.6 Sol as comparison baselines, evaluated at the same medium reasoning level. On TerminalBench 2.1, HydraFusion improved verified task quality by 4.9 percentage points at 67% lower estimated cost versus Opus 5. On DeepSWE, cost was 36% lower with quality down 1.5 points; on CheckpointBench, its internal benchmark based on real GitHub Copilot sessions, cost was 65% lower with quality down 0.1 points. GitHub said these controlled offline results are specific to the evaluated benchmark revisions, workflow configurations, model pool, and pricing assumptions, and that the preview will validate how they translate to real developer workloads.
HydraFusion currently shows workflow stages but holds intermediate drafts until returning one coherent result, and GitHub said it is exploring better progress updates. Feedback is requested in the GitHub Community.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
In controlled offline evaluations, HydraFusion’s selective coding workflows matched or exceeded the evaluated Opus 5 baseline while reducing estimated workflow cost. Now available as a research preview in GitHub Copilot. The post Project HydraFusion: Frontier quality via multi-model orchestration appeared first on The GitHub Blog .