Akka Tests Spec-Driven AI Delivery across 65 Open Source Projects

Akka ran a spec-driven workflow for AI-assisted software porting across 65 open source projects, examining how specification structure, context, model selection, automated validation, and delivery guardrails affect results, according to an account by Leela Kumili. The experiment measured time, token use, code size, test parity, and performance.
The first tranche took 99.3 hours and consumed 9.41 billion tokens, with Akka reporting a lines-of-code or performance improvement in 57 of the 65 ports. The work ran in two tranches: Akka generated specifications and implemented up to 10 percent of each of the 65 projects' surface area, then selected 10 for complete implementation based on system characteristics and measurable results.
The delivery harness cycled through discovery, specification, porting, benchmarking, and improvement. Discovery analyzed code, models, schemas, and runtime behavior, while Claude with Akka Specify handled implementation, testing, and review. A common benchmark runner compared tests, code size, and latency.
Akka found structured specifications with claims, evidence, and typed behavior improved first-pass implementations, while gaps in context files remained around cross-component decisions. Follow-up areas include interface enumeration, test ingestion, provenance tracking, differential testing, and adversarial testing.
Sonnet averaged 61 minutes per port versus 120 minutes for Opus, while Opus used about 40 percent fewer tokens; higher effort settings increased consumption without consistently improving efficiency. Validation used original unit and integration tests plus auditors checking serialization, security, error handling, PII, idempotency, and architectural boundaries. Akka added guardrails as failures exposed new issues and reported that additional exit conditions increased porting costs.
Performance varied by project: applications, frameworks, and libraries generally improved, while infrastructure and tooling showed median degradation. Akka reported a 143,333 times improvement for Dify but noted the compared workloads differed, while Netflix Metaflow was approximately 100 times slower.
Commenting on Akka CEO Tyler Jewell's LinkedIn post, Aaditya wrote that the smaller model's behavior matched modernization work he had observed, where the small model follows the spec while a larger model may improvise, and questioned whether reductions in lines of code resulted primarily from dead code removal or from differences in the target language. Rick Bryce, Head of Marketing at Avahi, suggested constraints could influence the result.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
Akka used 65 open-source projects to examine how specification structure, context, model selection, automated validation, and delivery guardrails affect AI assisted software porting. The experiment measured time, token use, code size, test parity, and performance, finding substantial variation across models, effort levels, and project types. By Leela Kumili