AivexaNewsSearch
AI news for builders and product teamsChecked every hour

Android Bench 2 Adds Support for Long-Horizon Tasks, Agentic Evaluation, and Continuous Scoring

Collected Oct 9, 2026

Google has released Android Bench 2.0, a major update to its benchmark framework for evaluating AI models and agents on Android development tasks. The framework first launched a few months ago and tests AI models against a set of common development tasks that incorporate Android best practices around permissions, navigation, and connectivity. Version 2.0 widens that scope considerably. The news was reported by Sergio De Simone for InfoQ.

The headline addition is a first set of long-horizon tasks, or LHTs. Google characterizes these as tasks of great complexity that would take an engineer multiple days or even a week to complete. Examples include upgrading dependencies, adding new features, building apps from scratch, and converting a cross-platform app to Android. That is a meaningful shift from the original version, which focused on incremental changes to existing repositories. In practice, an incremental change has a small blast radius: a model can often satisfy it by editing one or two files. A multi-day task instead requires the agent to hold a plan across many steps, keep intermediate state coherent, and recover from partial failures. This suggests LHTs are closer to how developers actually delegate work to coding agents, rather than how benchmarks traditionally score them.

The second change is agentic evaluation, which begins with agents from corresponding model providers. Details on the harness itself are thin, but the framing points to scoring a model as it operates through an agent loop rather than grading a single completion. A likely trade-off is that agentic results are harder to compare across vendors, because the surrounding scaffolding, tool access, and retry behavior can influence outcomes as much as the underlying model does.

The third change replaces binary pass/fail evaluation with continuous scoring. Under the old scheme, a complex task could be marked as failed because of a single failing edge-case assertion, even when the agent had met dozens of other requirements. Android Bench 2.0 instead computes a completion rate from a combination of factors: functionality, visual fidelity, and avoiding regressions. Google also applies objective scoring penalties when a solution deviates from evaluation instructions or structural constraints. For teams watching leaderboards, this matters because partial credit changes the ranking dynamics; a model that reliably gets most of a task right will now separate from one that either fully passes or fails outright.

The results are also meant to reveal where AI assistance is more likely to succeed. Google reports that AI does a better job at writing new code than at refactoring existing code, which the company links to the fact that refactoring and migrations require understanding the architectural complexity of a codebase. Conversely, AI performs well on what are described as well-established, deterministic transformations, even in larger codebases. Named examples include converting Java to Kotlin, swapping Retrofit for Ktor, and introducing a ViewModel layer. These are mechanical, well-documented changes with clear correct answers, which fits the pattern.

Failures cluster in a different set of conditions. Models still struggle with tasks that need runtime validation, such as missing dependency injection graphs, with breaking framework changes, and with knowledge gaps around unreleased libraries. Those three categories share a common property: the model cannot settle the question from static source alone, either because it must observe behavior at runtime, because the API contract shifted, or because the relevant knowledge is too new to be represented reliably. One concrete number stands out: the best-in-class model reaches only an 80% completion rate when porting a cross-platform app to Android, which Google treats as an overall open challenge.

The Android Bench 2.0 dashboard includes recent models such as Gemini 3.8 Flash, Gemini 3.7 Flash, OpenAI GPT-6, Anthropic Fable 5.1, Kimi K3, and Qwen 3.8 Max. At the time of the article, Claude Opus 5.5 sat at the top of the leaderboard with a 32% long-horizon task pass rate, followed by GPT 6 Astra at 28%. Those figures are worth reading carefully. A 32% pass rate on multi-day tasks is low in absolute terms, and my reading is that it reflects deliberately hard task selection rather than a straightforward verdict on these models' usefulness, since the earlier incremental task set is not what the LHT number measures.

Why it matters: teams choosing a coding model for Android work now have a benchmark that scores multi-step, multi-day tasks with partial credit instead of pass/fail, which should make comparisons more informative and less brittle. The reported gaps, weaker refactoring than greenfield writing, difficulty with runtime validation and breaking framework changes, and the 80% ceiling on cross-platform ports, point to where human review is still needed. My inference is that the 32% and 28% LHT figures will be the numbers vendors and engineering leads quote most, so it is worth understanding what a completion-rate score does and does not capture before drawing conclusions.

Read at InfoQ · AI, ML & Data Engineering

Based on reporting from the original publisher. Visit the source for full context and later updates.

Publisher excerpt

Google has released Android Bench 2.0, a major update to its benchmark framework for evaluating AI models and agents on Android development tasks. The update introduces long-horizon tasks (LHTs), agent-based evaluation, and continuous scoring to better assess performance on complex, multi-step development tasks. By Sergio De Simone