What you need to know
In a significant advancement for the Android development community, Google has unveiled Android Bench 2.0, an enhanced benchmark aimed at assessing the capabilities of large language models (LLMs) and AI agents in tackling intricate Android development tasks. This new iteration builds upon the original Android Bench introduced earlier this year, which focused on evaluating AI performance in real-world coding scenarios.
The upgraded benchmark is specifically designed to address long-horizon tasks (LHTs), which are considerably more complex and could require human engineers several days or even weeks to complete. Unlike its predecessor, which primarily measured smaller, incremental changes, Android Bench 2.0 raises the stakes by including challenging tasks such as:
- Upgrading dependencies
- Adding major new features
- Building Android apps from scratch
One of the notable changes in Android Bench 2.0 is its grading system. Moving away from a binary pass-or-fail approach, the new benchmark employs a “continuous scoring” method. This innovative scoring system offers a more nuanced understanding of a model’s performance, even if it does not fully complete a task.
As for the results, the leaderboard reveals that GPT-6 Astra currently leads with a commendable 28% pass rate, while Gemini 3.8 Flash trails behind with just 8%. Other models evaluated include Claude Fable 5.1, GPT-5.6 Sol, and Claude Opus 5, among others.
Google emphasizes that testing these models against the LHT dataset will provide deeper insights into their strengths and weaknesses, equipping developers with practical guidance on which models are best suited for various Android development challenges. The updated Android Bench 2.0 leaderboard is now accessible, and Google has plans to expand it further with additional models and results in the future.