scoring

AppWizard
September 19, 2026
Google has launched Android Bench 2.0, an upgraded benchmark for evaluating AI coding agents, which now includes 30 long-horizon tasks that can take human engineers days or weeks to complete. The best pass rate for these tasks is around 28%, down from approximately 91% in the previous version. The benchmark measures both pass rates and completion rates, acknowledging partial progress rather than just failures. It assesses various challenges, including app creation and feature integration, using a comprehensive scoring methodology that evaluates functionality, regression checks, and visual fidelity. AI models perform better with new code than with modifications to existing systems, facing challenges in runtime validation and cross-platform conversions. The current leaderboard shows OpenAI’s GPT-6 Astra leading with a 28% pass rate, while Gemini 3.8 Flash has an 8% pass rate. Developers using AI agents should be aware that while AI can generate significant portions of applications, further refinements will require human input. The benchmark is available on Google’s Android Bench leaderboard.
AppWizard
September 18, 2026
Google has released version 2.0 of Android Bench, which focuses on managing complex development tasks rather than minor adjustments. The new version evaluates tasks that may take engineers days or weeks to complete, such as adding features and building applications. A continuous scoring method has replaced the previous pass or fail system, assessing completion rates based on functionality, visual fidelity, and adherence to instructions. Various AI models have been tested, with GPT-6 Astra achieving a 28% pass rate, significantly lower than the previous scores around 90%. No model has achieved a 100% pass rate in porting cross-platform applications, with the best reaching 80%. AI performs better in writing new code than in refactoring existing code, facing challenges with architectural complexity and runtime validation. Android Bench utilized agents from model providers for evaluations and plans to incorporate various model combinations in future updates.
AppWizard
September 18, 2026
Google has launched Android Bench 2.0, a benchmark for evaluating large language models (LLMs) and AI agents on complex Android development tasks. This version focuses on long-horizon tasks (LHTs) that are more intricate than those assessed by the original benchmark. Key tasks include upgrading dependencies, adding major new features, and building Android apps from scratch. The new grading system uses continuous scoring instead of a binary pass-or-fail method, providing a more detailed performance evaluation. Currently, GPT-6 Astra leads the leaderboard with a 28% pass rate, followed by Gemini 3.8 Flash at 8%. Other evaluated models include Claude Fable 5.1, GPT-5.6 Sol, and Claude Opus 5. Google plans to expand the leaderboard with more models and results.
AppWizard
September 12, 2026
Insomniac's Wolverine has received mixed reviews, with GamesRadar rating it 3 out of 5, IGN giving it a score of 6 out of 10, Shacknews awarding it 8 out of 10, DualShockers rating it 8.5 out of 10, GameInformer also giving it 8.5 out of 10, TechRadar scoring it 4 out of 5, GameSpot rating it 6 out of 10, VGC giving it 3 out of 5, Eurogamer scoring it 3 out of 5, The Guardian rating it 3 out of 5, and Giant Bomb giving it a lower score of 2.5 out of 5. Critics praised the character of Wolverine and the voice acting but criticized the gameplay for being repetitive and lacking depth. The game is seen as falling short of expectations for a PlayStation 5 exclusive.
AppWizard
September 11, 2026
Devolver Digital, in collaboration with Dark Dark Goose, has announced HYPERDROP, a pachinko-inspired roguelite for PC. Players drop a ball onto a pegboard to create combinations for points, transforming the board by adding and upgrading pegs. The game features various upgrades that allow balls to split, return to the top, create gravity wells, or trigger chain reactions. Players can modify ball characteristics, introducing different physics behaviors. The game includes challenging stages, modifiers, and boss battles, with new challenges and board layouts unlocked as players progress. A free demo is available on Steam, but there is no confirmed release date.
Winsage
September 9, 2026
Microsoft's September 2026 security update revealed 973 vulnerabilities, with 113 classified as critical. Two actively exploited vulnerabilities are CVE-2026-81963 (Windows Update Stack, elevation of privilege, CVSS 7.8) and CVE-2026-85880 (Windows ALPC, elevation of privilege, CVSS 7.8). Among the 113 critical vulnerabilities, 82 are remote code execution (RCE) vulnerabilities. Notable vulnerabilities include: - CVE-2026-69676: RCE in Windows Kerberos, CVSS 8.8, authentication bypass. - CVE-2026-69852: RCE in Windows RRAS, CVSS 7.5, heap-based buffer overflow. - CVE-2026-72957: RCE in Windows Deployment Services, CVSS 7.8. - CVE-2026-69854: Elevation of privilege in Spring Cloud Azure, CVSS 9.0, improper authentication. - CVE-2026-83501: Information disclosure in Windows VBS, CVSS 5.5. - CVE-2026-69730: RCE in Windows DNS Server, CVSS 9.8. Less likely to be exploited vulnerabilities include: - CVE-2026-69845: RCE in Windows DHCP Server, CVSS 9.8, heap-based buffer overflow. - CVE-2026-65772: Vulnerability in Microsoft Dynamics 365 On-Premises, CVSS 8.8, deserialization of untrusted data. - CVE-2026-66302: RCE in Skype for Business, CVSS 9.8. Additional critical vulnerabilities include: - CVE-2026-62916: Elevation of privilege in Microsoft Entra ID, CVSS 9.1. - CVE-2026-83941: Elevation of privilege in Entra ID, CVSS 9.9. - CVE-2026-80098: Vulnerability in Copilot Studio, CVSS 9.3, improper verification of cryptographic signatures. Talos is releasing a new Snort ruleset to detect attempts to exploit these vulnerabilities, with specific SIDs for Snort 2 and Snort 3 rule coverage.
Search