How we found 24 Android vulnerabilities using our open source AI security agent

In a significant advancement for mobile application security, the GitHub Security Lab has unveiled its innovative Taskflow Agent, designed to assist security researchers in automating the identification of vulnerabilities within Android applications. This open-source tool leverages artificial intelligence to streamline the process of vulnerability detection, enabling researchers to package and share effective workflows and prompts.

By employing custom taskflow prompts, security experts can guide AI models through a structured approach, breaking down complex research into manageable steps. This methodology has proven effective, with over 20 vulnerabilities reported in various Android applications. The GitHub Security Lab has made these findings accessible through their advisories page, allowing developers to stay informed about potential security risks.

How to Run the Taskflows on Your Own Project

For those eager to implement these taskflows, the process is straightforward. However, it is important to note that a GitHub Copilot license is required, as the prompts utilize premium model requests. Running the taskflows may lead to numerous tool calls, which can consume a significant amount of tokens. Here’s how to get started:

  1. Navigate to the seclab-taskflows repository and initiate a codespace.
  2. Allow a few minutes for the codespace to initialize.
  3. In the terminal, execute ./scripts/audit/run_mobile.sh myorg/myrepo.

Upon completion, an SQLite viewer will display the results. Users can examine the “auditresults” table to identify vulnerabilities marked with a checkmark in the “hasvulnerability” column.

To enhance the taskflows’ effectiveness, specific adjustments were made to focus on the unique vulnerabilities associated with Android applications. For instance, the taskflow gathermobileentrypointinfo.yaml was introduced to categorize entry points in the code, distinguishing between mobile and non-mobile entry points. This refinement allows the AI to operate effectively across diverse application types, ensuring a comprehensive understanding of the attack surface.

Two Examples of Vulnerabilities Found by the Taskflows

The taskflows have successfully identified several critical vulnerabilities, two of which are highlighted below:

Tracking Users via OsmAnd

OsmAnd, a widely-used navigation app, was found to have a vulnerability that permits malicious applications to track user locations. The app exports an activity called MapActivity, which handles settings files and deeplinks. However, the way it processes intent extras allows any app to send arbitrary data to this exported activity. This oversight enables attackers to manipulate settings undetected, leading to potential privacy breaches.

Wikipedia Account Takeover via Deeplink

The Wikipedia Android app was also found to contain a logic flaw that could allow attackers to redirect users to malicious websites disguised as legitimate Wikipedia pages. By exploiting a deeplink mechanism, attackers can execute arbitrary JavaScript within the app’s WebView, creating a pathway for account takeovers. This vulnerability highlights the importance of rigorous security measures in mobile applications.

LLMs Are Good at Finding Vulnerabilities but Struggle at Estimating Severity

While large language models (LLMs) excel at identifying vulnerabilities, they often struggle to accurately assess their severity. Many findings require specific states that may be challenging to replicate in real-world scenarios, leading to potential false positives. Security researchers are encouraged to review AI-generated findings to ensure a comprehensive understanding of the actual impact of each vulnerability.

LLMs Have Great Knowledge of API Behavior

Despite some limitations, LLMs demonstrate a robust understanding of API behaviors across various programming languages. This capability allows them to generate effective proof-of-concept exploits with minimal modifications, showcasing their potential as valuable tools in the security research landscape.

Notes on the Results

As of the latest update, the GitHub Security Lab has identified 24 vulnerabilities in Android applications. These findings underscore the effectiveness of AI-driven security research in uncovering both common and critical vulnerabilities, reinforcing the need for ongoing vigilance in mobile application security.

AppWizard
How we found 24 Android vulnerabilities using our open source AI security agent