OpenMobile-2

Building Versatile Mobile Agents
with Scalable Environments and App-Native Tools

Mobile agents learn from interaction experience, and the trajectories behind the strongest agents are proprietary. Open efforts collect theirs on a few dozen emulator apps, mostly system utilities and open-source clients, because third-party apps rarely reset cleanly and the full Android stack is costly to replicate at training scale. OpenMobile-2 is an open stack for closing that gap, with the environments, data, recipes and trained agents all released.

  • MobileGym++: commercial apps rebuilt as high-fidelity simulated clients, with persistent business state, clean reset and cross-app workflows
  • OpenMobile-Data: demonstration trajectories for supervised fine-tuning and the largest open set of executable, verifier-backed RL tasks for mobile agents, released with the training recipes
  • App-native tools alongside the GUI: apps expose selected functions as tools that act on the same state, so an agent can interleave GUI actions and tool calls within one task, trained and evaluated on MobileGym++ Bench
  • OpenMobile-2 9B / 27B: open-data agents trained with GUI-only SFT, hybrid SFT and RL on the playground, competitive on established benchmarks and carrying the gain to real devices

Chenyang Yan, Yingying Zhang, Kanzhi Cheng, Qiushi Sun Zhengyuan Pan, Zheng Ma, Hang Yan, Yian Wang, Nuo Chen, Jialin Cao, Xingdong Gong, Wenpo Song, Han Chen, Zichen Ding, Fangzhi Xu, Shujian Huang, Xinyu Dai, Yichen Liu, Tiankuo Yao, Bo Wang, Ben Kao, Jianbing Zhang, Lewei Lu, Dahua Lin

Highlighted authors contributed equally.

Overview

Scaling mobile agents by scaling their environments

A mobile agent improves with the interaction experience it is trained on, and the open community has had little of it. Human demonstrations are fixed once recorded, emulator rollouts cover a few dozen utilities and open-source clients, and simulated apps have been few and shallow.

OpenMobile-2 puts environment scaling into practice. It rebuilds commercial apps as controllable simulated clients, collects training data on them, trains agents with that data, and releases all of it.

  1. Environment scaling with replicated applications. MobileGym++ rebuilds 35 commercial apps as simulated clients with persistent business state, clean reset and cross-app workflows, in a playground of more than 110 apps.
  2. Open training data and recipes at scale. OpenMobile-Data holds 11.6K demonstration trajectories and more than 2K executable RL tasks with verifiers, released with the SFT and RL recipes.
  3. A unified framework for GUI and app-native tool use. More than 300 app-native tools across 50+ apps write the same state as the GUI, and MobileGym++ Bench compares GUI-only and hybrid agents on the same 215 tasks.
  4. Open-data agents that transfer. OpenMobile-2 9B and 27B are competitive on AndroidWorld and MobileWorld and carry the gain to real devices on SPA-Bench.
Environment

A playground of commercial apps

MobileGym++ rebuilds commercial apps as simulated clients with persistent business state, clean reset and programmatic verification, next to a configured set of emulator apps on the same playground.

Apps by domain and runtime

Information moves between apps

How a commercial app is rebuilt

Against existing environments

Live demo

The playground, running in this page

The simulator behind the playground needs no server: the whole phone runs in the browser, its state is one snapshot that resets in a click, and every app-native tool can be called from the panel beside it.

Examples

Episodes on the playground, step by step

MobileGym++ Bench tasks run by Gemini 3.1 Pro Preview in hybrid mode, replayed tap by tap with the reasoning behind each action and every app-native tool call.

Open the trajectory viewer Every recorded episode, filtered by app, mode and tool use, plus samples from OpenMobile-Data.
Data

Trajectories and tasks, collected on the playground

Both runtimes feed one pipeline, and two resources come out of it.

One trajectory, from app graph to released data

App-native tools

The same task, with or without tools

Apps expose selected functions as tools that write the same state as the GUI, so one task suite compares GUI-only and hybrid agents under one checker.

One benchmark task, two ways

The same instruction run twice by Gemini 3.1 Pro Preview, first with taps only and then with the apps' tools switched on.

What a tool looks like

MobileGym++ Bench

Long-horizon tasks across the commercial apps, half of them crossing app boundaries, each checked over the before-and-after diff of business state.

Results

Two open-data agents, five benchmarks

GUI-only SFT, hybrid SFT and RL on the playground, from Qwen3.5-9B and Qwen3.6-27B, against commercial and open-weight agents on emulators, simulated apps and real devices.

Main results

What each training stage adds

Analysis

What drives the gains

Coverage over repetition

At a fixed budget of emulator trajectories, spreading them over more apps raises success on unseen apps while the familiar apps hold.

Tools pay off after hybrid training

Tool access barely moves success before hybrid SFT. After it, the hybrid lane gains the most and finishes in fewer steps.

Citation

BibTeX

openmobile2.bib
@article{openmobile2,
  title={OpenMobile-2: Building Versatile Mobile Agents with Scalable Environments and App-Native Tools},
  author={Chenyang Yan and Yingying Zhang and Kanzhi Cheng and Qiushi Sun and Zhengyuan Pan and Zheng Ma and Hang Yan and Yian Wang and Nuo Chen and Jialin Cao and Xingdong Gong and Wenpo Song and Han Chen and Zichen Ding and Fangzhi Xu and Shujian Huang and Xinyu Dai and Yichen Liu and Tiankuo Yao and Bo Wang and Ben Kao and Jianbing Zhang and Lewei Lu and Dahua Lin},
  journal={},
  year={}
}