Scaling mobile agents by scaling their environments
A mobile agent improves with the interaction experience it is trained on, and the open community has had little of it. Human demonstrations are fixed once recorded, emulator rollouts cover a few dozen utilities and open-source clients, and simulated apps have been few and shallow.
OpenMobile-2 puts environment scaling into practice. It rebuilds commercial apps as controllable simulated clients, collects training data on them, trains agents with that data, and releases all of it.
- Environment scaling with replicated applications. MobileGym++ rebuilds 35 commercial apps as simulated clients with persistent business state, clean reset and cross-app workflows, in a playground of more than 110 apps.
- Open training data and recipes at scale. OpenMobile-Data holds 11.6K demonstration trajectories and more than 2K executable RL tasks with verifiers, released with the SFT and RL recipes.
- A unified framework for GUI and app-native tool use. More than 300 app-native tools across 50+ apps write the same state as the GUI, and MobileGym++ Bench compares GUI-only and hybrid agents on the same 215 tasks.
- Open-data agents that transfer. OpenMobile-2 9B and 27B are competitive on AndroidWorld and MobileWorld and carry the gain to real devices on SPA-Bench.
The playground
Commercial apps rebuilt as simulated clients, with the emulator apps alongside.
The data
Demonstration trajectories and executable RL tasks, with the recipes that use them.
App-native tools
GUI actions and tool calls on one app state, compared on the same tasks.
Results
Open-data agents against commercial and open-weight models, on emulators, simulated apps and real devices.
A playground of commercial apps
MobileGym++ rebuilds commercial apps as simulated clients with persistent business state, clean reset and programmatic verification, next to a configured set of emulator apps on the same playground.
Apps by domain and runtime
Information moves between apps
How a commercial app is rebuilt
Against existing environments
The playground, running in this page
The simulator behind the playground needs no server: the whole phone runs in the browser, its state is one snapshot that resets in a click, and every app-native tool can be called from the panel beside it.
Episodes on the playground, step by step
MobileGym++ Bench tasks run by Gemini 3.1 Pro Preview in hybrid mode, replayed tap by tap with the reasoning behind each action and every app-native tool call.
Trajectories and tasks, collected on the playground
Both runtimes feed one pipeline, and two resources come out of it.
One trajectory, from app graph to released data
The same task, with or without tools
Apps expose selected functions as tools that write the same state as the GUI, so one task suite compares GUI-only and hybrid agents under one checker.
One benchmark task, two ways
The same instruction run twice by Gemini 3.1 Pro Preview, first with taps only and then with the apps' tools switched on.
What a tool looks like
MobileGym++ Bench
Long-horizon tasks across the commercial apps, half of them crossing app boundaries, each checked over the before-and-after diff of business state.
Two open-data agents, five benchmarks
GUI-only SFT, hybrid SFT and RL on the playground, from Qwen3.5-9B and Qwen3.6-27B, against commercial and open-weight agents on emulators, simulated apps and real devices.
Main results
What each training stage adds
What drives the gains
Coverage over repetition
At a fixed budget of emulator trajectories, spreading them over more apps raises success on unseen apps while the familiar apps hold.
Tools pay off after hybrid training
Tool access barely moves success before hybrid SFT. After it, the hybrid lane gains the most and finishes in fewer steps.
Environments, data, models and recipes
MobileGym++
The simulated clients, the app-native tools and MobileGym++ Bench.
OpenMobile-Data
The released trajectories and the verifier-backed RL tasks, on Hugging Face.
Models
Checkpoints of both agents after each training stage.
Code
Data collection, SFT and RL recipes, and the evaluation harness.
BibTeX
@article{openmobile2,
title={OpenMobile-2: Building Versatile Mobile Agents with Scalable Environments and App-Native Tools},
author={Chenyang Yan and Yingying Zhang and Kanzhi Cheng and Qiushi Sun and Zhengyuan Pan and Zheng Ma and Hang Yan and Yian Wang and Nuo Chen and Jialin Cao and Xingdong Gong and Wenpo Song and Han Chen and Zichen Ding and Fangzhi Xu and Shujian Huang and Xinyu Dai and Yichen Liu and Tiankuo Yao and Bo Wang and Ben Kao and Jianbing Zhang and Lewei Lu and Dahua Lin},
journal={},
year={}
}