Conceptual

Dynamic In-the-Wild Evaluation Platform for Mobile GUI Agents

An interactive benchmarking environment that evaluates mobile GUI agents by running them live inside real third-party Android apps against multi-step, goal-oriented tasks, rather than scoring single frozen screenshots. It provides a dataset-agnostic action space so agents trained on any data can be tested, and uses an automated large-language-model judge to grade task completion with minimal human labelling.