A group of about thirty researchers published RecreationWorld on 18 September, a test environment for computer-use agents, along with RecreationBench, a 250-task benchmark that runs on it. The setup is deliberately awkward. An agent is shown a working reference application and told to build a faithful copy, with no instructions describing what the application does; it has to work that out by using it. The environments span Ubuntu, macOS, Windows, Android and the web, with one toolset for driving a graphical interface and another for writing code.
The headline is a pair of numbers that sit oddly together. GPT-6 Astra scored 58.1% overall, and passed every programmatic test on 2.8% of tasks. Partial credit accumulates steadily; complete, working reproductions almost never arrive.
The paper also reports that agents reproduce static interface elements more reliably than they implement interactions or computed outputs, and that the applications they produce come out smaller and more monolithic than the originals. Separately, the authors report that models trained on trajectories from the task improved on five unrelated benchmarks.
This is a preprint. The figures are the authors’ own and nobody outside the group has reproduced them.
