ComputerWorld: A reproducible world for computer-use agents

Train, evaluate, and build agents across simulated computers and a connected internet. Fork checkpoints, replay actions, and inspect the state behind the screen.

01 / For researchers

Training environments for computer-use agents.

Fork world state for parallel rollouts, compute rewards from explicit state checks, and extract labelled frames.

Branch from the same checkpoint.

Take a snapshot, fork independent worlds, and try different action sequences. Each branch starts from the same state. Repeating the same actions with the same engine, world, seed, and viewport reproduces the run exactly.

Fork latency
35.0 µs
Retained per fork
20.9 KB

Measured natively on company-2026 after 1,000 actor steps. Retained memory is before a branch writes; subsequent writes have additional costs. Benchmark and methodology ↗

One checkpoint / three trajectories
The repository’s two open pull requests, the state every branch starts fromCheckpoint2 open pull requests
    1. The pull request “Sort refs before iterating in BFS”, still open, with a request for changes outstanding1Open #15
    2. The pull request’s two changed files, twelve additions and six deletions, still open2Files changed

    No policy broken

    1. The same pull request, reached from the same checkpoint in a world of its own1Open #15
    2. The same pull request closed by the agent, its badge reading Closed and the repository down to one open pull request2Close pull request

    P12 broken

    1. The repository’s stargazers, eight of them now that the agent has starred it, the button reading Starred1Star

    P15 broken

Actual simulator frames, one per action. Three worlds forked from the same checkpoint and driven apart: one reads the pull request and changes nothing; one closes a pull request somebody else opened, which is the counterexample the checker returned for P12; one stars the repository, which is its counterexample for P15. The two breaking paths are replayed from the report, so the picture cannot drift from the run.
Rendered frame / structured scene
A rendered frame from inside the world: a macOS desktop with the menu bar and dock, and Safari maximized on a GitHub pull request — toolbar and address bar, repository navigation, the pull request title and review conversation, and a sidebar of reviewers, assignees and labels. The same frame with the scene drawn over it: a hairline box around each of 96 text runs and a thick box around each of 67 enabled controls, sixteen of them labelled with the role and string the scene reports.
This frame contains 96 text runs and 67 enabled controls. Toggle the overlay to see their bounds.

State checks and frame labels.

Use explicit predicates over world state to compute per-step rewards. A check gives the same verdict on replay; what it measures depends on the specification you write.

scene() exposes text, widgets, and their bounds before rendering. Use those labels for perception training or enumerate the controls available to an agent.

Engine-labelled frames improved held-out-font OCR accuracy from 66.8% to 79.4%. OCR case study ↗

02 / For evaluation teams

Reproducible evaluations and bounded policy checks.

Replay failures exactly. Explore a defined action space and return the sequence that violates a policy.

Inspect the actions behind a violation.

Choose a start state, action set, and search depth. The checker explores the configured search space and tests policies against state. A counterexample records the actions needed to reproduce the violation.

In the example, two actions violate P2: open the pull request, then merge it while a request for changes still stands.

A result that holds is bounded by the chosen model, start state, action domains, and depth. It is not a guarantee about every possible action or a production application.

Recorded check + counterexample replay
The recorded search finds and replays the two-action violation in the simulated pull-request workflow.

Reference-world results

15 policies checked to depth 5.

Read the evaluation ↗

6held within the search

9reachable violations

77,748states checked

Simulated pull-request review and merge flow · engine 0.2.0 · seed 7 · 21 September 2026. All nine counterexamples reproduced from a fresh world.

View all policies and counterexamples
Checked against the company-2026 world. Depth 5 exhausted — every action sequence of five steps or fewer from the start state. 77,748 states checked in 2 h 59 m by scripts/checks/cw-check.mjs on engine 0.2.0, 21 September 2026. All nine counterexamples were replayed from a fresh world and reproduced.
PolicyStatementVerdict
The run's own output is data/policies.json; this table is that file, rendered.

03 / For developers

Programmable computers, applications, and services.

Define a world and control it through Python, JavaScript, or Rust. The same Rust engine runs natively and as WebAssembly.

Configure the world your agent uses.

World definitions describe machines, users, files, and network services. Connect desktops and phones to shared mail, chat, documents, and Git services, then give each agent a handle with explicit action and observation permissions.

Extend the simulation with custom applications and services. Supported operations and fidelity limits are documented per subsystem.

What you hold, and what it holds

You write

Your codePython · JavaScript · Rust
actions ↓↑ scenes · frames · state
ComputerWorldone world · snapshot · fork · replay

It holds

  • macOS
  • Windows 11
  • Ubuntu
  • iPhone
  • Android

on one network — mail, chat, docs, git, the web

You never leave your own process. One handle takes actions and gives back a scene, a frame and a state hash; everything past it is the world’s to run.

Start with an environment.

These examples use a world definition. Follow the first-episode guide for setup and a complete example. To work on the engine itself, see the contributor guide.

pip install computerworld
from computerworld import World

world = World(definition, seed=42)
env = world.environment({"actor": "alice", "machines": ["alice-mac"],
                         "actions": ["pointer.v1", "keyboard.v1", "application.v1"],
                         "observations": ["semantic.v1"]})

env.step([{"family": "application.v1", "op": "launch",
           "machine": "alice-mac", "payload": {"kind": "code"}}])

scene = env.scene(1440, 900)     # roles, names, geometry — no pixels needed
frame = env.render(1440, 900)    # or exact RGBA, if your model wants pixels