Training OCR with simulator-labelled frames

Held-out-font accuracy increased from 66.8% to 79.4% in one reported experiment.

Rendered frame / ground-truth bounds
Engine-rendered GitHub page with text and control bounds overlaid
An example of the scene labels available from ComputerWorld; this is not a sample from the reported training dataset.

How the labels are produced

An agent types text in the simulated world. The engine renders that text and exposes its string, bounds, and transform through scene(width, height). Pair each text region with the corresponding crop from render(width, height) to produce labelled OCR examples.

What changed in the experiment

One consumer trained an OCR model using its own agent’s typed text. Reported accuracy on held-out fonts rose by 12.6 percentage points, from 66.8% to 79.4%.

The project’s published account does not specify the model architecture, training-set size, split, or uncertainty estimates. This is a reported result from one experiment, not a general performance guarantee or a complete reproducibility package.

Using the same method

  1. Record an episode containing the text you want to render.
  2. Replay it with the same engine version, world, seed, and actions.
  3. Extract text nodes and render frames at the same viewport.
  4. Vary supported themes, viewports, and bundled typefaces, keeping evaluation fonts separate from training fonts.

Keep the engine version with the dataset: pixel output can change between releases.

Technical documentation and source account ↗