ComputerWorld

How it works

Determinism and checkpoints#

Reproduction requires the same engine, registered module versions, world definition, seed and ordered action sequence. Replay is a semantic guarantee for supported operations, not a promise that arbitrary future engine versions will interpret old checkpoints identically.

Time is unsigned logical microseconds. Clock rejects backward movement and checked overflow. Scheduled work is ordered by due time, phase and insertion sequence. Neither executing a command nor observing a frame consults wall time.

Determinism implements sha256-named-splitmix64-v1: each named random stream is derived from the seed and stream name. Using an unrelated stream does not perturb another subsystem. ID counters are namespaced, deterministic and checked for overflow. Streams, IDs and queued work are serializable state. Rust extensions must use the supplied deterministic context instead of host clocks or RNGs.

Snapshots retain simulation and environment state; the facade owns the complete checkpoint boundary, including browser/application sessions. In-memory snapshots and forks share immutable backing with copy-on-write roots where implemented. Portable JSON export/import traverses state and is a different, more expensive operation. Code registrations remain external and must match the checkpoint's module identities. A checkpoint with unresolved live host effects is rejected on restore; synthetic queued work is serializable. Inspecting or evaluating should not consume random values or advance logical time.

Applications do work between actions only in steps: after the actions of every step(), each native application with background work (a video export) advances by one bounded unit (NativeApp::background), so how far it has got is a function of the steps taken and replays exactly. Media playback reads the world clock: a playing timeline's position is derived from the logical time playback started and the time now, never from the host (see video-editing.md).

Record trajectories for explanation and the input sequence for replay. Event records include sequence, tick, kind, optional machine/actor and structured data. Use state hashes to compare reconstructed suffixes. A hash mismatch is diagnostic failure, not something to hide by dropping disagreeing event fields.

The renderer uses fixed scene geometry, bundled font data and explicit raster inputs. Host font lookup and wall-time animations are not part of canonical output: CSS transitions and animations, timers and animation frames all resolve on the world clock. Web layout is deterministic too — cw-web lays out in app units (1/64 px integers, never floats), measures through the same cw_scene::metrics the native renderer draws from, and sorts anything that would otherwise depend on hash order — so a page produces the same fragment tree and the same pixels on every platform. What is not promised is that a given page lays out exactly as Chromium would; parity with Chromium is measured on a fixture corpus and reported, not guaranteed.

The corpus#

crates/computerworld/tests/determinism_corpus.rs is the regression evidence for all of the above. It replays five scenarios over two worlds, five seeds, four machines (including a Windows box, for a second shell dialect) and all seven built-in action families, 300 to 360 steps each, against goldens recorded from a known-good commit. It pins state_hash() after every step — so a divergence names the step that caused it, not just the run — plus a digest of every StepResult, the complete event log with its chained hash, sampled scene() digests at two viewports, the action journal's head hash, the exact byte length and SHA-256 of an exported snapshot, two forks driven apart, a replay of the recorded journal, a snapshot fixture exported at the baseline commit, and the negative half: every case validate_snapshot must still refuse, with the same error code and the same message. It refuses to be regenerated by accident.

Run it as an independent check when changing an engine hot path, rather than after: the fork and step optimizations since 0.2.0 left every hash byte-identical, and the corpus is what established that.

Rendered frames as labelled data#

Because a frame is an exact function of (engine, world, seed, action sequence, viewport), the text an agent types is also a label for the pixels that text produces. Replay a recorded episode, render at whatever viewport you want, and you have glyph-accurate ground truth for every string on screen — for free, at whatever volume you are willing to render, with no annotation pass and no labelling error.

The scene is the index. scene(width, height) gives each text node's string, bounds and transform before rasterization, so a (crop, string) pair falls out of walking the scene and cropping the frame from render(width, height) at the same viewport. Varying the theme, viewport and bundled typeface varies the rendering of identical text, which is the axis that matters for OCR generalization.

One consumer used this to train an OCR model on its own agent's typed text and measured held-out-font accuracy rising from 0.668 to 0.794.

This only holds within one engine version: pixel output is not stable across releases (the alpha shell work changed it), so keep the version with the data.

What that buys, beyond replay: frames that are their own labels, above; and bounded model checking, which needs a hashable state, a deterministic transition relation and cheap backtracking, and would not be possible without them. The cost of a fork and of a step is measured in performance.