What it costs to fork a world

35.0 µs median fork latency and 20.9 KB retained per fork after 1,000 actor steps.

Measured on a working world

The benchmark starts the company-2026 reference world, drives an actor through real steps, then measures snapshots and forks in a native release build. The current measurements use three runs of 300 timed forks, with 50 warm-up forks and 200 live forks for the retained-memory measurement.

Accumulated stepsMedian fork latencyForks / second
015.9 µs62,753
1,00035.0 µs28,508
4,000116.4 µs8,511

Forking and writing have different costs

A live fork retains 20,863 bytes: 5,975 bytes of heap plus the 14,888-byte World value. This is the incremental cost of retaining an unmodified branch, not its eventual memory footprint.

At 1,000 steps, the measured first write in a branch allocated 30.5 MB and took 8.11 ms. At 4,000 steps it allocated 88.1 MB. Copy-on-write postpones copying shared state until mutation; the first write is a separate cost.

Measurement conditions

Measurements ran on an ARM Cortex-X925 / Cortex-A725 host with 121 GiB of RAM and Ubuntu 24.04, pinned to one performance core. Other workloads were active. At 1,000 steps, median fork latency across the three runs ranged from 34.8 to 35.0 µs; throughput ranged from 28,412 to 28,659 forks per second.

Reproduce the benchmark

BENCH_RUNS=3 BENCH_SAMPLES=300 BENCH_WARMUP=50 \
BENCH_LIVE_FORKS=200 BENCH_FORK_STEPS=0,1000,4000 \
cargo run --release -p cw-benchmarks --bin fork_throughput

The earlier five-run dataset in the repository predates the optimizations and does not contain these current figures. These results come from the subsequent measurements in the benchmark report.

Performance guide · Full engineering report and history ↗