We reverse-engineer the world, then build the ground where AI gets tested.
One model dreams up what comes next. Another rebuilds it for real: deterministic physics, verified rules, running at 60 fps. What's left is a real, working world to test what an AI actually understands, not a screenshot of one.
How it works
Three roles, run on repeat, until the build holds up.
A diffusion model designs what comes next: one new mechanic, one hardware era, a bolder look. A coding model builds it for real, with deterministic physics and telemetry-verified mechanics, at 60 fps or it doesn't ship. Judges and discriminators score the result and catch what's faking it: a similarity score that looks high but doesn't hold up as real behavior. Then the loop runs again.
Custom evals, not synthetic benchmarks
We don't generate our evals from a prompt. We reverse-engineer the world itself, then test against what we rebuild.
Published work
From Diffusion Dreams to Playable Games is the formal writeup: target-faithful world implementation as constrained program search, run through four generative roles and measured across three studies on one deterministic engine. Submitted to IAAI‑27, Deployed Applications track. Currently under review.
Request the preprint →