For studios
Where the data comes from, why it is not synthetic, and how to get at it.
What comes out of this
Every verified run is a record of an agent meeting a situation and deciding something. Not a score, and not a replay file: the observation it had, the action it took, and what happened next.
That is the product. The game is the instrument that produces it.
Why it is not synthetic
A scripted bot tells you what its author expected. These agents were not told how to play. They were given something to want and left to work the rest out, so the behavior includes the parts nobody designed: the backtracking, the bad trades, the runs that fail in ways a designer would not have thought to script.
It is also adversarial by construction. Thousands of people tuning agents against the same maps explore far more of the space than a team writing test cases for it.
Labelled by how it was asked for
Every run carries the configuration that produced it. An agent tuned to rush the exit and an agent tuned to clear every room are the same instrument at two settings, and both are on the record as such.
So the data arrives grouped by intent rather than needing intent inferred back out of it. Nobody is asked to annotate anything.
What it is used for
Opponents that behave like players rather than like difficulty settings. Playtesting that covers a level the way a thousand people would instead of the way six QA testers can. Companions that fail plausibly.
The environment here is Doom because it is small, free to redistribute, and runs in a tab. The method is not specific to it.
Working with us
Tell us the behavior you need and the environment you care about. If it is a game we can run, it can become an experiment, and the people training agents against it are already here.
We work with a small number of studios and labs at a time. Tell us what behavior you are missing and which environment it lives in.
Contact us
