regact is a research framework for agents that reason about an unknown
game and act in it. It drives a code-writing agent (Claude Code, codex,
or Alan) that plays an environment (ARC-AGI-3, MiniGrid), writes an
act(obs) -> action controller, and submits it for evaluation. The controller is
always-on; optional features such as Code World Model add capabilities. Controllers
can retain state between actions; evaluation creates a new controller for each episode.
The agent, the environment (a "problem"), and the optional features have separate extension interfaces within this controller-evaluation workflow. The agent reaches the environment through a localhost HTTP boundary. With sandboxing enabled, the game source is hidden from the agent, preventing it from bypassing exploration by reading the implementation.
A code-writing agent probes an unknown game, writes an act(self, obs) controller, and gets scored — browsed in the visualizer (sped up 2x). Full-quality clip.
Python 3.11 or 3.12 (not 3.13). Create a venv, install the core, then add only the extras you need.
python -m venv .venv && . .venv/bin/activate
make install # core framework + dev/lint/test tooling (pinned)Add a game engine and/or an agent backend:
make install-arc # the ARC-AGI-3 engine (problem=arc_agi)
make install-minigrid # the MiniGrid envs (problem=minigrid)
make install-agents # the Alan code agent (agent=alan)The two cloud CLI agents are external programs you install and authenticate once —
see docs/agents.md for claude and codex setup (one command
each). scripted (the deterministic test backend) needs nothing.
Three diagnostics, each reports only on what you installed:
make doctor # is the machine ready? (python, agent CLIs, sandbox, game extras)
make probe # does the OS sandbox actually confine here? (the R1-R6 contract)
make agentcheck # do the installed agent backends launch — bare and sandboxed?A run is composed from config groups you pick by name, plus fields you override
on the CLI. The defaults live in src/regact/conf/config.yaml:
agent: scripted # who writes the code - scripted | claude | codex | alan
problem: arc_agi # the environment - arc_agi | minigrid
controller: default # always-on: the agent writes + submits a policy (knobs: controller.*)
features: none # OPTIONAL extra capabilities - none | cwm
sandbox: true # confine the agent + block egress (false = off)
limits:
max_turns: 350 # agent turns per task
max_seconds_per_task: null # wall-clock per task
max_actions_per_env: null # env.step cap per env instanceThe always-on controller's eval knobs live under controller.* (e.g. controller.n_episodes);
each optional feature owns its knobs under features.<name>.*. A few examples:
# smoke test: scripted agent, no LLM; runs ARC ls20 (requires make install-arc and game data):
make run ARGS="experiment=dev"
# MiniGrid with Claude:
make run ARGS="agent=claude problem=minigrid"
# ARC-AGI-3 with Alan, add the Code World Model feature, 3 eval episodes:
make run ARGS="agent=alan problem=arc_agi features=cwm controller.n_episodes=3"See a config composed without running it: make run ARGS="... --cfg job".
Explore saved experiments in the local browser viewer:
make viz EXP=experiments # browse all experiments and benchmarks
# Or open one experiment's latest run:
make viz EXP=experiments/<experiment_name>/latestOpen localhost:8030. Set PORT=8031 to use another port. You can then navigate to an experiment and browse its tasks. In each of them you have access to panels Overview, Conversation, Artifacts (files and videos) and Graphs (metrics).
| Conversation | Overview |
|---|---|
![]() |
![]() |
| Follow the agent's messages and tool calls, and jump between submissions. | Inspect the game preview, scores, resource usage, and run configuration. |
If you performed your experiments in experiments/<benchmark_name>/, you can also compare multiple runs of the same benchmark and see their metrics side by side with the Graphs panel of a benchmark interface.
| Benchmark experiments | Benchmark graphs |
|---|---|
![]() |
![]() |
| Browse experiments and compare their task results and resource usage. | Compare metrics across experiments, choose an aggregation, and filter crashed runs. |
| Guide | What it covers |
|---|---|
| Overview | The three seams and how a run flows through them |
| Agents | Use an agent backend · add a new one |
| Environments | Use a problem · add a new one |
| Features | Use a feature · add a new one |
| Experiments | Launching runs, outputs, and the visualizer |
| Sandboxing | How isolation works and how it is verified |
make check # the CI gate: ruff + mypy + unit tests
make test-all # every test, including the live ones (needs alancode / a game)



