Survive Flappy Bird with no reinforcement learning, no reward channel, no policy training, and no repeated episodes to learn from. The agent embodies exactly three concepts:
| concept | computational primitive |
|---|---|
| SELF | an indexical (which signals are my body) + an efference copy (what I just did) → gives (state, action) → state' pairs, so its own dynamics are identifiable by supervised prediction |
| PAIN / DEATH | contact ends the future (absorbing), and nociception intensity is to be minimised. This is the entire objective. It is innate and not learnable — only its scale is |
| OTHER | not-self occupied space, which persists, is spatially extended, drifts at a learnable speed, and hurts on contact |
Everything else — gravity, its own flap strength, its own terminal velocity, its own body radius, its own pain scale, the world's speed, the world's shape — is identified online from prediction error. The reason no RL is needed is that the three concepts supply the objective. What RL spends millions of samples discovering (what is good) is given; what remains (what my actions do) is dense, self-supervised, and available every single frame.
Receives, per frame: 16 ray distances over the full circle from its body centre to the
nearest not-self surface; a scalar nociception intensity in [0,1] whose mapping from
distance is not given; and a dead flag. Nothing else.
Never receives: its own position or velocity (it infers them from rays), the score, any reward, any physics constant, the pipe geometry, or the fact that "pipe", "ground" and "ceiling" are different kinds of thing.
Exact backward dynamic programming on a (y, vy) × time lattice:
Z[H][s] = alive(s, H)
Z[t][s] = alive(s,t) · exp(−β · pain(s,t)) · ( Z[t+1][flap(s)] + Z[t+1][fall(s)] )
a* = argmax_a min over plausible worlds log Z[1][ successor(s₀, a) ]
log Z is a free energy: it counts how many futures from here stay alive, discounted by
predicted discomfort. Death sets it to zero, so death is absorbing and dominant. Maximising it
is staying deep inside the viability kernel — "keep the most ways to keep living". No
preference for the middle of a gap is written anywhere.
Identification, from ~60 frames of ordinary flight and zero deaths:
| parameter | truth | identified |
|---|---|---|
gravity g |
1.0 | 1.000 ± 0.05 |
flap impulse J |
9.0 | 9.000 ± 0.10 |
terminal velocity vmax |
12.0 | 12.00 ± 0.20 |
world speed s |
3.0 | 3.000 ± 0.03 |
body radius r |
12.0 | 11.7 – 12.4 ± 0.7 |
nociceptive range λ |
25.0 | 23 – 24 |
Survival, 5 seeds × 20 000-frame cap (one continuous life each, learning carried across deaths):
seed 0 died at 6 767 111 pipes
seed 1 survived 20 000 332 pipes
seed 2 died at 7 966 131 pipes
seed 3 died at 11 202 185 pipes
seed 4 survived 20 000 332 pipes
66 135 frames, 3 deaths → one death per ~22 000 frames (~364 pipes). A reactive contact-avoider that has all the same senses but no world model dies after ~100 frames. Cost: ~15 ms/frame, single-threaded numpy.
| ablation | deaths | median frames | max pipes |
|---|---|---|---|
| full model | 1/3 | 3000 (cap) | 48 |
| − self-model identification | 3/3 | 136 | 5 |
| − option value (myopic, H=3) | 3/3 | 163 | 1 |
| − pessimism under uncertainty | 1/3 | 3000 | 48 (but min 81 — one life died during bootstrap) |
| − object permanence (map memory) | 3/3 | 223 | 10 |
| − pain gradient in plan (β=0) | 2/3 | 1540 | 48 |
Pessimism does not change the median at all — it changes the worst case, and only during the first ~100 frames while the self-model is still unidentified. That is exactly the role it was designed for, and it is invisible in any average.
With the nociception term switched off entirely (β=0), so that value is nothing but the count of surviving futures, at the moment a pipe arrives:
y= 25 logZ = −690 (gap edge: absolute death)
y= 41 logZ = −3.13
y= 89 logZ = −1.69 <-- geometric centre of the gap
y=129 logZ = −0.85 <-- actual maximum
y=137 logZ = −1.77
y=145 logZ = −690
Two things fall out with nobody writing them down:
- The alive band is the 130 px gap shrunk by exactly the body radius on each side. The agent discovered the size of its own body from pain and applied it.
- The maximum is not the geometric centre — it sits low. Flapping gives an instant −9 (strong upward authority); gravity gives +1/frame (weak downward). So low positions have far more surviving action sequences. "Maximise options" prefers the centre of the reachable set, weighted by control authority, not the centre of the hole.
That is a falsifiable prediction about behaviour, and it holds:
| plan objective | signed offset at pipe crossing (+ = below centre) |
|---|---|
| pure future-count (β=0) | +14.4 px |
| count + nociception (full) | +1.7 px |
So gap-centring is not option-counting in disguise. Counting alone flies low-biased; the nociception gradient is what cancels the actuation asymmetry and also what buys margin against residual model error (β=0 survives 1540 frames median vs 3000). The two innate terms have distinct, measurable, complementary roles.
Each of these was a real failure found by diagnosis, not a hypothetical.
- A ray distance is an upper bound on the true nearest-surface distance (corners between rays are invisible). Least-squares body calibration is therefore biased outward: the agent concluded it had radius 38 when it had 12, became too timid to enter any gap, and died. Fix: fit the lower envelope per pain level — a bound that never crosses the truth from below. Lesson: with a one-sided sensor, use a one-sided estimator.
- A forward-only ray fan is inconsistent with whole-body nociception. Pain from a pipe already passed had no ray pointing at it. Fix: sense in every direction the body can be damaged from. Lesson: proximity coverage must match damage coverage.
- Leaving the occlusion shadow unknown-and-free is fatal. Rays stop at surfaces, so the map held a pipe's 4 px front face and nothing behind it; the agent cleared the face and drifted into the 52 px body. Fix: fill the shadow along the ray — surfaces bound solid, extended bodies. Lesson: "OTHER is a thing" has to include thickness.
- An unbounded fall speed in the self-model looks conservative and is actually optimistic. Predicting harder falls means believing you can dive to a low gap faster than your body can. The agent committed to descents it could not finish. Fix: identify terminal velocity as a third self-parameter (a no-flap frame where velocity stops growing is the terminal velocity). Lesson: check which direction of model error is fatal, not which sounds cautious.
- Vertical-only clearance ignores the body's horizontal extent. The agent collides with a
wall
rpx before that wall reaches its column; the planner only checked the arrival frame. Fix: inflate the obstacle by the body over every column the body spans. Lesson: the configuration space is the body's, not the centroid's. - A filter whose variance has collapsed cannot correct itself. One noisy velocity
estimate pushed
vmaxto 15.16 permanently, and that seed died first. Fix: variance floors, plus refusing one-sided updates larger than one physically possible step. Lesson: on a robot the same floor is what lets damage be noticed at all. - Pure maximin degenerates while the posterior is wide. Some plausible world always says every future dies, which flattens the minimum and makes the action arbitrary — the agent died on frame 23. Fix: a tiny weight on the mean over worlds keeps the ranking informative. Lesson: robustness needs a graceful floor, or it becomes paralysis.
The three concepts are not literally sufficient. These are the extra commitments, stated plainly rather than hidden in the code:
- Object permanence and spatial extension. The map persists across frames and shadows are filled. Arguably contained in "OTHER is a thing", but it is a real inductive bias and removing it kills the agent in 223 frames.
- Finite-variance innate priors on
g, J, vmax, r, λ. With no prior at all the agent cannot act before its first observation. These are body-scale priors, not knowledge of the world. - The world is vertically bounded, and that extent is measurable from the up/down rays. Used as a safety fallback where the map has holes.
- Efference copy — it knows which action it emitted.
- Nociception is affine and monotone in the true distance. Used for calibration. The valence (pain is bad) is innate and unlearnable; only the scale is learned.
- Discretisation:
yat 1 px,vyat 1 px/frame, map at 4 px, horizon 75 frames.
All three remaining deaths have an identical, reproducible signature: log Z sits healthy
(−0 to −2, i.e. almost every action sequence survives) for ~20 frames, then drops to −691 (no
surviving future at all) in a single frame, with pose error 0.00 px and every parameter
correctly identified. A correct viability computation cannot do that unless the map changed —
the agent discovered occupancy it had been treating as free.
The obvious explanation is that counting futures through unobserved space credits ignorance as safety. Measured: over a whole life 20 % of planning-horizon columns contain no known obstacle at all, rising to 26 % at the collapse frame, concentrated in the far columns (t ≳ 60) — consistent with the story.
Three interventions aimed squarely at that explanation all failed:
| intervention | frames | deaths | rate |
|---|---|---|---|
| baseline (16 rays, unknown treated as free) | 66 135 | 3 | 1 / 22 045 |
| 32 rays — double the angular resolution | 42 429 | 4 | 1 / 10 607 |
| truncate the horizon at the knowledge frontier | — | — | myopic: died at frame ~100, map still filling |
mild discomfort on unobserved cells (ubeta=0.7) |
65 400 | 3 | 1 / 21 800 |
Doubling the rays did improve coverage (20 % → 12 % of columns looking empty) without improving survival at all. The frontier truncation traded false optimism for honest myopia and was strictly worse. The soft penalty changed nothing.
So the honest position: the collapse signature is real and reproducible; its cause is not established. And the experiment cannot settle it — with 3 deaths per condition the 95 % interval on the rate spans roughly a factor of three, so nothing smaller than a ~3× effect is detectable here. Establishing the cause needs either many more seeds or a targeted replay of the collapse frame with the map replaced by ground truth. That is the next thing to do, and it is deliberately not claimed as done.
What is established: with all seven fixes below in place, ~330 pipes of continuous survival comes out of three innate concepts plus online identification, with no training loop and no deaths spent on learning.
Transfers as-is: viability-kernel volume as the sole objective; identification by prediction error rather than reward; pessimism over a model posterior producing caution as a derived behaviour; a nociception gradient buying margin against model error; damage showing up as self-model prediction error (which needs the variance floors from finding 6).
Needs real work: continuous, high-dimensional action (the exact lattice DP has to become
sampling-based — MPPI/CEM — and the exactness of Z is lost); partial observability beyond
occlusion; and composition with a task objective. For that last one the recommended shape is
lexicographic: leaving the viability kernel is not a cost to be traded off, it is removal
from the candidate set; the task objective is maximised only inside what remains. That keeps
the minimal-variable structure — survive, avoid damage, achieve the goal — with survival as a
constraint rather than a competing reward.
| file | role |
|---|---|
env.py |
the world. Ground truth lives here and never leaves it |
models.py |
StateEstimator (self-localisation from rays), SelfModel (g, J, vmax), BodyModel (r, λ from pain), OccGrid (OTHER, with permanence and shadow fill), WorldModel (s) |
planner.py |
the Z recursion, the world ensemble, maximin |
agent.py |
glue + the ablation switches |
run.py |
python3 run.py --lives 5 --frames 3000 [--ablate no_pessimism] |
longrun.py |
ABLATE=... python3 longrun.py <seed> <cap> — one continuous life |
ablate.py |
python3 ablate.py 3 3000 — the ablation table |
render.py |
python3 render.py [seed] — world on the left, the agent's belief on the right |
probe_*.py |
the diagnostics that found findings 1–7, kept because they are the evidence |
Switches: no_self_model, no_option_value, no_pessimism, no_memory,
no_pain_gradient are the ablations in the table. frontier and unknown_penalty are the
two attempted fixes for the remaining deaths — both off by default because neither worked;
they are kept so the negative result is reproducible. N_RAYS=32 changes sensor resolution
(also no effect on survival).
Reproduce the headline number:
for s in 0 1 2 3 4; do python3 longrun.py $s 20000 & done; wait