Skip to content
marohanPublic

About

flappy_WM is an agent that survives continuously in a Flappy Bird environment without reinforcement learning (RL), reward channels, policy training, or episode iterations.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

flappy_WM — a world-model agent with three innate concepts

Survive Flappy Bird with no reinforcement learning, no reward channel, no policy training, and no repeated episodes to learn from. The agent embodies exactly three concepts:

concept computational primitive
SELF an indexical (which signals are my body) + an efference copy (what I just did) → gives (state, action) → state' pairs, so its own dynamics are identifiable by supervised prediction
PAIN / DEATH contact ends the future (absorbing), and nociception intensity is to be minimised. This is the entire objective. It is innate and not learnable — only its scale is
OTHER not-self occupied space, which persists, is spatially extended, drifts at a learnable speed, and hurts on contact

Everything else — gravity, its own flap strength, its own terminal velocity, its own body radius, its own pain scale, the world's speed, the world's shape — is identified online from prediction error. The reason no RL is needed is that the three concepts supply the objective. What RL spends millions of samples discovering (what is good) is given; what remains (what my actions do) is dense, self-supervised, and available every single frame.

What the agent receives, and what it never receives

Receives, per frame: 16 ray distances over the full circle from its body centre to the nearest not-self surface; a scalar nociception intensity in [0,1] whose mapping from distance is not given; and a dead flag. Nothing else.

Never receives: its own position or velocity (it infers them from rays), the score, any reward, any physics constant, the pipe geometry, or the fact that "pipe", "ground" and "ceiling" are different kinds of thing.

The single innate quantity

Exact backward dynamic programming on a (y, vy) × time lattice:

Z[H][s] = alive(s, H)
Z[t][s] = alive(s,t) · exp(−β · pain(s,t)) · ( Z[t+1][flap(s)] + Z[t+1][fall(s)] )
a* = argmax_a  min over plausible worlds  log Z[1][ successor(s₀, a) ]

log Z is a free energy: it counts how many futures from here stay alive, discounted by predicted discomfort. Death sets it to zero, so death is absorbing and dominant. Maximising it is staying deep inside the viability kernel — "keep the most ways to keep living". No preference for the middle of a gap is written anywhere.

Results

Identification, from ~60 frames of ordinary flight and zero deaths:

parameter truth identified
gravity g 1.0 1.000 ± 0.05
flap impulse J 9.0 9.000 ± 0.10
terminal velocity vmax 12.0 12.00 ± 0.20
world speed s 3.0 3.000 ± 0.03
body radius r 12.0 11.7 – 12.4 ± 0.7
nociceptive range λ 25.0 23 – 24

Survival, 5 seeds × 20 000-frame cap (one continuous life each, learning carried across deaths):

seed 0  died at  6 767   111 pipes
seed 1  survived 20 000  332 pipes
seed 2  died at  7 966   131 pipes
seed 3  died at 11 202   185 pipes
seed 4  survived 20 000  332 pipes

66 135 frames, 3 deaths → one death per ~22 000 frames (~364 pipes). A reactive contact-avoider that has all the same senses but no world model dies after ~100 frames. Cost: ~15 ms/frame, single-threaded numpy.

Every mechanism is load-bearing (3 lives × 3000-frame cap)

ablation deaths median frames max pipes
full model 1/3 3000 (cap) 48
− self-model identification 3/3 136 5
− option value (myopic, H=3) 3/3 163 1
− pessimism under uncertainty 1/3 3000 48 (but min 81 — one life died during bootstrap)
− object permanence (map memory) 3/3 223 10
− pain gradient in plan (β=0) 2/3 1540 48

Pessimism does not change the median at all — it changes the worst case, and only during the first ~100 frames while the self-model is still unidentified. That is exactly the role it was designed for, and it is invisible in any average.

The emergent behaviour is not the one that was expected

With the nociception term switched off entirely (β=0), so that value is nothing but the count of surviving futures, at the moment a pipe arrives:

y= 25  logZ = −690   (gap edge: absolute death)
y= 41  logZ = −3.13
y= 89  logZ = −1.69   <-- geometric centre of the gap
y=129  logZ = −0.85   <-- actual maximum
y=137  logZ = −1.77
y=145  logZ = −690

Two things fall out with nobody writing them down:

  1. The alive band is the 130 px gap shrunk by exactly the body radius on each side. The agent discovered the size of its own body from pain and applied it.
  2. The maximum is not the geometric centre — it sits low. Flapping gives an instant −9 (strong upward authority); gravity gives +1/frame (weak downward). So low positions have far more surviving action sequences. "Maximise options" prefers the centre of the reachable set, weighted by control authority, not the centre of the hole.

That is a falsifiable prediction about behaviour, and it holds:

plan objective signed offset at pipe crossing (+ = below centre)
pure future-count (β=0) +14.4 px
count + nociception (full) +1.7 px

So gap-centring is not option-counting in disguise. Counting alone flies low-biased; the nociception gradient is what cancels the actuation asymmetry and also what buys margin against residual model error (β=0 survives 1540 frames median vs 3000). The two innate terms have distinct, measurable, complementary roles.

Seven ways this broke, and what each one means for a robot

Each of these was a real failure found by diagnosis, not a hypothetical.

  1. A ray distance is an upper bound on the true nearest-surface distance (corners between rays are invisible). Least-squares body calibration is therefore biased outward: the agent concluded it had radius 38 when it had 12, became too timid to enter any gap, and died. Fix: fit the lower envelope per pain level — a bound that never crosses the truth from below. Lesson: with a one-sided sensor, use a one-sided estimator.
  2. A forward-only ray fan is inconsistent with whole-body nociception. Pain from a pipe already passed had no ray pointing at it. Fix: sense in every direction the body can be damaged from. Lesson: proximity coverage must match damage coverage.
  3. Leaving the occlusion shadow unknown-and-free is fatal. Rays stop at surfaces, so the map held a pipe's 4 px front face and nothing behind it; the agent cleared the face and drifted into the 52 px body. Fix: fill the shadow along the ray — surfaces bound solid, extended bodies. Lesson: "OTHER is a thing" has to include thickness.
  4. An unbounded fall speed in the self-model looks conservative and is actually optimistic. Predicting harder falls means believing you can dive to a low gap faster than your body can. The agent committed to descents it could not finish. Fix: identify terminal velocity as a third self-parameter (a no-flap frame where velocity stops growing is the terminal velocity). Lesson: check which direction of model error is fatal, not which sounds cautious.
  5. Vertical-only clearance ignores the body's horizontal extent. The agent collides with a wall r px before that wall reaches its column; the planner only checked the arrival frame. Fix: inflate the obstacle by the body over every column the body spans. Lesson: the configuration space is the body's, not the centroid's.
  6. A filter whose variance has collapsed cannot correct itself. One noisy velocity estimate pushed vmax to 15.16 permanently, and that seed died first. Fix: variance floors, plus refusing one-sided updates larger than one physically possible step. Lesson: on a robot the same floor is what lets damage be noticed at all.
  7. Pure maximin degenerates while the posterior is wide. Some plausible world always says every future dies, which flattens the minimum and makes the action arbitrary — the agent died on frame 23. Fix: a tiny weight on the mean over worlds keeps the ranking informative. Lesson: robustness needs a graceful floor, or it becomes paralysis.

Honest list of assumptions beyond the three concepts

The three concepts are not literally sufficient. These are the extra commitments, stated plainly rather than hidden in the code:

  • Object permanence and spatial extension. The map persists across frames and shadows are filled. Arguably contained in "OTHER is a thing", but it is a real inductive bias and removing it kills the agent in 223 frames.
  • Finite-variance innate priors on g, J, vmax, r, λ. With no prior at all the agent cannot act before its first observation. These are body-scale priors, not knowledge of the world.
  • The world is vertically bounded, and that extent is measurable from the up/down rays. Used as a safety fallback where the map has holes.
  • Efference copy — it knows which action it emitted.
  • Nociception is affine and monotone in the true distance. Used for calibration. The valence (pain is bad) is innate and unlearnable; only the scale is learned.
  • Discretisation: y at 1 px, vy at 1 px/frame, map at 4 px, horizon 75 frames.

The limit that remains — signature solid, cause NOT established

All three remaining deaths have an identical, reproducible signature: log Z sits healthy (−0 to −2, i.e. almost every action sequence survives) for ~20 frames, then drops to −691 (no surviving future at all) in a single frame, with pose error 0.00 px and every parameter correctly identified. A correct viability computation cannot do that unless the map changed — the agent discovered occupancy it had been treating as free.

The obvious explanation is that counting futures through unobserved space credits ignorance as safety. Measured: over a whole life 20 % of planning-horizon columns contain no known obstacle at all, rising to 26 % at the collapse frame, concentrated in the far columns (t ≳ 60) — consistent with the story.

Three interventions aimed squarely at that explanation all failed:

intervention frames deaths rate
baseline (16 rays, unknown treated as free) 66 135 3 1 / 22 045
32 rays — double the angular resolution 42 429 4 1 / 10 607
truncate the horizon at the knowledge frontier — — myopic: died at frame ~100, map still filling
mild discomfort on unobserved cells (ubeta=0.7) 65 400 3 1 / 21 800

Doubling the rays did improve coverage (20 % → 12 % of columns looking empty) without improving survival at all. The frontier truncation traded false optimism for honest myopia and was strictly worse. The soft penalty changed nothing.

So the honest position: the collapse signature is real and reproducible; its cause is not established. And the experiment cannot settle it — with 3 deaths per condition the 95 % interval on the rate spans roughly a factor of three, so nothing smaller than a ~3× effect is detectable here. Establishing the cause needs either many more seeds or a targeted replay of the collapse frame with the map replaced by ground truth. That is the next thing to do, and it is deliberately not claimed as done.

What is established: with all seven fixes below in place, ~330 pipes of continuous survival comes out of three innate concepts plus online identification, with no training loop and no deaths spent on learning.

Extending to a robot

Transfers as-is: viability-kernel volume as the sole objective; identification by prediction error rather than reward; pessimism over a model posterior producing caution as a derived behaviour; a nociception gradient buying margin against model error; damage showing up as self-model prediction error (which needs the variance floors from finding 6).

Needs real work: continuous, high-dimensional action (the exact lattice DP has to become sampling-based — MPPI/CEM — and the exactness of Z is lost); partial observability beyond occlusion; and composition with a task objective. For that last one the recommended shape is lexicographic: leaving the viability kernel is not a cost to be traded off, it is removal from the candidate set; the task objective is maximised only inside what remains. That keeps the minimal-variable structure — survive, avoid damage, achieve the goal — with survival as a constraint rather than a competing reward.

Files

file role
env.py the world. Ground truth lives here and never leaves it
models.py StateEstimator (self-localisation from rays), SelfModel (g, J, vmax), BodyModel (r, λ from pain), OccGrid (OTHER, with permanence and shadow fill), WorldModel (s)
planner.py the Z recursion, the world ensemble, maximin
agent.py glue + the ablation switches
run.py python3 run.py --lives 5 --frames 3000 [--ablate no_pessimism]
longrun.py ABLATE=... python3 longrun.py <seed> <cap> — one continuous life
ablate.py python3 ablate.py 3 3000 — the ablation table
render.py python3 render.py [seed] — world on the left, the agent's belief on the right
probe_*.py the diagnostics that found findings 1–7, kept because they are the evidence

Switches: no_self_model, no_option_value, no_pessimism, no_memory, no_pain_gradient are the ablations in the table. frontier and unknown_penalty are the two attempted fixes for the remaining deaths — both off by default because neither worked; they are kept so the negative result is reproducible. N_RAYS=32 changes sensor resolution (also no effect on survival).

Reproduce the headline number:

for s in 0 1 2 3 4; do python3 longrun.py $s 20000 & done; wait

About

flappy_WM is an agent that survives continuously in a Flappy Bird environment without reinforcement learning (RL), reward channels, policy training, or episode iterations.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages