Skip to content

Phase 1: integrate self-hosting progress and welcome new Aero users - #91

Merged
RobVanProd merged 51 commits into
masterfrom
agent/phase-1-integration
Sep 6, 2026
Merged

RobVanProd merged 51 commits into
masterfrom
agent/phase-1-integration

Conversation

@RobVanProd

@RobVanProd RobVanProd commented Sep 6, 2026 •

Copy link
Copy Markdown
Owner

Phase 1: integrate existing work, refresh the visitor experience, then stop

Protected merge; accepted-head replay in progress

PR #91 merged normally at df0fecea13e9f28c5065e092973aa34d5ac055c9 on
2026-09-06 at 04:33:04 UTC. Its tree is 92e70fd5d5308390efad95cb1c7dc9526b60a6a1,
identical to the reviewed candidate. Ordered parents are accepted base b987cd2
then candidate dfbb5d4. Local master is synchronized and tracked files are clean;
pre-existing branches, stash and untracked folders were preserved.

All 13 candidate checks passed before merge. Accepted-head
evidence capture
and CodeQL/security
completed successfully. Accepted-head CI
and Rust CI
remain in progress at handoff; no completed result is claimed for these replays.
The requested Phase 1 changes are merged and local master is synchronized.
Work stops here at the user's Phase 1 boundary, with post-merge verification
explicitly pending. Phase 2 is not started.

Candidate: dfbb5d4; accepted base: b987cd27f18df752e9c89dd951daae89ef5b854f.
Preserves and merges self-hosting branch 2d99ca7e3f791295db1d1fd4f933cfb650b05e10
(47 previously unmerged commits). The source history is retained, not squashed or rewritten.

Included

  • CORE-093 entry-block allocation repair and its original regression tests.
  • H1A self-source ingestion, bounded H1B grammar/capacity work, CAP-056/057 module
    parsing, and CAP-058 bounded multi-function semantic and checked-IR construction.
  • Existing off-system-drive gate output and bounded build/test parallelism controls.
  • A welcoming README with the tested quick start, navigation, contribution guidance,
    clear current limitations, and collapsed historical evidence.
  • Current integration headers on the state, capability, backend, and roadmap documents,
    preserving historical assertions and evidence rather than rewriting old checkpoints.

Explicit boundary

The compiler source can be parsed, but still fails semantic analysis. Small
multi-function probes produce checked IR and still fail verification. Connected
syntax representation and full self-source semantics/emission are not complete.
CAP-059/H1M-3 remains a contract only. No phase 2 implementation, self-hosting,
general standard library, release, stability, safety, performance, or GPU execution
claim is added. Compiler production and Aero examples match 2d99ca7 exactly.
Validation now also covers the hoisted workflow layout, with companion assertions
adapted to entry allocation plus unchanged dataflow rather than inline allocation.

Completed pre-merge verification and integration history

  • Reviewed the original CORE-093 red-first failures and retained all regressions.
  • Documentation replay completed with 31 passed and exit 0 before the final status
    headers. The original missing experimental-label failure was fixed in prose only.
  • GitHub Markdown rendering checked: visitor content, quick start, navigation, and
    two balanced collapsed historical sections.
  • Focused stack-allocation tests: 3/3 passed. Self-source ingestion/module tests:
    63/63 passed. The unchanged compiler production is still the source-branch tree.
  • Candidate 67ba195 exposed stale scorer/inference allocation-order assumptions
    in Rust CI 34007823211. The repair preserves all dependency identities and
    native assertions while requiring the moved allocations in the entry block.
    New real-output/mutation regression tests: 2/2 passed. The exact Windows
    fixed-array workflow replay passed at O0/O2. This is an integration-oracle repair,
    not a compiler or phase-2 change.
  • Complete repository-root validation on dfbb5d40c7805866f2cd571fc37c83dbeed8b3a4,
    tree 92e70fd5d5308390efad95cb1c7dc9526b60a6a1, completed with exit 0:
    1,018 passed, 0 failed, 16 existing ignored, 118 targets, plus formatting and
    correctness Clippy. LLVM 22.1.8 was explicit and generated output stayed on D:.
    No tracked file changed during that final replay. Required public gates subsequently passed before merge.
  • The first local root replay and CI 34008617659 stopped at two obsolete
    companion text fragments (19/20 fixed-array tests passed). Those fragments now
    require both entry placement and the original adjacent dataflow. The completed
    focused replay is 53/53 passed across six targets. The final full gate is
    green on dfbb5d4; earlier failed/cancelled candidates are not acceptance.
  • All 13 candidate checks completed successfully: stable/nightly, Windows native
    job 101421565365, both compiler CI jobs, CodeQL aggregate 101421618872,
    all four language/security analyses, and evidence capture/aggregate 101422168941.
  • Required protections passed normally; no direct master push or bypass was used.

The protected merge identity and completed/pending replay statuses are recorded
above. The visitor README was checked through GitHub Markdown rendering. Local
master is synchronized. Phase 2 will be discussed separately with the user;
post-merge CI completion is the only outstanding verification item.

codex and others added 30 commits August 17, 2026 17:18
CORE-093. The code generator emitted each value's storage slot inline at the
point the value was produced, so every checked ByteBuffer result temporary
inside a loop became a non-entry alloca. LLVM never reclaims one of those before
the function returns and mem2reg cannot promote it, so an Aero loop over a
ByteBuffer grew the stack once per iteration.

The accepted bounded corpus never exposed this because its canonical input is 34
bytes. Feeding the compiler its own 241,918-byte source terminated with
STATUS_STACK_OVERFLOW before any diagnostic: the emitted module placed 1,035 of
its 1,116 allocas outside the entry block, 423 of them inside loop bodies.

Every alloca this generator emits has a static type, a constant alignment, and
no dynamic element-count operand, so generate_hoisted_function_body now lifts
them all to the top of the entry block. Relative order is preserved, no
instruction text changes, and an alloca carrying a dynamic count never moves.

The tracked examples/loop_stack_stability specimen survives 400,000 checked
bytes_push and 400,000 checked bytes_get operations at O0 and O2. The structural
rule is asserted for the specimen and all eight accepted .aero products, each of
which still passes required LLVM verification.

Five digest sentinels pin MD5 hashes of emitted LLVM and therefore moved. Before
re-freezing them, each of the four distinct programs they cover was compiled by
a binary built from the exact pre-change code_generator.rs and again by the
fixed binary: for all four the line multiset, the alloca line multiset, and the
relative order of every non-alloca line are identical. Placement is the only
difference. No assertion was removed, relaxed, or skipped.

Co-Authored-By: Claude Opus 5 <[email protected]>
CAP-049 / H1A, the first H1 prerequisite. examples/aero_self_host_v0/compiler.aero
is a copy-derived successor of accepted CAP-047/B1C differing only in six
mechanically reconstructable ways: three raised ingestion bounds (1,048,576
source bytes, 262,144 token records, 16,384 names), a new lexical token kind 37
for a lone `&`, the matching token-record validator bound, and one
quadratic-to-linear rewrite of the located-token re-derivation.

Fed its own exact bytes the compiler now consumes all 241,918, interns 571
names, and records 31,062 located token records, then stops at the
independently predicted first unsupported parser construct: status 10 at offset
16, line 1, column 17, expecting `)` and finding an identifier. That is the
`result` parameter of `fn result_value(result: Result<int, int>)`, the first
construct outside the frozen `fn NAME ( ) -> int { return` skeleton. Every
downstream phase reports not-attempted, so all 67 independently derived
expectation values match at O0 and O2.

The focused target derives every expectation from its own oracle rather than
from observed Aero output, and proves the boundaries in order: the accepted
product stops at byte 8,192 having consumed exactly 8,193; raising only that
bound reaches the lone `&` of bytes_len(&source) at offset 17,681; admitting
that one lexical form completes the stream. The accepted 34-byte canonical
program is preserved exactly - exit 91 and the identical 144-byte module at O0
and O2 - and self-input produces no output byte and no artifact.

Reaching this required the separately scoped CORE-093 code-generator fix. H1A is
ingestion and tokenization only: the compiler reads its own source, it does not
parse, type, check, verify, or lower it. This is not H1B, H1, H2, stage
convergence, or any self-hosting claim.

Co-Authored-By: Claude Opus 5 <[email protected]>
Ledger entries for CORE-093 and CAP-049 with their red and green checkpoints,
the exact measured boundaries, the re-frozen digest justification, and the
complete gate result: ./tools/test.sh exits zero with 312 library tests, 36
binary tests, all 117 integration/native/system targets, and doc tests green.

BOOTSTRAP_CONVERGENCE_READINESS.md records the moved boundary and decomposes
H1B from H1A's token census. The self-source grammar is now measured and closed:
23 fn items, 469 let bindings, 935 if, 82 while, 2,756 assignments, 417
references, and exactly one match - with no `[`, `.`, `%`, or `!` token anywhere,
so H1B needs no array syntax, field access, modulo, or negation.

It also records a boundary that is not the parser's alone: the semantic,
checked-IR, verifier, and emitter phases all assume exactly one function, while
the canonical source has 23. Admitting a second fn item changes four downstream
authorities at once and gets its own ordered gate rather than being absorbed
into a parser checkpoint.

Co-Authored-By: Claude Opus 5 <[email protected]>
CAP-050 freezes the first H1B checkpoint: the parameter-list grammar the
canonical source actually uses, with its exact predicted stop.

Both were measured rather than assumed. The 23 signatures all return int and
declare 99 parameters, of which 98 are int and exactly one is Result<int, int>;
two take none and the widest takes 67. No parameter is a ByteBuffer or a
reference - those appear only as local binding types and call arguments - so
reference syntax belongs to the later call checkpoint, and the readiness table
is corrected accordingly.

The predicted stop is derived from the canonical token stream, not guessed: with
the parameter list admitted, tokens 3 through 14 parse, the body's leading
`match` identifier reduces into one name-reference node, and the frozen `; } EOF`
closing sequence rejects the identifier `result` at offset 68, line 2, column 18,
expecting `;`.

The contract also freezes what a parameter may not become. The semantic,
checked-IR, and verifier phases require root == node_count, one symbol, and one
fact per node, so emitting a syntax node for a parameter would silently cross
four downstream authorities. Parameters get their own bounded store instead.

Co-Authored-By: Claude Opus 5 <[email protected]>
CAP-050's red-first requirement is that the next checkpoint's stop is predicted
from the canonical token stream, not read off the parser once it changes. This
adds that derivation as a passing oracle test so the implementation cannot be
graded against its own output.

The oracle now models the frozen signature grammar - `fn NAME ( params? ) -> int {`
where a parameter is `IDENT : TYPE` and TYPE is `int` or the exact sequence
`Result < int , int >` - plus enough of the accepted expression grammar to reach
the frozen `; } EOF` closing sequence.

Applied to the canonical source it yields the exact target: with signatures
admitted, `fn result_value(result: Result<int, int>) -> int {` parses, one
parameter is recorded, `return match result {` reduces the `match` identifier to
a single name-reference node, and the closing sequence rejects the identifier
`result` at offset 68, line 2, column 18, expecting `;`. Parameters produce no
syntax node.

The existing test still proves the parser stops at the CAP-049 boundary today,
so this freezes a target rather than claiming progress.

Co-Authored-By: Claude Opus 5 <[email protected]>
An H1B-1 implementation was written, exercised, and reverted rather than
committed, so the tree stays at the accepted CAP-049 product. This records what
it proved and what remains, so the next attempt does not repeat the work.

Proven: the parameter sub-machine fits inside the existing skeleton_step == 3
slot with no new token-read state pair, and the modified source checks,
compiles, verifies, links, and still returns 91 with the identical 144-byte
canonical module. That last result validates the new 989 separator, the
parameter-region fold, the recomputed canonical checksum 810191, and the
68-value entry point - a single wrong word there would have returned 80.

Remaining: self-input returns 80. Because the canonical run passes, the
divergence is confined to values only self-input produces - the node record the
leading `match` identifier appends, the resulting node_count, or the located
closing-sequence diagnostic - not the parameter store or checksum layout.

The recommended next step is field-level resolution inside the parse group: the
return codes separate phase groups but not fields, so a harness that tries
several candidate vectors in one linked binary would isolate it in a single
compile instead of repeated whole-suite runs.

Co-Authored-By: Claude Opus 5 <[email protected]>
The diagnostic harness recommended in the previous entry was built and run: it
embeds many complete expectation vectors in one linked binary, calls the entry
point once per candidate with the stream reset between calls, and reports which
one the product agrees with. One compile covers the grid; 46 candidates run in
about a minute.

None of the 46 matched. The grid covered every stop position from token 3 to
token 18 with the code the frozen grammar expects at each, statuses 10 and 12 at
each position, internal statuses 16/14/8/13, node counts 0 and 1, parameter
counts 0 and 1, and both parameter type codes.

That is still progress, because it narrows the cause. The canonical 34-byte run
returning 91 already proved the checksum layout, the 989 separator, the
parameter fold, and the 68-value entry point are correct. A null result across
one-at-a-time variations means at least two parse-group fields differ
simultaneously, which those families cannot express.

Next: a genuine cross product of stop position, status, node count, parameter
count, and type code - about 2,000 candidates, still one compile. If that is
also empty, the divergence lies in a value every candidate held fixed, so the
product's ingestion of the enlarged source should be re-checked directly.

Co-Authored-By: Claude Opus 5 <[email protected]>
The cross-product probe was built and run: stop position over tokens 3 to 20,
status in 10/12/16, node count 0 and 1, parameter count 0 and 1, both parameter
type codes, and five diagnostic codes at each position - roughly 2,000
internally consistent candidates in one compile. It matched nothing.

That is a real narrowing rather than another dead end. Among the candidates was
the exact shape of "the parameter sub-machine never engaged" - stop at token 3,
status 10, code 11, zero nodes, zero parameters, the accepted H1A behavior. Its
failure rules out the entire family of parser-stop explanations.

Every candidate held name_count, token_count, root, the parameter record's name
id, and the source/name/token word streams fixed, so the divergence is in one of
those. Together with the canonical 34-byte run still returning 91, that points at
ingestion or token-record production for the enlarged 250,370-byte source
differing from the oracle - which the canonical program is too small to expose,
and which the CAP-049 product-level ingestion check would have caught had it not
been the test replaced during the attempt.

Next: restore that ingestion assertion against the modified source first, then
bisect the six source edits, and only then return to the grammar.

Co-Authored-By: Claude Opus 5 <[email protected]>
…ference

A store-only variant was built and run: the parameters owner, the 68th
expected_parameters value, the 989 checksum region with its validation, and the
parse-group comparison - but none of the parser sub-machine.

It passes. The product ingests its own modified 243,693 bytes, stops at the
unchanged H1A construct, and matches all 68 expectation values. The canonical
34-byte program still returns 91 with its exact module.

That corrects the previous entry. Ingestion, token-record production, the
parameter store, its validation, the checksum region, the widened entry point,
and the recomputed canonical constant 810191 are all proven correct. The claim
that ingestion diverged was wrong. The defect is confined to the roughly sixty
lines of param_mode dispatch, alternation, type matching, and advance added to
skeleton_step == 3.

It also explains the null cross-product result: the grid only covered stop
positions through token 20. A mis-advancing sub-machine does not stop early - it
accepts tokens it should reject and runs deep into the body, far outside the
grid.

Next: land the store-only variant as CAP-050a (green apart from the exact-equality
derivation test, which needs the six store transformations added), then add the
sub-machine alone on that proven base so any failure is unambiguously the grammar.

Co-Authored-By: Claude Opus 5 <[email protected]>
The store-only bisection is landed as its own checkpoint. The canonical source
gains a parameters owner, a parameter_count counter, the 68th
expected_parameters value and its guard, a validated 989 checksum region, the
parse-group comparison, and the two canonical vector constants - but no parser
rule changes.

The source stays exactly reconstructible from accepted B1C: six CAP-049
ingestion differences plus seven CAP-050a store differences, asserted
byte-for-byte, so the diff cannot widen silently.

Evidence: the focused target passes 9/9. The product ingests its own complete
243,693 bytes and matches all 68 expectation values at O0 and O2, and the
accepted 34-byte canonical program still returns 91 with the identical 144-byte
module. The recomputed canonical checksum 810191 is derived as
step(step(586661, 989), 0), not observed.

This is proven infrastructure, not a capability: the store records zero
parameters because no rule produces one yet. Separating it means that when the
sub-machine is added, any failure is unambiguously the grammar.

Co-Authored-By: Claude Opus 5 <[email protected]>
The mode transitions were traced token by token against the canonical source's
first signature and are correct at every step, and the empty-parameter case is
already proven by the canonical program. So the defect is most likely not in the
transition table but in its interaction with the surrounding skeleton block: the
shared param_alternate rejection bypass, the reuse of lexer scratch registers
b0-b5 and word/push_result inside a parser state, or the changed skeleton_step
advance condition.

Recorded so the next attempt instruments those three rather than re-reading the
transitions.

Co-Authored-By: Claude Opus 5 <[email protected]>
Re-applied on top of the accepted CAP-050a base rather than all at once. The
canonical 34-byte program still returns 91 with its exact 144-byte module at O2,
so mode 0 seeing a closing parenthesis and completing an empty parameter list
works with the sub-machine present. Only the nonempty path remains unverified.

That isolation is exactly what splitting CAP-050a out was for, and it holds.

Co-Authored-By: Claude Opus 5 <[email protected]>
The canonical self-source is a single opaque pass/fail, which is why the
first CAP-050 attempt burned two probe grids and matched nothing. Nine
complete probe programs now exercise one signature-grammar rule each and
stop inside the parse phase, so a sub-machine defect localises to one rule.

The oracle is extended rather than duplicated: Ingestion carries the
parameter store and the node arena, parse_checksum folds both, and
signature_parser_stop is the single model of the parser CAP-050
authorizes, with the accepted signature_grammar_stop projected out of it.

Red-first. All nine probes run against the real linked product and return
91 against the accepted CAP-049 boundary; the CAP-050 target for the same
bytes is derived separately and not yet claimed.

Co-Authored-By: Claude Opus 5 <[email protected]>
Between the `(` and the `->` the parser now accepts either an immediate
`)` or a nonempty `IDENT : TYPE ( , IDENT : TYPE )*`, where TYPE is the
identifier `int` or the exact sequence `Result < int , int >`. Each
parameter appends one record to the CAP-050a store. No syntax node is
created for a parameter, so no downstream authority is crossed.

Four transformations inside the frozen skeleton block: the mode latched
once per token (as the driver already latches parser_state), a mode-driven
expected-kind table with one `param_alternate` second admissible kind
cleared on every token, the transitions and inline store append, and a
`param_hold` that suppresses the skeleton_step advance while the list is
open.

The oracle also gained the parser's origin sidecar. The semantic group is
never entered here but still folds origins and reports origin_count, and
the model assumed zero because H1A never produced a node. No Aero source
changed for that.

Fed its own source the product records one parameter, reduces the leading
`match` identifier to one node, and stops at offset 68, line 2, column 18
expecting `;`, matching all 68 derived values at O0 and O2. The canonical
34-byte program still returns 91 with the identical 144-byte module.

Co-Authored-By: Claude Opus 5 <[email protected]>
The ledger claimed the latching hazard was consistent with the first
attempt's split between the working empty path and the failing nonempty
path. That is unverifiable: the prior sub-machine was applied and reverted
twice and never committed, so it survives in no git object. All 1,066
dangling blobs were scanned; one holds a pre-CAP-042 ancestor of the
parser, none holds param_mode or param_cycle_mode. A plausible story
sitting in the ledger as history is worse than a gap, because a later
session reads it as established and narrows toward it.

Also adds the session handoff near the top of the CAP-050 section: base
commit, what was proved by running, what was ruled out, and the probe
table as the instrument H1B-2 should extend before touching compiler.aero.

Co-Authored-By: Claude Opus 5 <[email protected]>
Ledger-first contract for the next checkpoint, per
BOOTSTRAP_CONVERGENCE_READINESS.md:289. Not started: no product change is
authorized until the red-first oracle derivation it specifies is done.

Resolves both ambiguities rather than leaving them for the next session.

Ambiguity 1, the naming rule: the rule holds in substance, since `match`
is what forces the checkpoint and the stop sits one token later only
because `match` was consumed as a name reference. Refined in one respect:
H1B-2 must not retract the node CAP-050 created, because the arena is
append-only and origin_count != node_count is a hard failure, but that
node does not survive into H1B-2's parse either. The mechanism is to
dispatch before appending, not to append and undo.

Ambiguity 2, node kinds: recorded as open, with the observation that
1..=19 are fully allocated and the `> 19` bound is in the parse group,
so raising it is inside the parser's authority at this checkpoint's stop
but creates a debt H1C pays. Flagged as an observation about the code,
not something the documents answer.

Candidate stop derived from the canonical bytes: the second `fn` item at
offset 146, line 8, column 1, expecting EOF and finding `fn`. Offered as
a prediction to confirm, explicitly not frozen, because freezing it
without the oracle would violate the red-first requirement.

Also records the branch push as durability only: not published, not
accepted, no capability claim, no PR, ci.yml the only workflow that runs.

Co-Authored-By: Claude Opus 5 <[email protected]>
Dispatch on the leading token of the return expression, before the operand
reduction runs, so the identifier `match` opens the admitted construct
`match IDENT { IDENT ( IDENT ) => EXPR , IDENT ( IDENT ) => EXPR , }` instead
of reducing to a name-reference node. The node arena is append-only and every
append is mirrored by an origin record, so nothing is appended for `match` and
nothing has to be retracted.

Ambiguity 2 from the contract resolves to the cheapest option: the construct
appends no node of its own and the arm bodies go through the already-accepted
expression grammar, producing four nodes of existing kinds on the canonical
source. The `kind <= 0 || kind > 19` validator bound is unchanged and no
authority is widened.

Red-first: the oracle was extended and exercised before the parser moved, and
the red was observed - 11 passed / 2 failed, the two product-graded CAP-051
targets returning 80. Thirteen focused probes were hand-derived from the frozen
grammar and confirmed against the oracle by a test that touches no product; all
thirteen agreed on the first attempt.

Fed its own bytes the compiler now parses the whole body of its first function
and stops at the independently derived next construct: status 10, offset 146,
line 8, column 1, expecting end of input and finding the second `fn` item, with
four nodes and one parameter. That stop is the expected result, not a defect.

The full repository gate is green: 117 test binaries, zero failures. The
canonical source stays exactly reconstructible from accepted B1C, asserted byte
for byte, and the accepted 34-byte canonical program still returns 91 with the
identical 144-byte module at -O0 and -O2.

Not H1B completion, H1, H2, stage convergence, or any self-hosting claim.

Co-Authored-By: Claude Opus 5 <[email protected]>
Ledger-first, no product change. Three things it settles that the readiness
document does not.

The decision CAP-051 left open is derived, not asserted: the statement grammar
owns the return statement. Keeping return outside it still forces the frozen
skeleton's fixed `return` step to dissolve, so it saves nothing; treating
statements as a prefix before a terminal return fits all 23 canonical functions
but is falsified by function 2, where a return sits inside an `if` with a
statement after it, so it would be undone at H1B-4. Under the adopted design
`;` is demoted from a closing token to the return statement's own terminator,
and CAP-051's two closing-sequence entry points collapse to one.

Ambiguity 1: this checkpoint cannot move the canonical stop and must not pretend
to. CAP-051 stops at the second `fn` item, which every parser checkpoint is
excluded from admitting, and function 1 is now parsed completely, so H1B-3,
H1B-4 and H1B-5 all leave the stop at offset 146. Forward evidence is entirely
focused probes; the self-ingestion target becomes a regression guard and may not
be cited as progress.

Ambiguity 2: the document's claim that the checkpoint order is the one the
source forces is false as measured - function 2 opens with `if`, and zero of the
23 functions have statements without control flow or a call. The order still
stands, on grammar dependency instead, and this contract records the corrected
justification rather than reordering the table.

Ambiguity 3 is left open with evidence for the next session: a statement
sequence probably does need a new node kind, which is the opposite of CAP-051's
answer, and overloading an arithmetic kind is not a cheap alternative because
the origin sidecar records the token kind that produced each node.

CAP-051's four orphan arm-body nodes are carried forward explicitly so H1B-3
cannot adopt them silently: any chaining that walks the arena by index would
sweep them in, so the canonical assertion must keep stating node_count == 4.

Measured statement grammar: 502 typed initialised bindings (471 `let mut`), 2,298
assignments, every target a bare identifier, and `int` the only binding type
whose initialiser is ever call-free.

Full repository gate green: 117 test binaries, zero failures.

Co-Authored-By: Claude Opus 5 <[email protected]>
A function body is now `{` followed by one or more statements followed by `}`,
and a statement is exactly one of `let IDENT : int = EXPR ;`,
`let mut IDENT : int = EXPR ;`, `IDENT = EXPR ;`, or `return EXPR ;`. The
skeleton's fixed `return` step is dissolved into the statement loop and `;` is
demoted from a closing token to the return statement's own terminator, so
CAP-051's two entry points into the closing sequence collapse into one rule
inside the loop and that sequence shrinks to `}` then end-of-input, entered
once. Three new parser states carry it: 45 dispatches a statement on its own
leading token, 47 runs the binding and assignment sub-machine, and 49 is the
shared terminator both the ordinary return expression and the closed match
construct return to.

Ambiguity 3 from the contract resolves to no new node kind and no raised bound,
and it was written before the parser was edited as that section required. A
statement produces no syntax node, exactly as a CAP-050 parameter does not. The
derivation is not aesthetic: the frozen canonical assertion of four nodes
forbids appending a statement node at `;`, so any statement node must defer to
the module's end, which only a complete parse reaches - and no multi-statement
program can complete a parse while the semantic phase is outside this
checkpoint's authority. Every line of sequence-building code would therefore
have been unreachable by every test available here. A binding's or an
assignment's initializer nodes are unowned, joining the four CAP-051 left; H1C
adopts them.

This checkpoint deliberately does not move the canonical self-ingestion stop.
Function 1 already parses completely and admitting a second `fn` item is
excluded from every parser checkpoint, so nothing CAP-052 admits is reachable in
the canonical source at all. The stop stays at status 10, offset 146, line 8,
column 1, expecting end of input and finding the second `fn` item, with four
nodes and one parameter, and it is asserted as a regression guard rather than
cited as progress.

Red-first: the oracle statement model and the shared `parse_expression`
extraction landed before the parser moved, with the ten CAP-050 signature probes
and thirteen CAP-051 match probes staying green against the refactored oracle.
Eighteen focused statement probes - six positive shapes and twelve negatives -
were hand-derived from the frozen contract and confirmed against the oracle by a
test that touches no product; all eighteen agreed on the first attempt, no hand
derivation needed correction. The red was then observed from the real linked
product, which returned 80 for them before `compiler.aero` was edited.

The full repository gate is green: 114 test binaries, zero failures. The
canonical source stays exactly reconstructible from accepted B1C, asserted byte
for byte, and the accepted 34-byte canonical program still returns 91 with the
identical 144-byte module at -O0 and -O2.

Not H1B completion, H1, H2, stage convergence, or any self-hosting claim.

Co-Authored-By: Claude Opus 5 <[email protected]>
Adds two things to the CAP-052/H1B-3 record that were missing from it, and
changes no product.

The gate figure. The complete repository-root gate is green on the accepted
tree: 117 test results - 114 integration binaries, the `src/lib.rs` and
`src/main.rs` unit targets, and the doc-test target - with zero failures. This
is an addition rather than a correction, because no record here has ever carried
a gate count at all: 114 exists only in commit `084cb1a`'s own message and 117
only in commit `6a2278e`'s. The two describe the same shape of run under
different denominators, so what is reconciled here is two commit messages
against each other, not a record against a run. The count also did not move
between `25fa375` and `084cb1a` - `src/compiler/tests/*.rs` is 114 files at both
commits and the diff lists no deletion - so no test was weakened, skipped, or
deleted. `084cb1a` is left unrewritten: it is pushed, and rewriting published
history is forbidden here.

The provenance. Three duplicate sessions were accidentally started on this
worktree at once. One wrote the whole CAP-052 implementation and was blocked
before it could commit, as was a second; a third adopted the uncommitted tree,
verified it rather than trusting it, and committed it as `084cb1a`. The
adoption was checked by structural inspection - parser states 44 through 49 and
each new register appearing exactly once, all ten patch constants defined once
and used once, no superseded marker surviving, no pre-existing test removed, and
`SIGNATURE_PROBES` and `MATCH_PROBES` byte-identical to `25fa375` - and settled
by `canonical_self_host_source_is_a_copy_derived_successor`, which asserts the
canonical source is exactly accepted B1C plus the enumerated deltas byte for
byte and which interleaved writes from two authors could not pass. Recorded so a
later reader does not find an implementation whose author left no report and
have to guess why.

Co-Authored-By: Claude Opus 5 <[email protected]>
Records a second, independent execution of the complete repository-root gate on
the accepted tree at 1066e83, from a clean working tree byte-identical to the
pushed commit. It reproduces the recorded figure exactly - 117 test result
lines, 969 passed, 0 failed, 16 ignored, exit 0 - with fmt and correctness
clippy green ahead of it, and self_host_source_ingestion_tests 16/16 including
the byte-for-byte B1C reconstruction and the unmoved canonical stop.

Also records the structural checks made on the adopted tree before the gate was
spent, since CAP-052 was written by one session and committed by another during
an accidental three-session overlap: states 44-49 each defined once, no
duplicated test name, 23 fn items, clean termination, and the benign brace-count
imbalance explained by three comment lines that quote a brace.

Changes no product and no test. The recorded gate figure is now reproduced
rather than trusted.

Co-Authored-By: Claude Opus 5 <[email protected]>
Measures what BOOTSTRAP_CONVERGENCE_READINESS.md:310 says H1B-6 raises the
bounds to, which had never been measured, and authors the H1B-4 control-flow
contract ledger-first ahead of any parser edit.

The measurement, over the 264,163-byte canonical source: 13,190 node records,
13,144 value records and 4,157 operator records as a measured floor, and
23,509 / 14,697 / 5,710 once the shapes H1B-4 and H1B-5 admit are costed. The
512 bound is exceeded by 11x to 51x.

Two findings that change how the bound should be read. value_records and
operator_records are never decremented, so each counts every push over the whole
parse rather than stack depth; the deepest either stack reaches on the complete
source is 5. And a fourth literal 512 lives at compiler.aero:4852 in the
verifier group, which BOOTSTRAP_CONVERGENCE_READINESS.md:246-248 forbids H1B to
widen, so H1B-6 should raise the three parse-group bounds and record that one as
debt.

The pull-forward rule does not fire at H1B-4 or H1B-5: both leave the canonical
stop at offset 146 with four nodes and are proven by focused probes. It fires at
the module-shape gate, where a 512 bound exhausts the node arena inside function
8 at line 154 of 6,085.

CAP-053 resolves two ambiguities before any parser edit. No node kind is added
and the 1..=19 bound is untouched, because an `if` node could reference its
condition but not its body and would assert at H1C that the conditional has
none. Nested blocks use a linked block-record store, a fourth monotonic counter
whose canonical requirement (1,197 records, peak depth 10) is measured here so
H1B-6 covers it.

Records two corrections to CAP-052. Its assignment figure of 2,298 is 2,432,
reached three independent ways; seven of its eight other measured figures
reproduce exactly. And its frozen rule that "after the return statement's `;`
the only admissible token is `}`" is not implemented - body_root is written at
compiler.aero:1791 and read only at :2470 and :2507 - so H1B-4 implements it per
block and confirms it red-first.

No product change. The complete repository-root gate is green on this tree:
117 test result lines, 969 passed, 0 failed, 16 ignored, process exit 0.

Co-Authored-By: Claude Opus 5 <[email protected]>
…ty figures

BOOTSTRAP_CONVERGENCE_READINESS.md:223 defines H1B as emitting a validated flat
AST for every construct actually present in compiler.aero. Five of the six H1B
checkpoints admit a construct without representing it - the parameter, the match
construct, the binding, the assignment, the statement sequence, and under
CAP-053 the conditional and the loop - and only H1B-5 is scheduled to create a
node. Each deferral is defensible on its own record; the effect compounds.

Measured on the complete canonical source: the accepted accounting produces
13,190 node records of which 154 are reachable from a root. 13,036 are orphans,
98.8%. run_runtime_ascii_llvm_emitter produces 12,020 nodes of which 3 are
reachable. Function 1 parses completely and all four of its nodes are orphans. A
representation that discharged :223 needs 23,509 nodes, so the obligation is
10,319 nodes: 4,186 sequence positions, 2,505 assignments, 1,553 calls and
references, 1,026 conditionals, 512 bindings, 252 else arms, 201 returns and 84
loops.

:223 is not wrong and is not weakened here. The checkpoint table at :324-329 is
missing a row, and no checkpoint in it can discharge :223. Recorded rather than
decided: H1B is not complete when H1B-6 is green; an explicit representation
checkpoint belongs after H1B-5 rather than absorbed into H1C, for the reason
:367 already refuses to absorb the single-function coupling; H1B-6 must precede
it, because a representing parser is the one that produces 23,509 nodes; and the
orphan census is that checkpoint's acceptance criterion.

Two corrections to the capacity measurement in 95a6aa8, left visible rather than
restated. The accepted parser appends its kind-18 return node once per function
at compiler.aero:2507, not once per return statement, so the node floor is
13,190 rather than 13,391; the 201-node difference moves into the projected
column, where a full representation needs it, and the projected total of 23,509
is unchanged. And operator records exceed 512 by 8x to 11x, not the 11x to 14x
first recorded: 4,157 / 512 is 8.1.

No product change; compiler.aero and the ingestion tests are byte-identical to
f416067. CAP-053's scope is unchanged. All readiness line citations rebased
after the insert shifted them. The complete repository-root gate is green on
this tree: 117 test result lines, 969 passed, 0 failed, 16 ignored, exit 0.

Co-Authored-By: Claude Opus 5 <[email protected]>
CAP-053/H1B-4 is authored as a contract and is not implemented. This is a
deliberate stop rather than an exhausted one: the H1B-4 product edit needs five
parser edits plus an oracle extension and a probe table ahead of it, and an
uncommitted parser tree is the failure mode that has already lost three sessions
in this worktree. compiler.aero and self_host_source_ingestion_tests.rs are
byte-identical to f416067, verified by hash before each of this session's three
commits.

Records for the next session: the base commit; what is proved, which is
measurement only and no control-flow behaviour; what is ruled out, with its
number - capacity is not H1B-4's or H1B-5's problem, no new node kind is needed,
and CAP-052's unreachability argument cannot be reused because a control-flow
node would be appended where probes assert exact counts; the four ordered next
steps, beginning with confirming the oracle extension is behaviour-preserving
against all forty-one existing probes before a single new one is written; and the
cost data - a 3.5-minute focused cycle, a 30-to-35-minute gate, and the four
environment variables the gate needs.

Also names the one claim in the contract still resting on reading rather than a
run: that CAP-052's stated rule about a statement following a completed `return`
is unenforced. Two of the required probes are expected red for that reason
rather than for the ordinary one, and observing that difference is the
confirmation.

No product change. The complete repository-root gate is green on this tree:
117 test result lines, 969 passed, 0 failed, 16 ignored, exit 0 - the third such
run this session, all three independent.

Co-Authored-By: Claude Opus 5 <[email protected]>
…ammar

CAP-053/H1B-4. Two control-flow statement forms are added to CAP-052's four,
each over the already-accepted expression grammar with no new expression form
and no new node kind. The `1..=19` node-kind bound is untouched: an `if` node
would have to reference a statement sequence that has no representation in the
accepted arena, and one carrying only its condition would assert at H1C that
the conditional has no body.

The block stack is a fourth bounded parse-group arena, appended through
`parser_append_target = 5` and read through `parser_record_target = 3`, one
three-word record per nested block holding its kind, the enclosing block's
statement state, and the link to the enclosing record. It carries the same 512
bound and the same `status = 15`, `diagnostic_code = 512` exhaustion diagnostic
as the value and operator stores, and no new status code. It is folded into no
checksum and adds no expectation value, because a block record is a parser
register rather than AST.

This also implements one rule CAP-052 froze and never implemented: after a
return statement's `;` the only admissible token is that block's `}`. That the
accepted parser admitted `return 1; return 2;` - parsing both and orphaning the
first return's expression - was confirmed by running the accepted product
against the CAP-052 model before this change, not by reading.

Red-first, in this order. The oracle was extended and the extraction confirmed
behaviour-preserving with compiler.aero byte-identical: the ten CAP-050, the
thirteen CAP-051 and the eighteen CAP-052 probes all green before one new probe
was written. Twenty-five control-flow probes were then hand-derived from the
frozen contract and independently confirmed by the oracle, twenty-five of
twenty-five agreeing on the first run with no probe expectation corrected.
Twenty-four returned 80 from the real linked product beforehand; the
twenty-fifth, `cf-else-without-if`, was already 91 because CAP-052 rejects it
identically, which corrects the contract's claim that all of them must be red.

Canonical function 2, `is_identifier_start`, now parses as a standalone probe
at 21 nodes and one parameter, lifted verbatim and asserted equal to
compiler.aero[146..315] byte for byte. It still does not parse in situ.

The canonical self-ingestion stop is deliberately unmoved and is a regression
guard rather than progress: status 10, offset 146, line 8, column 1, code 0,
actual 3, four nodes, one parameter, at O0 and O2. The accepted 34-byte program
still returns 91 with the identical 144-byte module, and the source remains a
copy-derived successor of accepted B1C asserted byte for byte.

Repository-root gate green: 117 test results, 973 passed, 0 failed, 16 ignored.

Co-Authored-By: Claude Opus 5 <[email protected]>
…lsified

The implementing session's outcome for CAP-053/H1B-4, plus the checkpoint's row
in the readiness table, the H1B-6 bound list, and PROJECT_STATE's current
position. The contract itself is unmodified: it was frozen before any product
change, and both session-outcome sections are kept so the authoring session and
the implementing session read in order.

Two claims are corrected rather than smoothed, and where each was caught is the
point of recording it.

The contract required every control-flow probe to return 80 from the accepted
product before the parser changed. Twenty-four of twenty-five did. The
twenty-fifth, a leading `else` with no preceding `if`, was already 91 - CAP-052
rejects it at the same offset with the same expectation - so its expectation is
genuinely unchanged by this checkpoint and it cannot be red. It is kept as a
lock on a rule that must not move, not cited as evidence of the change.

The oracle extension passed all forty-one inherited probes while having silently
changed what the CAP-052 model predicts for a shape no CAP-052 probe covers. It
surfaced only because the red observation graded an out-of-table shape against
the previous checkpoint's model and the product disagreed. The generalization is
recorded for H1B-5: a probe suite passing is evidence about the probe suite, not
about the extraction, so plan one out-of-table grading deliberately.

Also records the three decisions the contract left open - the expectation code
for an empty nested block, that the block store is folded into no checksum and
adds no expectation value, and that `else` is dispatched inside the statement
loop rather than as a statement - and the empirical confirmation that the
accepted parser admitted both shapes CAP-052's text said it rejected.

Repository-root gate green over these records: 117 test results, 973 passed,
0 failed, 16 ignored.

Co-Authored-By: Claude Opus 5 <[email protected]>
…t represents

H1B-5 is the first checkpoint whose construct is a syntax node
(BOOTSTRAP_CONVERGENCE_READINESS.md:328), so this contract does not reuse any of
the four admit-without-representing arguments CAP-050 through CAP-053 made. Four
node kinds are added and the `1..=19` bound becomes `1..=23`: kind 20 the call,
carrying its callee as payload and its argument list as left; kind 21 one
argument-list cell; kinds 22 and 23 the two references.

Each representational choice is derived rather than asserted. The callee is a
payload because a kind-2 node in callee position would assert the program reads
a variable `f`, which the self-source cannot do. The argument list chains
through nodes rather than a bounded side store because `left = 0` on a call node
would assert the call has no arguments - the same falsehood CAP-053 refused for
an `if` with no body. The chain is built right to left because the arena is
append-only and has no write-at-index path.

Measured independently over the 273,968-byte source, with all four token
censuses reconciling exactly: 1,113 calls, 1,725 arguments, widest list 68,
nesting depth 3, and 451 references of which 451 are a whole call argument over
a bare identifier.

Two corrections to the accepted record. `emitter_fixed_byte` needs 394 nodes,
not the 474 recorded here and in the readiness document; the function is
byte-identical between f416067 and 7b0e929 and 394 falls out of its own token
histogram, so no canonical function lifted verbatim can reach the 512 bound at
H1B-5. And the capacity section's projection 1 for calls is not implementable as
written, so the upper projection's shape for arguments is the real one.

The orphan census is stated honestly rather than claimed as progress:
representing calls moves it from 154 of 13,190 to 240 of 16,819, 98.83% to
98.57%. The orphan problem is statements, and this checkpoint does not touch
statements.

No product file changed. `examples/aero_self_host_v0/compiler.aero` and
`src/compiler/tests/self_host_source_ingestion_tests.rs` are byte-identical to
7b0e929 and were hash-verified so immediately before this commit.
`./tools/test.sh` green: 117 test result lines, 973 passed, 0 failed, 16
ignored, exit 0.

Co-Authored-By: Claude Opus 5 <[email protected]>
…kpoint that represents

CAP-054/H1B-5, implemented from the contract committed unmodified at 7a0fd5d
before any product change. A call is `IDENT ( ARGS )` where the callee is an
operand-position identifier immediately followed by `(`; an argument may begin
with `&` or `& mut` and may do so nowhere else, which is the measured shape -
all 451 references in the canonical source are a whole call argument over a bare
identifier.

Four node kinds take the node-kind bound from `1..=19` to `1..=23`: kind 20 the
call, carrying its callee as payload and its argument list as left; kind 21 one
argument-list cell; kinds 22 and 23 the two references. The callee is a payload
rather than a name-reference child because a kind-2 node in callee position
would say the program reads a variable named `f`, which this source cannot do.
The argument list chains through nodes rather than a bounded side store because
`left = 0` on a call node would say the call has no arguments - the falsehood
CAP-053 refused for an `if` with no body. The chain is built from its last
element, because the arena is append-only and has no write-at-index path.

Open calls are carried by a fifth bounded parse-group arena on the CAP-053 block
store's shape and plumbing, with the same 512 bound, the same status 15 /
diagnostic 512 exhaustion, and no new status code. H1B-6's bound list is now
five and every figure it needs is measured.

Ordering, which is what makes the numbers mean anything. The oracle was extended
first and confirmed behaviour-preserving with compiler.aero byte-identical and
SHA-256-verified before and after: 20/20 green, all ten CAP-050, thirteen
CAP-051, eighteen CAP-052 and twenty-five CAP-053 probes unchanged, before one
new probe was written. Forty-four probes were then hand-derived from the frozen
contract and all forty-four agreed with the oracle on the first run, no
expectation corrected. The red was measured rather than asserted: thirty-six of
forty-three returned 80 from the unedited product and seven returned 91
correctly, because their located rejection is identical under CAP-053; those
seven are kept as locks on rules this checkpoint must not move.

The anti-fitting check CAP-053 asked for was run deliberately. Seven shapes no
probe table covers were frozen with expectations hand-derived under both models,
two of them shapes the models must decide differently, and graded under the
CAP-053 column against the real product before the parser changed. All seven
agreed, so the extraction did not silently move the previous checkpoint's model
where nothing was looking.

Two corrections to the accepted record, both derived rather than asserted.
`emitter_fixed_byte` needs 394 nodes, not 474; the function is byte-identical
between f416067 and 7b0e929 and 394 falls out of its own token histogram, so no
canonical function lifted verbatim can reach the 512 bound at H1B-5. And the
capacity section's projection 1 for calls is not implementable as written.

Calls are represented and the orphan census barely moves: 154 of 13,190 becomes
240 of 17,621, 98.83% to 98.64%. A call's subtree is reachable only when the
call is, and 98.6% of the source's calls sit inside statements that have no
representation. The debt stands where the representation gap put it, and a
second gap is now recorded beside it: no checkpoint owns the ByteBuffer and
Result<int, int> binding types, whose only blocker this checkpoint removed.

The canonical self-ingestion stop is unchanged and that is the correct result:
status 10, offset 146, line 8, column 1, code 0, actual 3, four nodes, one
parameter, at O0 and O2. Three canonical functions parse as verbatim probes -
word_byte_1 at 5 nodes, is_identifier_continue at 15, and main at 144 with the
source's widest argument list at 68. No canonical function containing a
reference can be lifted at all, and that is a measured negative rather than an
omission.

The canonical source is 293,592 bytes, 7-bit ASCII, and remains exactly
reconstructible from accepted B1C byte for byte, now with sixteen further
differences. It also still parses under the grammar it admits: the post-edit
source was run through the measuring instrument and consumes completely, so the
parser this checkpoint wrote is inside the grammar this checkpoint reads.

`./tools/test.sh` green on this exact tree: 117 test result lines, 979 passed, 0
failed, 16 ignored, exit 0. The 979 is six above CAP-053's 973, exactly the six
tests this checkpoint adds. The focused target is 26/26 green in 233 seconds.

Co-Authored-By: Claude Opus 5 <[email protected]>
The CAP-054 handoff said the base commit was "recorded in the commit message
rather than guessed", which is not a base commit. It is `bf4fc97`, and both this
session's commits are now named with what each one was gated against: `7a0fd5d`
carries the contract and was gated with `compiler.aero` and the focused test file
byte-identical to `7b0e929`; `bf4fc97` carries the implementation. Both were
gated on the exact tree committed.

Ledger text only. `examples/aero_self_host_v0/compiler.aero` and
`src/compiler/tests/self_host_source_ingestion_tests.rs` are byte-identical to
`bf4fc97` and were hash-verified so immediately before this commit.
`./tools/test.sh` green on this exact tree: 117 test result lines, 979 passed, 0
failed, 16 ignored, exit 0.

Co-Authored-By: Claude Opus 5 <[email protected]>
… it found

Ledger-first, from locally green CAP-054/H1B-5 at 466701c, confirmed on the
remote by git ls-remote rather than by push output. compiler.aero and the
focused test file are byte-identical to the base; ./tools/test.sh is green on
this exact tree with 117 test result lines, 979 passed, 0 failed, 16 ignored.

The contract raises five parse-group record bounds from 512 to 65,536 and
changes no grammar. Three results were derived before any product edit and are
recorded rather than left to be rediscovered.

The independent oracle models no record bound of any kind. oracle::Bounds
carries source, token and name; neither status 14 nor status 15 occurs anywhere
in it, and no test asserts either. H1B-6's charter sentence is therefore not
satisfiable by editing a literal, and building the model is the checkpoint's
substance. Twenty-six probes passed green around a wholly unmodelled behaviour.

The value bound cannot fire at any uniform bound. Every value push is paired
with a node append and three node appends have no value push, so
value_records <= node_count, and at each value check the paired node increment
has already happened while the node check that guards the same path passed.

The block bound cannot be reached at 65,536. An empty block is rejected, so no
source reaches a block push for under 6.5 tokens, and 65,537 pushes need at
least 425,990 tokens against the frozen 262,144-token bound.

The bound list is five, not three, following the readiness document's own
correction at :352-359 rather than its stale row at :329. The verifier's fourth
512 at :5557 is left alone and recorded as debt, per the standing instruction,
which was checked rather than obeyed on sight.

Co-Authored-By: Claude Opus 5 <[email protected]>
codex and others added 21 commits August 19, 2026 00:14
CAP-055/H1B-6, from the contract frozen and gated at 481f688, which this
commit does not modify. ./tools/test.sh green on this exact tree: 117 test
result lines, 988 passed, 0 failed, 16 ignored. The focused target is 35/35,
nine above CAP-054's 26, which is exactly the nine tests added here.

The literal change is the small half. The independent oracle modelled no record
ceiling of any kind, so the charter's "same independent-oracle proof H1A used
for tokens" required building the model. oracle::Caps carries the five
ceilings, oracle::Counts mirrors the product's never-decremented push counters,
and oracle::Reject carries a located stop whose diagnostic_actual is
independent of the current token, which the product requires because it locates
the reduction stop at a pending operator and the call stop at a held callee.

call_parser_stop was not copied. It is capacity_parser_stop at Caps::UNBOUNDED,
so CAP-054's model survives as an instance rather than a copy that can drift,
and the focused target was 26/26 green with the model in place and the bounds
still at 512 before any capacity test existed.

The model was validated against the old product before the product moved. With
the bound pinned to 512 the four boundary probes were graded against the real
linked product and all four matched hand-derived predictions on the first
attempt, with no correction. It therefore cannot have been fitted to the raised
product. The red is product-confirmed: given the 32,768-leaf chain the base
product appends 512 of 65,535 nodes and stops at status 14, code 512, actual
20, offset 534.

The storage was raised, not only the guard, and this is measured on all five
arenas rather than argued from bytes_new: real product runs append 65,535 node,
origin and value records, 65,536 operator and call records, and 1,300 block
records. The block store is the one ceiling that cannot be reached, so it is
proven from the other side by a probe above the canonical requirement of 1,289,
with the count pinned by bracketing the model rather than taken on its word.

The out-of-table grading against CAP-054's model shows more than a wrong
detail: on the over-bound shapes that model cannot report a capacity stop at
all and returns a grammar stop where the product returns an exhausted arena.

The canonical stop is unchanged at offset 146, line 8, column 1, four nodes,
one parameter, at -O0 and -O2. The canonical source is still exactly
reconstructible from the accepted B1C product byte for byte, with the raise
expressed as one counted transform asserting 16 conditions, 17 compared stores
and 16 diagnostic codes; anchoring each site would have accepted a missed site
silently as no difference.

Three corrections are recorded rather than smoothed. The contract's site count
was half the truth: 33 occurrences over 32 lines. The measurement's "65,536 is
2.5x the upper projection" is 2.4888x, found by a test failing. And the five
ceilings live in a new oracle::Caps rather than in oracle::Bounds, which is a
departure from the contract, because Bounds is ingestion policy.

The verifier's 512 at :5557 is untouched and is now the only one left in the
product, asserted by exact list. Its debt is recorded at the module-shape gate
rather than in a capacity paragraph, with the timing sharpened: it does not
fire at module shape, and it fires at H1C/H1D by a factor of 24.

Co-Authored-By: Claude Opus 5 <[email protected]>
The outcome section's evidence paragraph claimed "the focused target is 34/34
green" and "./tools/test.sh from the repository root, green" before either run
had happened. It recorded an expectation in the past tense. The next
./tools/test.sh invocation returned exit 1, stopping at cargo fmt --check, so
the claim was false when written and became true only after a fix and a second
gate. The figure was then edited to 35/35 when the block-storage probe was
added, again ahead of the run that would confirm it.

The commits were correctly gated: git commit followed a read EXIT=0 in both
cases, so nothing red was ever committed and the rule "green before every
commit" held. What failed is the stricter rule the project runs on, that an
entry means what it says at the time it is written. "It turned out to be true"
is not that standard, and a later session reading a silently corrected entry
could not tell the difference.

The entry now records the defect rather than overwriting it, and carries the
full run table with exit values - including the exit 1 row, labelled as the run
that falsified the claim written above it. Both cited figures now derive from
the final gate rather than from arithmetic. The procedural fix is stated for
the next outcome section: author the evidence paragraph after reading the exit
status, never before.

This commit changes documentation only; examples/ and src/ are byte-identical
to 2426071. ./tools/test.sh was run on this exact tree and its exit status read
before this commit was made: exit 0, 117 test result lines, 988 passed, zero
failed, 16 ignored.

Co-Authored-By: Claude Opus 5 <[email protected]>
The handoff named 2426071 as the next session's base, and 1efc041 overtook it
within the same session, so the paragraph was stale the moment the correction
commit landed. A commit cannot contain its own hash, so the last commit of a
session is always the one its own handoff cannot name; naming a specific base
therefore guarantees staleness whenever a further documentation commit follows.

The handoff now names the three CAP-055 commits by role - 481f688 the contract,
2426071 the implementation, 1efc041 the evidence correction - and defers the
head itself to git ls-remote, which is the only thing that can be right. It
also records that examples/ and src/ are byte-identical between 2426071 and
1efc041, so a reader knows the product and tests are entirely 2426071's.

Documentation only; examples/ and src/ are unchanged from 2426071.
./tools/test.sh was run on this exact tree and its exit status read before this
commit was made: exit 0, 117 test result lines, 988 passed, zero failed, 16
ignored.

Co-Authored-By: Claude Opus 5 <[email protected]>
…indings it derived

Ledger-first, from locally green CAP-055/H1B-6 at 815d162, confirmed on the remote
by git ls-remote rather than by any hash written in a prior handoff. examples/,
src/ and tools/ are byte-identical to the base; this commit changes documentation
only. ./tools/test.sh was run on this exact tree and its exit status read before
this commit was made: exit 0, 117 test result lines, 988 passed, zero failed, 16
ignored.

The gate is split into three checkpoints rather than one, and the split is derived
rather than chosen. The question that fixes it is what refuses a second fn item
once the parser admits one. A parse-group refusal is refuted by compiler.aero:3679,
which requires root == 0 when status != 0 and so would discard root, the item chain
and the root == node_count invariant - the entire result the gate exists to produce.
No refusal at all is refused because :4251 and :4443 have no defined behaviour for a
second function node. What is left is that the downstream phases already refuse:
:4251-4257 asserts semantic_node == root rather than assuming it, and the fact loop
reaches item 1's kind-19 node before it reaches root. So H1M-1 crosses the parse
group alone and predicts four downstream refusals it does not modify; H1M-2 takes
semantic and checked IR; H1M-3 takes the verifier and emitter.

Module shape represents the module, and the alternative was measured rather than
argued. Merely admitting - N function nodes, root the last one - reads 146 reachable
of 17,621, against 240 of 17,621 for a represented item list, so it would take the
orphan census backwards for the first time in the project. The list is a reverse
chain through kind-19 right, because a forward chain is refuted twice by the
product: the node arena has no write-at-index path, and :3657 requires every
reference to point backwards. No new node kind, no new arena, and single-item
behaviour is byte-identical.

Three findings, each derived by a third independent counting instrument that
reproduces eight previously recorded figures exactly, including 17,621/15,842/6,030/
1,289/1,120, emitter_fixed_byte's 394, and CAP-054's census of 240.

The canonical stop moves for the first time since CAP-051, from offset 146 to offset
5,203, line 232, column 15: status 12, code 102, actual 1, on the Result binding
type in read_input_value, with 14 of 23 items parsed. Every function before it
parses completely.

The gate does not exercise the raised bounds. It holds 486 node, 449 value, 169
operator, 54 block and 9 call records - 0.74% of the node arena, and 26 records
inside the 512 bound H1B-6 replaced. The measurement's prediction that 512 would
bite inside function 8 at line 154 was computed under a projected policy the product
does not implement; the real figure at line 154 is 325. The first checkpoint to put
real volume in the arenas is the one that admits the two binding types.

The claim that the five parser checkpoints plus module shape suffice for the
canonical source is false for the product. Both prior instruments modelled a
binding's type as any identifier; the product accepts only int. The non-int binding
type is the only construct in the whole 293,658-byte source that the accepted
grammar plus module shape does not admit - 19 sites in two functions.

The verifier's 512 is not this gate's and is not any H1M checkpoint's. It is owned by
the first checkpoint that drives a complete status == 0 pipeline over a canonical
function larger than 512 nodes, which is gated on the binding-type checkpoint. Its
recorded overrun of "about 24" is corrected to 31.9x for
run_runtime_ascii_llvm_emitter's 16,355 nodes and 34.4x for the whole module.

CAP-055's evidence rule is carried into the method section verbatim: a ledger entry
must be written after reading a completed exit status, never before.

Co-Authored-By: Claude Opus 5 <[email protected]>
… the item list

CAP-056/H1M-1, the parse-group half of the module-shape gate. A function item
now closes at its own '}' and the module then takes another 'fn' item or
end-of-input. The item list is represented rather than merely admitted: a
kind-19 function node's 'right', previously required to be 0, carries the
previous item's node id, so every item is reachable from 'root' and
'root == node_count' is preserved exactly.

Parse group only. No new node kind, no new arena, no new bound, no new checksum
input, and not one line inside the semantic, checked-IR, verifier or emitter
groups. Their refusal of a multi-item module is predicted and asserted, never
edited: semantic_status 27 / semantic_code 3 at the first item's function node
for a module with no identifier, and 17 / 2 at the first identifier use for one
with an identifier.

The canonical self-ingestion stop moves for the first time since CAP-051, and
because the grammar admits more rather than because a bound was relaxed. Fed its
own bytes the compiler parses fourteen complete function items and stops at
status 12, diagnostic_code 102, diagnostic_actual 1, offset 5,203, line 232,
column 15, on 'Result' in a non-int binding type - the exclusion CAP-052 froze.
Every figure was hand-derived before the run and none moved. The five arenas
hold 486 / 449 / 169 / 54 / 9 records, inside the bound CAP-055 replaced, so
this checkpoint does not exercise the raised ones.

No hand-derived node count in any inherited probe table was edited. Each table
still grades its own checkpoint's model and the product is graded against the
module model, with the difference asserted to be either nothing or exactly the
item's own two nodes.

Also corrects, in place and visibly, two records that asserted this checkpoint
green before any run said so - PROJECT_STATE.md at 10:53 and the readiness
document before 10:52 - and restates the evidence rule to bind any record a
later reader could cite, not only the ledger.

Gate: ./tools/test.sh returned exit 0 at 13:23:48 on the exact tree committed,
read before this commit - 117 'test result:' lines, 998 passed, 0 failed,
16 ignored. cargo fmt --check and cargo clippy -D clippy::correctness green.

Co-Authored-By: Claude Opus 5 <[email protected]>
The rule lived only in TASK_LEDGER.md, twice, both times worded to bind "a
ledger entry". CAP-055 and CAP-056 each broke it in a file that wording did not
name - CAP-056 writing "locally green" into PROJECT_STATE.md sixty seconds after
reading a completed red, and into the readiness document while the first run on
the changed tree was still executing. TASK_LEDGER.md stayed clean both times,
which is the only reason either inversion was detectable. A rule followed only
where it is spelled out is a lookup, not a discipline.

Restated in AGENTS.md to bind any record a later reader could cite: the ledger,
PROJECT_STATE.md, a readiness document, a commit message, a handoff, or a status
reported to a human.

Two procedures follow. Write a run's row with its result column empty and fill
it only from a read exit status, because nothing else prevents a summary being
written while the gate is still running. And put the last gate's totals in the
commit message rather than in the files it covers, because recording a gate
edits the tree that gate verified - so exactly one unrecorded run is in flight
at commit time by construction, not by oversight.

Scope: this is narrower than the version first gated. The "read the exit status
rather than a harness's report of it" line and the TMP/TEMP environment
correction were cut as a different subject, not authorized with this change, and
are proposed separately. AGENTS.md still requires all task output on D: while
the recipe it gives does not achieve that, which is an open defect.

Gate: ./tools/test.sh returned exit 0 at 14:51:57 on the exact tree committed,
read before this commit - 117 'test result:' lines, 998 passed, 0 failed,
16 ignored.

Co-Authored-By: Claude Opus 5 <[email protected]>
…raction

Two arguments in the CAP-056 outcome have the same surface shape - "I predicted
it, so proceeding was sound" - and one of them is the failure that cost this
project two retracted records in two days. Left adjacent and unqualified, a
later session could cite the upheld one as precedent for the retracted one.

Records what separates them. The stop-condition-6 prediction is a derivation
from a cost change already made and readable in the product, written down before
the run and confirmed afterward by a completed exit status. The "locally green"
claims were bets on runs still executing, or made against a completed red in the
expectation it would clear. A prediction confirmed by a finished run is
evidence; a prediction standing in for a finished run is the inversion.

Also states in prose what previously had to be read out of a diff: the old
`node-under` probe was retained in both tests that consume it and removed from
neither. Its table entry is byte-identical to the base commit. The model-only
test is untouched and still grades it against CAP-055's model, under which it is
still a grammar stop at 65,535 nodes. Only the product-grading test changed, and
only in which model it derives from; the new status-14 expectation is computed
by the oracle from the two node-ceiling guards and was never typed into a table.
`node-under-with-item` at 65,533 records is an addition, not a substitution.

Records the review outcome: the judgement to continue past the predicted
stop-condition-6 violation was reviewed after commit and upheld.

Gate: ./tools/test.sh returned exit 0 at 15:36:01 on the exact tree committed,
read before this commit - 117 'test result:' lines, 998 passed, 0 failed,
16 ignored.

Co-Authored-By: Claude Opus 5 <[email protected]>
Decision 4 records that the projection "overshoots what this gate actually
produces by a factor of 36 - 17,621 projected against 486 produced". Every
number is right and the inference a reader reaches for is wrong: that the
projection is unreliable, that CAP-055's raise was waste, or that capacity is
solved. The 36x is a prefix-versus-whole artifact - a whole-source projection
divided by an actual measured over 1.74% of the source.

The 14 functions that parse are 61% of the items and 1.74% of the bytes, because
they are the small ones; 97.2% of the module's nodes lie past the stop, 92.8% in
run_runtime_ascii_llvm_emitter alone. The parsed prefix is node-denser than the
module average - 10.61 bytes per node against 16.83 - so 486 over-represents
node production per byte rather than under-representing it. At uniform density
the prefix would hold about 306.

The real discrepancy is about 1.5x, not 36x: on those same fourteen functions
the projected policy crosses 512 inside function 8 while the product holds 339
at the end of function 8. That gap is the node-producing statement policy
CAP-053 declined to implement - a decision, not a measurement error.

So the raise stands on the figure that governs it and that this checkpoint
leaves untouched: 17,621 node records against a bound of 512, a factor of 34.4.
Capacity is untested at scale, not solved. The same qualification is attached to
the arena row in PROJECT_STATE.md and the readiness document, because those are
the records a later session reads: 486/449/169/54/9 fit inside the replaced 512
bound only because the parse stops before the nine expensive functions.

Gate: ./tools/test.sh returned exit 0 on the exact tree committed, read before
this commit - 117 'test result:' lines, 998 passed, 0 failed, 16 ignored.

Co-Authored-By: Claude Opus 5 <[email protected]>
…lly filled it

Operations note OPS-001, not a checkpoint. No product change; examples/ and
src/ are untouched.

The premise that this project's build output filled C: is wrong. Measured
before anything was touched, C: held 460 GB of 461 GB, and this project's
leftovers on it were three zero-byte files - two of them the exact intermediates
from the 11:08 gate failure. There is no cargo target directory anywhere on C:.
The consumers are the OS-managed pagefile at 43.8 GB allocated against 7.4 GB
peak use, hiberfil at 13 GB, and 24 GB of Claude desktop VM bundles. All three
are system or application data, none is a build artifact, and all are left for
Rob. Total freed by this work: 0 bytes, because there was nothing of ours to
free.

The convention failure is real but transient and was not the cause. clang on
Windows does not honour TMPDIR; it reads TMP and TEMP. Proved with clang -###:
with TMPDIR on D: and TMP/TEMP unset it writes to C:\Users\usa50 itself, and
with them inherited from Windows it writes to AppData\Local\Temp. Either TMP or
TEMP alone is sufficient to redirect it; TMPDIR alone does nothing. The root
cause is that nothing enforced the convention - tools/test.sh set no output
location at all and left it to each operator.

tools/test.sh now defaults CARGO_TARGET_DIR and TMP/TEMP/TMPDIR to
repo-relative locations, respects values already exported, converts to a native
path via cygpath, and aborts if any of the four resolves onto C:. Proved by
observation from a hostile environment rather than by reading the script, and
the guard was exercised in both directions.

Checked the three targets that assert content inside the record files -
version_claim, cli_status and cap024_claim_verification - all green, 25 tests.
A full gate was not required for this task.

Co-Authored-By: Claude Opus 5 <[email protected]>
Record corrections only. No product byte moves; examples/ and src/ are
untouched, and the canonical source stays at a839ff37.

The CAP-057 contract (committed in 4608d8a) built a counting instrument and
validated it against six results it did not choose before using any output:
all 70 cells of CAP-056's per-item table, CAP-056's full canonical stop vector,
the standing whole-source requirement at 466701c on all five arenas, the
readiness document's hand-derived 394 for emitter_fixed_byte, the 1,093 product
token count, and the 62-of-486 orphan census. Three standing figures did not
survive it, and one of this session's own hand-derivations did not either.

1. The five-arena requirement 17,621 / 15,842 / 6,030 / 1,289 / 1,120 is not
   wrong, it is stale. The instrument reproduces all five exactly on the
   293,592-byte source at 466701c, which is where it was measured. CAP-056 then
   added 2,926 bytes to compiler.aero - which IS the measured source - so the
   current 296,584-byte tree needs 17,700 / 15,921 / 6,051 / 1,293 / 1,120.
   Any record citing the requirement should name its tree.

2. "Without H1B-6 the binding-type checkpoint would exhaust 512 inside function
   22" is right for three arenas and wrong for the one that fires first, so the
   parse would never reach function 22. Node crosses 512 inside item 16,
   binary_precedence (492 -> 547); value inside item 17 (505 -> 558); operator,
   block and call inside item 22. The node arena governs and fires six
   functions earlier. The conclusion about H1B-6's ordering is unchanged and is
   strengthened.

3. "16 ByteBuffer and 2 Result<int, int> bindings" totals 18 while the same
   document says 19 sites two paragraphs earlier. The source carries 17 and 2.
   The 16 predates CAP-054, whose calls arena added the seventeenth at :521.
   Enumerated: :232, :515-531, :6761.

Also recorded in the contract, because it was this session's own error rather
than an inherited one: the representation gap's "[function 1] yields four
nodes, and all four are orphans" is wrong - exactly one is. The product latches
body_root = expression_root at the return's ';' (compiler.aero:1915), and for a
match return expression_root is the second arm's root. The census derived from
that prose predicted 59 reachable against CAP-056's product-measured 62; the
model was corrected at the mechanism rather than tuned to close the gap, and
then reproduced 62 and 87.24% exactly.

Gate: ./tools/test.sh from the repository root on the exact tree committed,
exit 0 read from the pipeline's own $?, completed 21:23:39 UTC. 117 `test
result:` lines, 998 passed, 0 failed, 16 ignored, cross-checked against the
log's own totals rather than a harness report. Tripwire over compiler.aero,
TASK_LEDGER.md, PROJECT_STATE.md and the readiness document verified unchanged
across the gate.

Two earlier gate attempts died on environment faults and neither is a test
failure. Exit 127: ~/.cargo/env does not exist on this machine, so
tools/test.sh's `. "$HOME/.cargo/env"` guard is a no-op and cargo was never on
the Git Bash PATH - a gap 1e9dee7 does not close, because it hardens TMP/TEMP
and CARGO_TARGET_DIR but not PATH. The gate above ran with cargo and the LLVM
22.1.8 bin directory exported explicitly.

Note on provenance, recorded because a later reader would otherwise be misled:
4608d8a carries this session's 444-line CAP-057 contract as well as the CAP-056
prefix-versus-whole correction its message describes, because a concurrent
session committed the shared working tree mid-authorship. The contract is
intact at TASK_LEDGER.md:490-933. AGENTS.md requires concurrent writing agents
to have non-overlapping files; two sessions were editing TASK_LEDGER.md.

Co-Authored-By: Claude Opus 5 <[email protected]>
… the grammar work

CAP-057/H1M-1b, implemented from bb7f7e4 which `git ls-remote origin
claude/self-hosting-analysis-be3f72` confirms was both the local HEAD and the
remote head. Five files, all authorized by the contract: the parse group of
examples/aero_self_host_v0/compiler.aero, the oracle and its probes, and the
three records.

The canonical source parses end to end, for the first time. status = 0,
root == node_count, 23 items walked from the root through `right` in reverse
order. The canonical parser stop that pinned every checkpoint's evidence from
CAP-051 through CAP-056 no longer exists.

What replaces it was in place and asserted before anything relied on it: the
complete-parse vector (root == node_count is the one assertion a quietly
truncated parse cannot satisfy, because compiler.aero:3680 forces root = 0 on
any stopped parse), the item chain walked rather than counted, and the stop
relocated one phase later to the semantic group's own already-implemented
refusal - semantic_status = 17, semantic_code = 2, node 1, offset 98, line 3,
column 22, the arm-1 body `value`. Predicted in the contract and not modified.

Arenas, hand-derived from the diff before any run, then graded against the model
and then against the linked product at -O0 and -O2. All five exact:

  arena     pre-edit    delta   predicted   observed
  node        17,700     +285      17,985     17,985
  value       15,921     +237      16,158     16,158
  operator     6,051     +114       6,165      6,165
  block        1,293       +9       1,302      1,302
  call         1,120      +32       1,152      1,152

The baseline was verified rather than assumed: this checkpoint's model was run
over the pre-edit bytes first and reproduces the contract's Decision 4
projection exactly on all five arenas, from an instrument that never saw it.

This is the first checkpoint at which any of CAP-055's five raised bounds is
exercised by more than 1%, and the first evidence the raise was necessary rather
than merely ordered correctly: at 512 the parse cannot complete on any of the
five. The node arena holds 27.4% of the raised bound and 35.1x the replaced one.

Three corrections, reported rather than smoothed.

1. The first arena hand-derivation said 289/241/114/9/32 - operator, block and
   call exact, node and value each 4 high. Before changing anything the baseline
   was re-measured (exact) and all eight per-construct unit costs were priced
   individually against the model (all eight exact), which localised it to a
   miscounted unit rather than a mispriced one: the replacement register block
   contains thirteen `let mut stmt_*` lines and the diff adds nine, because
   stmt_b0..stmt_b2 already existed. Fixed at the count, pricing untouched.

2. The probe table's model-separation count said 5 and the instrument said 9.
   The reasoning was incomplete rather than the number wrong - four probes are
   refused by both models at different tokens, because CAP-056 stops at the type
   spelling while CAP-057 walks into Result< , > and stops inside it. The fix
   was not to write 9: the criterion was replaced by one that partitions the
   whole table (5 admitted, 4 refused later, 3 identical). The first attempt at
   that partition discriminated on `status` and got (8,1,3), which is also a
   real error since status cannot separate the groups; the discriminator was
   replaced by whether this checkpoint produced nodes the older model never
   reached.

3. The contract's own byte-proportional estimate does not survive. It projected
   roughly 27-81 nodes for an edit of 1,000-3,000 bytes; this diff is 3,887
   bytes and costs 285 nodes, 13.6 bytes per node against CAP-056's 37. Node
   cost tracks expression structure, not bytes - a ten-way byte comparison is 39
   nodes in one condition. No future checkpoint should size an arena delta from
   a byte count.

Census: 240 reachable of 17,985, 17,745 orphans, 98.665%. The 240 was predicted
exactly - reachability per item is bounded by the last completed return's
expression subtree plus the item's own two nodes, and this diff adds no return
statement, so all 285 nodes it adds are orphans. This figure is comparable to no
earlier one in this ledger and may not be cited as progress or regression
against CAP-056's 87.24%, which measured 1.72% of these bytes.

Out-of-table grading, both halves plus four more than required. Zero churn on
all seven MODEL_LOCK_SHAPES, in every folded field and all four counted arenas.
The product contradicts CAP-056's model on 9 of the 12 binding-type probes. And
on the whole canonical source the product now contradicts CAP-052's, CAP-053's,
CAP-054's and CAP-055's models, each asserted in that checkpoint's own test.

Five inherited tests had their premise expire and none was weakened, skipped or
deleted. The four `..._leaves_the_canonical_stop_unmoved` tests were inverted -
each older model still produces exactly the stop it always produced, asserted,
and the product must now contradict it - and two statement-probe rows
(stmt-bytebuffer-binding, stmt-result-binding) went from one assertion to two:
contradict CAP-052's model and agree with CAP-057's. Their correctly costed
replacements are binding-bytebuffer and binding-result in BINDING_TYPE_PROBES.
every_statement_probe_expectation_is_derived_twice is untouched and still
asserts both lifted rows in full against CAP-052's model.

Decision 2 holds without exception: no node kind, no arena, no bound, no
checksum input, no store. The binding type is checked and discarded exactly as
`mut` is, so the parse cannot distinguish `let x: int = f();` from
`let x: ByteBuffer = f();` in any observable output. Probes G and H - ByteBuffer
in a parameter and in a return position - are both still refused at status 12 /
code 102; only parser_cycle_state == 47 was changed and no shared classifier, so
CAP-050's authority was not crossed. The canonical source remains exactly
reconstructible from accepted B1C byte for byte.

A parse is not a compile, and the_end_to_end_parse_is_not_a_compile asserts it
rather than leaving it to prose: the semantic phase refuses at node 1, zero
facts are appended, and 98.6% of the arena is unreachable. This is grammar
coverage reaching 100% of the canonical source. It is not H1B's :223
obligation, not stage convergence, and not self-hosting.

Gate: ./tools/test.sh from the repository root on the exact tree committed,
GATE_EXIT=0 read from the pipeline's own $?, completed 03:18:56 UTC. 117 `test
result:` lines totalling 1,005 passed, 0 failed, 16 ignored, summed from the
log's own lines rather than from a harness report, with zero FAILED strings.
That is CAP-056's 998 plus exactly the 7 tests this checkpoint adds. A tripwire
of SHA-256 over all 493 tracked files plus HEAD was taken before any edit and
re-verified before and after every run: the tree was byte-identical across the
gate, exactly five files differ from the session base, and HEAD never moved.

Two earlier full gates are recorded in the ledger rather than hidden, because
they covered different trees: run 4 at 01:56:02 UTC on the product and oracle
before the records were written, and run 5 at 02:39:54 UTC on the records tree.
Four test targets check the content of TASK_LEDGER.md, PROJECT_STATE.md and
BOOTSTRAP_CONVERGENCE_READINESS.md, so writing the outcome changed a gated
input. Run 6's timestamp is here and not in the ledger because a gate row naming
its own run cannot be written into the tree that run covered.

Environment note, not a test failure: the first cargo build died with "memory
allocation of 3670016 bytes failed". System commit is nearly exhausted - an
81.8 GB limit with 1.27 GB free - because the pagefile cannot grow with C: at
291 MB free, down from the 1.3 GB OPS-001 measured. CARGO_BUILD_JOBS=2 works
around it and every run above used it. Still not this project's doing; OPS-001's
finding that this project has no build leftovers on C: is unchanged.

Co-Authored-By: Claude Opus 5 <[email protected]>
…e OOM

OPS-002. Tooling and docs only: tools/test.sh and TASK_LEDGER.md. No product
change, no capability claim; examples/ and src/ are untouched.

tools/test.sh now defaults CARGO_BUILD_JOBS and RUST_TEST_THREADS to 2,
respects any value already exported, and exports both.

What it fixes. During CAP-057 a full-parallelism `cargo build` died with
`memory allocation of 3670016 bytes failed`, taking rustc down with internal
compiler errors in crates this project does not own - `cannot find trait
'Default' in this scope` in ryu, `could not resolve trait item being
implemented` in anstyle. Read cold that is a broken toolchain or a poisoned
target directory. It is neither. Measured at the time: 32 GB of RAM with 4.7 GB
free, a system commit limit of 81.8 GB with 1.3 GB available, and C: at 291 MB.
The machine had free physical memory and no free commit, because the OS-managed
pagefile could not grow on a full system drive and the commit limit is RAM plus
pagefile.

Why OPS-001 did not already cover it. 1e9dee7 keeps CARGO_TARGET_DIR, TMP, TEMP
and TMPDIR off the system drive and aborts if any resolves onto C:. The pagefile
is not one of those variables and lives on the system drive regardless, so the
only thing a gate can do about it is generate less pressure. OPS-001 fixed where
the gate writes; this fixes how hard it pushes.

Cost. The full gate goes from roughly 25-40 minutes to roughly 40. That is the
correct trade: a gate that dies after half an hour costs more than a slower one
that finishes, and it costs it twice, because an OOM inside a dependency reads
like a product regression until somebody measures the commit limit. A machine
with headroom raises or removes the cap without editing the script:

    CARGO_BUILD_JOBS=8 RUST_TEST_THREADS=8 ./tools/test.sh

Override semantics verified directly rather than assumed: unset yields 2/2, and
CARGO_BUILD_JOBS=8 RUST_TEST_THREADS=9 yields 8/9.

A correction to OPS-001's premise, from measuring it again. OPS-001's
attribution holds and is reconfirmed - this project still has no target
directory, no incremental directory and no build leftovers anywhere on C:. What
OPS-001 did not establish is that the figure moves on its own. Observed across
2026-08-19 to 2026-08-20 with no deliberate reclamation between the first two
readings: C: free went 1.3 GB -> 291 MB -> 20.9 GB -> 36.5 GB. The pagefile
shrank from 47.2 GB to 35.8 GB once the CAP-057 gates ended, returning about
10.3 GB, and roughly a further 15 GB returned during a window in which this
session ran read-only scans only and can attribute nothing. C: free is not a
reliable standing figure on this machine; the diagnostic that actually predicts
a failed build is the commit limit, via Win32_OperatingSystem's
TotalVirtualMemorySize and FreeVirtualMemory.

Reclamation recorded in the ledger, each figure measured immediately before and
after the removal rather than inferred: 192.97 GB from 28 of 30 stale
per-checkpoint roots under D:\Aero-build-targets, and 283.9 MB from
C:\Users\usa50\.cargo\registry. D:\Aero-build-targets\cap057 is the live root
and was deliberately kept. Nothing else on C: was removed, and the ~20 GB of C:
recovery is explicitly NOT attributed to this session.

Gate: ./tools/test.sh from the repository root on the exact tree committed, run
with CARGO_BUILD_JOBS and RUST_TEST_THREADS unset so the shipped default is the
one exercised. GATE_EXIT=0 read from the pipeline's own $?, completed 15:24:43
UTC. 117 `test result:` lines totalling 1,005 passed, 0 failed, 16 ignored,
summed from the log's own lines rather than from a harness report, with zero
FAILED strings - identical to CAP-057's totals at 68e8341. The run also
re-downloaded 74 crates, which is the whole cost of having deleted the registry
cache.

Co-Authored-By: Claude Opus 5 <[email protected]>
…dence it

CAP-058/H1M-2, authored ledger-first from aaaf6a8 which `git ls-remote origin
claude/self-hosting-analysis-be3f72`, run from the worktree, confirms was both
the local HEAD and the remote head. Three files, all authorized:
TASK_LEDGER.md, PROJECT_STATE.md and BOOTSTRAP_CONVERGENCE_READINESS.md.

No product line is changed. H1M-2 is contracted, gated and NOT implemented,
and every record here says so.

The contract's first job was to say what replaces the canonical stop for a
semantic phase over N items, and the honest answer is partly negative. The
semantic phase runs four passes; pass 3 (compiler.aero:4173-4216) refuses ANY
kind-2 identifier node outright, and the canonical source's node 1 is one - the
arm-1 body `value` at offset 98, line 3, column 22, verified against the bytes
for this contract. The two passes H1M-2 generalizes, symbol emission at :4116
and the fact loop at :4218, sit either side of that refusal. So this is the
first checkpoint in the project whose capability the canonical source cannot
demonstrate at all. Its role here is a negative control: a 23-item,
17,985-node stress input whose located refusal must not move, which is a real
guard because pass 2 runs before pass 3 and pass 2 is one of the two rewrites.

What pins the checkpoint instead is the refusal relocating one authority
further down, to the verifier at :5555, which rejects
verified_function_count != 1 with status 1 / word_index 1 / code 2 /
expected 1 / actual N. Fixed location, value derivable from the source by
counting `fn` keywords, and it fires before the emitter. That is the canonical
stop's own property, moved rather than lost. It is red today by construction:
:4576 gates checked_attempted on semantic_status == 0, so a multi-item module
never reaches the checked group at all.

Three findings against standing records, reported rather than smoothed:

  1. BOOTSTRAP_CONVERGENCE_READINESS.md:504 names three single-function
     assumptions to generalize. There are eight. The two it does not name are
     the checked-IR result-derivation loop at :5229-5257, which assumes result
     i is instruction record i and breaks the moment per-item Return
     instructions interleave, and the instructions == results + 1 invariant at
     :5305. A session generalizing only the three named would have met the
     first in implementation rather than in the contract.
  2. The same row cites the node_count - 2 arithmetic at :4480. It is at
     :4619; :4480 is inside the semantic checksum.
  3. H1M-2 discharges NONE of the representation debt and cannot - the census
     is parse-group authority. Reachable stays exactly 240, over a node count
     that grows by the cost of the checkpoint's own diff, so the ratio gets
     worse. New: the semantic phase is linear over the arena, not a walk of the
     tree, and :4444 requires one fact per node record, so the canonical
     refusal is a refusal OF AN ORPHAN and any future representation checkpoint
     is coupled to the semantic group through that line - making it a
     two-authority checkpoint, which no record held before.

Gate, on the exact tree committed:

  ./tools/test.sh from the repository root returned exit 0, read from the
  gate's own log at 01:57:46 UTC on 2026-08-21. 117 `test result:` lines
  totalling 1,005 passed, 0 failed, 16 ignored, summed from the log rather
  than from a harness report - the harness reported this session's first
  two reds as exit 0, because the shell wrapper exited 0.

  The tripwire over all 493 tracked files is byte-identical to the tree that
  run covered, and HEAD never moved from aaaf6a8.

Runs 1 through 4 are tabulated in the contract with their read exit statuses.
Two of them are worth naming here. Run 1 returned 101 on an environment fault -
LLVM 22.1.8 absent from PATH, which reads as a product regression and is not
one - and the contract now records the fix. Run 3 returned 101 on a timing
flake, llvm_verifier's inherited-pipe deadline test, which reads no repository
markdown and which run 2 passed on a byte-identical src/ tree; run 4 on the
byte-identical tree returned 0, and that re-run is the evidence rather than the
assumption.

One correction to the run-table template itself, inherited from CAP-057 and
caught by this contract's own stop condition 10: the final row originally read
"green; its timestamp is in the commit message", which is a result written
before the run existed. Run 3 then returned 101, so it was false as well as
premature. The row now claims no result at all. This commit message is the only
record written after the tree was fixed, which is why the result is here.

Not claimed: no module of N functions is accepted by anything yet, no
identifier is resolved, no function calls another, and the canonical source is
still refused at its first node. This is a contract.

Co-Authored-By: Claude Opus 5 <[email protected]>
Implemented from 529e931, which `git ls-remote origin
claude/self-hosting-analysis-be3f72`, run from the worktree and querying that
branch by name, confirms was both the local HEAD and the remote head. The
session prompt named 529e931 and warned not to trust it; the warning did not
fire, for the third consecutive time, and it was verified rather than assumed.

THE CHECKPOINT IS INCOMPLETE AND IS RECORDED AS INCOMPLETE. The contract's
Decision 7 stages H1M-2 as 2a then 2b and says in terms that if only 2a lands
the checkpoint is not green. Only 2a landed. The verifier refusal that pins
H1M-2 - verified_actual = N at compiler.aero:5555 - has NOT been observed, and
nothing in these records may be cited as if it had.

What stage 2a changes, Decision 4 only, at three sites in the semantic group:

  S1  one symbol read out of `root` becomes one per item. The item chain is
      walked from `root` for its count and a separate ascending scan appends
      [1, payload(F_i), F_i, 1] per kind-19 node, so `symbol_count !=
      item_count` is a real check rather than a tautology.
  S2  the kind-19 fact rule becomes a chain rule. `semantic_right` must name
      the previous kind-19 node met in the same loop, and the payload
      comparison becomes a read of symbol record i back out of the symbols
      arena - a cross-check between pass 2 and pass 4 over the same item,
      where the accepted rule compared against a single register.
  S3  `1` and `16` become `N` and `16N`, plus the two post-loop assertions.

Pass 1 and pass 3 are untouched, `fact_count == node_count` is not weakened,
and no parse-group, checked-IR, verifier, emitter or driver line moves.

A multi-item module now reaches semantic_status = 0 and is refused one
authority down by C1 - compiler.aero:4583, symbol_count != 1 - which is
predicted and NOT modified, so this stage crosses exactly one authority.

What the probes establish: B, C, D and E reach the checked group and are
refused there; F is refused by pass 4 at item 2's return node on item 2's own
expression type; G stays refused by pass 3 at item 2's identifier, with
symbol_count still 2. What they do NOT establish: nothing about the checked-IR
group beyond its own existing refusal - E's division by zero is never reached
at this stage, so E is currently carrying no more weight than B - nothing about
the verifier, and nothing about a module of N functions being compiled.

The canonical source is the negative control and its located refusal is
UNCHANGED: 17 / 2, node 1, offset 98, line 3, column 22, checked_attempted = 0.
One field moves and was predicted: symbol_count goes 1 -> 23, because pass 2
emits one symbol per item and completes before pass 3 refuses. CAP-056's model
is kept verbatim, still says 1, and the product now rejects its vector.

Census: 240 reachable of 18,650, against 240 of 17,985. The ratio worsens from
98.665% to 98.713% and A WORSENING RATIO HERE IS EXPECTED, not a regression:
all 665 nodes the diff adds are orphans by construction. The five-arena delta
(665, 569, 285, 19, 64) is hand-derived by an independent cost instrument that
reproduces the fourteen per-item rows and the whole pre-edit file exactly
before it was used to price anything.

Four hand-derivation corrections, all fixed at the mechanism and none by
tuning a number:

  1. The cost instrument had three wrong RULES, not three wrong constants: an
     assignment target is free, a `match` scrutinee is free, and a grouping `(`
     or a call pushes an operator record but no node and no value.
  2. A correction to the contract. It predicts CAP-056's semantic model is
     undefined on D, E and F for carrying unseen node kinds. It is undefined
     only on D: the model returns at item 1's function node at id 3, before
     item 2's operators at node 6. The property is "an unseen kind BEFORE item
     1's function node".
  3. A second correction to the contract. Half one of its out-of-table grading
     cannot be executed as written - neither model is defined on the shapes it
     names. What replaces it is stronger and is stated as a replacement.
  4. `item_previous` already exists in run_runtime_ascii_llvm_emitter and the
     Aero subset scopes a function body as one scope, so the new registers are
     named semantic_item_*.

Four inherited tests had their premise expire and were INVERTED, not weakened:
CAP-056's model is kept asserting exactly what it asserted, and the product is
required to contradict it. No test was weakened, skipped or deleted.

The canonical source stays exactly reconstructible from accepted B1C: the
reconstruction gains four anchored transforms and no more.

Gate, both rows written from a read exit status and not before, UTC on
2026-08-21. Run 5, `./tools/test.sh` from the repository root: exit 1 at
04:56:09, on `cargo fmt --check`, before one test ran. Formatting only - the
tree it covered differs from the committed one solely by rustfmt's own reflow of
six statements in the test file, and no product byte and no record byte moved.
Recorded rather than discarded because it is the run that produced the reflow.
Run 6, `./tools/test.sh` from the repository root on THIS EXACT TREE: exit 0 at
05:45:36, with 117 `test result:` lines totalling 1,013 passed, 0 failed, 16
ignored, summed from the log's own lines and cross-checked against the
pipeline's exit status.

Runs 1 through 4 are tabulated in the ledger with their read exit statuses.
Run 1 is the red-first run against the unmodified product at 529e931: exit 101,
with `two-items` and `two-items-second-returns-bool` returning 90, the
semantic-group mismatch code. The tripwire over all 493 tracked files was taken
before anything was read and re-verified before this commit; exactly five files
differ and HEAD never moved from 529e931.

Not claimed: no module of N functions is verified or emitted, no identifier is
resolved, no function calls another, and the canonical source is still refused
at its first node.

Co-Authored-By: Claude Opus 5 <[email protected]>
One file, TASK_LEDGER.md, authorized by the contract. No product line and no
test line moves; `compiler.aero` is byte-identical to c1076e7.

This is a deliberate stop at the boundary Decision 7 exists to provide, not an
exhausted one. H1M-2 stays INCOMPLETE and stage 2b is NOT STARTED.

The handoff names the base and how to confirm it, what stage 2a proved, and
what it ruled out - including the negative that matters most, that probe E is
currently inert. E exists to prove the checked-IR group evaluates item 2's
expressions, and C1 refuses before the expression loop runs, so at stage 2a E
is indistinguishable from B. Making E mean something is stage 2b's first job.

It also names the thing that decides how long stage 2b takes, which is not the
product edit: checked_checksum folds every word of checked_ir and
verified_checksum folds them again, so the oracle has to construct the module
word for word rather than assert counts. The layout to transcribe, the order to
do the work in, and five operational facts this session paid for - including
that `cargo fmt --check` runs first in the gate and exits before one test, and
that the Aero subset scopes a function body as one scope - are recorded so the
next session does not rediscover them.

Gate, written from a read exit status and not before. Run 7,
`./tools/test.sh` from the repository root on THIS EXACT TREE: exit 0 at
06:36:05 UTC on 2026-08-21, with 117 `test result:` lines totalling 1,013
passed, 0 failed, 16 ignored, summed from the log's own lines and cross-checked
against the pipeline's exit status. The tripwire over all 493 tracked files was
re-verified before this commit: exactly one file differs from c1076e7 and HEAD
never moved.

Co-Authored-By: Claude Opus 5 <[email protected]>
Stage 2b generalizes the checked-IR group to N function items at C1 through
C8, across sixteen anchored sites inside compiler.aero:4552-5320 and no
others. The refusal relocates from the checked group's own symbol_count != 1
to the verifier at :5555, which refuses verified_function_count != 1 with
verified_status 1, word_index 1, code 2, expected 1 and actual = N.

That vector was written down before the product was edited and was not
adjusted afterwards. Predicted and observed agree field for field, on B and C
at N of 2 and 3, and `actual` is re-derived inside the test from the probe's
own raw bytes by counting `fn ` occurrences rather than read from the probe
row. Stop condition 8 did not fire and is asserted rather than assumed.

Probe E was inert at stage 2a - C1 refused before the expression loop ran, so
B and E were refused identically with checked_value_count 0 on both. It now
discriminates: B completes a 63-word module and reaches the verifier, E is
refused at checked_status 2 / code 6 at node 6, offset 52, line 1, column 53,
inside item 2's own bytes at a division only item 2 contains. Both are graded
against the linked product, not compared as models.

Stage 2a's C1 vector is kept verbatim, still asserted, and now contradicted by
the product with code 92 - the checked-group comparison and nothing else.

Corrects a counting error the contract, this ledger and the readiness document
all carried: there are eleven single-function assumptions, three semantic and
eight checked-IR, not eight. "Eight" was the size of the checked-IR table alone
promoted to a total of both groups; stage 2a then computed 8 - 3 = 5 and
reported five open sites where its own table showed eight. Corrected in place.

Two hand-derivation corrections, both fixed at the mechanism: the placeholder
guard was written onto the left child lookup and not the right, which is why
every N >= 2 probe failed while the single-item path stayed byte-identical;
and an anchor stopped one line short of its block's closing brace.

Negative control unmoved: 17/2 at node 1, offset 98, line 3, column 22,
checked_attempted 0, symbol_count 23. The census is 240 of 18,718 - 98.718%
orphaned against 98.713% - which is expected by construction and may not be
cited as progress or as decay. Stage 2b's arena delta is measured by the
instrument rather than hand-derived per construct, and is recorded as weaker.

Gate: ./tools/test.sh from the repository root on this exact tree returned
exit 0, 117 test result lines totalling 1,016 passed, 0 failed, 16 ignored,
read from the log's own EXIT line at 16:15:00 UTC on 2026-08-21. Clippy is
clean of correctness lints and this change adds no warning of any category.
The tripwire over all 493 tracked files was taken before anything was read and
re-verified before this commit; exactly the five files this session edited
differ and HEAD never moved.

Co-Authored-By: Claude Opus 5 <[email protected]>
…oduces

Ledger-first, no product line edited. `examples/` and `src/` are untouched;
the whole change is `TASK_LEDGER.md` and `BOOTSTRAP_CONVERGENCE_READINESS.md`.

Authored from 7d810d7, confirmed the local and remote head by
`git ls-remote origin claude/self-hosting-analysis-be3f72` run from the
worktree. Five consecutive non-firings of the stale-id warning now, still a
reason to verify rather than a prediction.

Decision 1 answers the session prompt's first question. The located-refusal
chain does end here, and that is established rather than assumed: below the
emitter the B1C driver makes only internal checksum authentication and refuses
no module shape, so there is no next authority. Two distinctions the premise
compresses: the canonical source's stop died at CAP-057 and the probe's refusal
does not die here - E, F and G stay refused and must be unmoved - so what ends
is the refusal as the *positive* instrument. Three replacements graded; the
emitted LLVM bytes are load-bearing, and they are recorded as *stronger* than
what they replace rather than as a loss, because a refusal cannot distinguish a
compiler that declined correctly from one that cannot do the work.

Decision 3 transcribes both groups and finds thirty-two single-function
assumptions where the readiness document names one. Three are not count
generalizations. The worst is the emitter's register naming, which names a
definition by instruction id and a reference by result id: those coincide only
while one Return exists and is last, and on a two-item module the emitter would
define %r3 and reference %r2. The product would not refuse that - it exits 91
and prints LLVM referencing an undefined register, invisible to every exit code
this product has. That is why Decision 1 requires a second grader the checkpoint
did not author.

Decision 4 raises the verifier's 512 to 65,536 under the verifier group's own
derivation rather than by analogy: the value is a parse-group node id, the parse
group refuses an append at `node_count >= 65536` and issues `node_id =
node_count` immediately afterwards, so 65,536 is the largest id a well-formed
producer can emit. Falsifier stated - had the guard been `> 65536` the bound
would be 65,537. `verified_header_instructions <= 510` is deliberately not
raised.

Decision 5 evidences `entry_function = N` rather than replacing it. Both
replacements are refuted: "named main" refuses the frozen accepted module
`fn score`, and "item 1" contradicts a rule already in the product. The verifier
derives word 5 from the function records instead of trusting it, a fault at
word 5 falsifies the rule, and the emitter makes the entry observable in the
bytes.

Decision 10 declines to stage verifier then emitter, unlike CAP-058: a
verifier-only intermediate reaches the unmodified emitter and produces LLVM with
two terminators in one basic block, which is a wrong product rather than a
refusing one. Only the bound raise stages separately.

Two inherited line citations are corrected at the source rather than restated:
the verifier's 512 is at compiler.aero:6018, not :5557 as
BOOTSTRAP_CONVERGENCE_READINESS.md said in three places, and
`verified_function_count != 1` is at :5878, not :5555.

Census: 240 of 18,718, 98.718% orphaned. H1M-3 discharges none of the
representation debt and the figure may not be cited as progress or as decay.
The 240 has held across three checkpoints partly by constraint, and that is
recorded so it is not read as a measurement.

Gate on this exact tree: `./tools/test.sh` exit 0, 117 targets, 1,016 passed,
0 failed, 16 ignored. Started 17:39:42 UTC and the exit status was read at
18:27:47 UTC on 2026-08-21, 48m05s wall clock under OPS-002's two-job cap. The
result is stated here rather than in the contract's own gate row, because a row
naming its own run cannot be written into the tree that run covered.

The tripwire over all 493 tracked files was taken before anything was read and
re-verified before this commit. No file changed that this session did not
change.

Co-Authored-By: Claude Opus 5 <[email protected]>
Integrate 2d99ca7 without changing its compiler production or tests. Preserve original red-first contracts and phase-2 exclusions. Documentation replay completed: 31 passed, exit 0 before final historical-status headers. Full exact-candidate local and public acceptance gates remain pending; this commit makes no acceptance or self-hosting claim.
INTEGRATION-001-W1: reproduce scorer red (0 matches instead of 1) and inference payload red before adapting allocation placement. Preserve dependency/guard/native assertions and require entry-block storage on both platforms. New workflow-regex mutation tests: 2 passed, exit 0. Exact Windows fixed-array workflow replay: exit 0 at O0/O2. Existing focused stack/self-source tests: 3 + 63 passed. Full unchanged-tree root gate and exact-candidate public checks remain pending. No compiler production or phase-2 feature change.
Completed exact-candidate red in local root gate and CI 34008617659: 19/20 fixed-array tests, obsolete inline-allocation companion fragments only. Replace those two fragments with explicit entry-block prefixes plus unchanged adjacent dataflow requirements. Final focused replay: 53 passed across six targets, exit 0; full root gate replay remains pending. No compiler or phase-2 implementation changes.
@RobVanProd
RobVanProd marked this pull request as ready for review September 6, 2026 04:10
@RobVanProd
RobVanProd merged commit df0fece into master Sep 6, 2026
13 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants