Skip to content

The durable journal can replay but nothing can resume: crash recovery stops one API call short of the payoff #21

Description

@minerva-sky

Hypothesis

Wiring ExecutionJournal.replay into PlanOrchestrator#execute_plan (skip completed tasks, seed their recorded outputs) would turn the journal from an audit artifact into actual crash recovery, at a cost of roughly one small method and a CLI flag. For long plans that die at task 9 of 10 — process kill, rate-limit exhaustion, Ctrl+C — re-running currently re-executes (and re-bills) everything.

Evidence

  • The replay half is built and hardened: ExecutionJournal.replay(path:, mode:) reconstructs completed_task_ids, outputs, failures, with tolerant/strict modes (lib/agentic/execution_journal.rb:171). ReplayedState#completed?(key) exists (:66) and matches by id or description — which only makes sense as a resume primitive, since a fresh process constructs tasks with new ids.
  • Nothing consumes it: zero callers of completed?/replay anywhere in lib/ outside the journal itself (grep). The orchestrator has no way to accept a prior state; agentic execute has no --resume/--journal option (lib/agentic/cli.rb:95-134).
  • The CLI already interrupts gracefully on Ctrl+C (execute_tasks rescues Interrupt, exits 130) — a clean half-finished journal is a state the product already produces on purpose, then can do nothing with.

Why this is Direction

WORLD.md: "Depth here wins: … better failure recovery, better execution history." The journal (Stack 5, #5) built the history; resume is the recovery that history was for. It also compounds with #10 — once dependents-of-failed-tasks get a terminal state, "fix the input and resume the rest" becomes the natural workflow.

Experiment that would validate it

  1. execute_plan(agent_provider, resume_from: replayed_state) — before scheduling, mark tasks whose id/description is in completed_task_ids as completed and seed results/outputs so needs: wiring and dependents see real values.
  2. CLI: agentic execute --plan plan.json --resume journal.jsonl.
  3. Test that proves it: run a 5-task plan with an agent that raises on task 4, kill, resume — assert tasks 1-3 execute zero times on the second run and the plan completes. Plus a tolerant-mode resume over a torn journal (the damaged-line fixtures already exist).

Measurable claim: on an N-task plan failing at task k, resume cost drops from N tasks to N−k+1. With LLM-backed tasks that's directly dollars and minutes.

Cost

  • One orchestrator entry-point change (kwarg, default nil — no behavior change for existing callers), one CLI option, specs. No new dependencies.
  • Open design question worth deciding at review: description-matching resume across edited plans (plan changed between runs) — I'd scope v1 to unchanged plans and let completed?'s description fallback be the escape hatch, not the contract.

Class code-feature → L0 per autonomy policy, so proposing here rather than as a PR, though this one is small and reversible enough that I'd take it to a PR on a nod.


Origin: loop:self session 2026-08-21. Nearest neighbors: #10 (terminal states for stranded dependents — complementary, not overlapping) and PR #5 which built the journal. PR #7 touches observability but not journal resume (verified against its file list).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    loop:selfForward-thinking improvement loop (WORLD.md-aligned)status:analyzedAnalyzed, awaiting decision

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions