fix(runtime): record the failure Claude Code states instead of only "Claude run failed" - #1224
Conversation
…Claude run failed" Smithers' ClaudeCodeAgent reports a failed result that has no `error` field as "Claude run failed" and drops the `result` text, which is where Claude Code states the cause, e.g. a contended OAuth refresh (#1084). The cause survived only in the session transcript. Smithers classifies the thrown message (quota park, auth disable, ENOENT, session loss), so the text must not reach it. Instead: - The Claude adapter keeps the result text of a generically failed result and attaches it to the thrown error as `details.agentStatedFailure`, leaving message and code unchanged. - The workflow's agent-failure normalizer passes that one field through, redacted and capped like `failure_message`. It is never a control input: a quota- or 402-worded statement does not park the run. - Run synchronization appends it to the node's `last_error` and the attempt ledger's `failure_message`, so `status` and `why` show "Claude run failed: <stated cause>". The claude.tsx adapter now declares the output-interpretation responsibility (#1084) in the boundary gate and its reference doc. Refs #1084 Co-Authored-By: Claude Opus 5.5 <[email protected]>
…d-failure # Conflicts: # CHANGELOG.md # packages/runtime/test/generated-workflow-verifier.test.ts
This comment has been minimized.
This comment has been minimized.
…in the stated-failure sync test Co-Authored-By: Claude Opus 5.5 <[email protected]>
|
Priority note for maintainers: this is not a stability fix, so merging it can wait. This PR makes failures easier to diagnose; it doesn't change whether runs fail. It records the cause Claude Code states next to the generic The stability half of #1084 (immediate retries that re-raced the OAuth refresh until stall detection killed the run) was already fixed by #1171, merged 2026-09-29. Please prioritise PRs that change run outcomes ahead of this one. The same applies to #1225, which is stacked on this PR and does the same for Both branches are up to date with |
…d-failure # Conflicts: # CHANGELOG.md
|
Thanks @mrthankyou, this is a careful fix. Keeping Smithers' thrown message and code unchanged, and carrying Claude Code's I traced every classifier in the pinned Smithers 0.35, and none of them reads the new field:
So retry, fallback and quota-park behaviour is unchanged, including the 60 s backoff from #1171. Redaction and the 1,000-byte cap apply in the normalizer and again when the text is written. CI is green. Merging. We'll open a small follow-up PR before v0.1.3 that:
Showing the statement in |
…ike ClaudeAgent (#1225) `DeepSeekAgent` runs Claude Code against DeepSeek's Anthropic-compatible endpoint through its own `ClaudeCodeAgent` subclass (`DeepSeekClaudeCodeAgent`). So it has the same gap #1224 fixes for `ClaudeAgent`: a failed result without an `error` field is reported only as `Claude run failed`, and Claude Code's `result` text, which states the cause (e.g. an auth error from the DeepSeek route), is dropped.
Refs #1084 (the remaining "opaque error" half; the retry half landed in #1171)
Problem
A failed Claude attempt is recorded only as
Claude run failed. Smithers'ClaudeCodeAgentmaps a failed result that has noerrorfield to that generic string and drops theresulttext, which is where Claude Code states the cause, for example:In #1084 that text survived only in the session transcript under
CLAUDE_CONFIG_DIR.Why not put the text in the error message
This follows the recommendation in the #1084 status update rather than cherry-picking
2c7fa0d1. Smithers classifies the thrown message:BaseCliAgentruns its quota, config-invalid and session-loss classifiers on it;/invalid_authentication|401|…/,engine.js:8748) disables the agent for the run;ENOENTas terminal (failureClassification.js).An expired-OAuth
resultin the message would therefore end the node on its first attempt. So the stated cause travels beside the message, in a field nothing classifies.Change
agents/claude.tsx: when aresultline produces Smithers' genericClaude run failed, the adapter keeps Claude Code'sresulttext and attaches it to the thrown error asdetails.agentStatedFailure. The message and code are unchanged.workflows/workflow.tsx: the agent-failure normalizer passesagentStatedFailurethrough. It gets the same secret redaction and 1,000-byte cap asfailure_message. It is never a control input: the 402 promotion still reads only the message, so a quota- or 402-worded statement does not park the run.workflow-sync.ts:errorText()appends the statement. A node'slast_error(shown bystatusandwhy) and the attempt ledger'sfailure_messagereadClaude run failed: <stated cause>.docs/reference/agent-adapter-boundaries.mddeclareclaude.tsx's newoutput-interpretationresponsibility, linked to Concurrent agents race on OAuth token refresh; three immediate retries all re-race and kill the run with an opaque "Claude run failed" #1084. There is also a CHANGELOG entry. Existing projects pick up the adapter by re-runningultrafuzz init.Not changed: retry, park, auth-disable and stall behaviour; no new flags, config or failure categories.
Known cosmetic effect: the existing secret redactor reads
OAuth token: anotheras an assignment, so the recorded #1084 text showsOAuth token: <redacted> Claude Code process is refreshing it…. The cause is still recognisable, and the redactor is untouched.Verification
claudeCLI prints anis_errorresult and exits 1. The thrown message staysClaude run failed…, anddetails.agentStatedFailurecarries the text for the OAuth-race case and for an expired-token 401 (the message never contains401). A result with no text attaches nothing.NodeFailedwhose error carries the statement shows it inlast_errorand in the ledger'sfailure_message.errorToJsoncopies an error's own keys. The allowlisteddetailsof normalized errors already reachNodeFailedpayloads in real runs, e.g. thefailureQuotadetails in a local metamorpho run'sstream.ndjson.inittemplate test and the failed-node sync test. Lint, strict lint and Prettier are clean.smokeprofile against Damn Vulnerable DeFisrc/backdoorwithClaudeAgent/claude-opus-4-8(subscription). 24/24 succeeded and the report was verified. No attempt failed, so this confirms no regression but doesn't exercise the new path./var,ENAMETOOLONG,EILSEQ, forgeulimitand a gitmasterdefault. The governance, lifecycle and audit-profile files pass with the ambient Codex config isolated. I haven't run the fullruntime.test.tslocally.Follow-up
DeepSeek runs Claude Code through its own adapter (
deepseek.tsx) and has the same gap. That's in a separate stacked PR.🤖 Generated with Claude Code
This PR appears safe to merge; no new actionable issue was identified in its changes.
Summary
The PR preserves Claude Code’s stated cause alongside a generic thrown failure, then redacts and includes that cause in durable node and attempt records without changing the message used for scheduling decisions.
Diagram
%%{init: {'theme': 'neutral'}}%% flowchart LR A[Claude failed result] --> B[Generic thrown error] A --> C[details.agentStatedFailure] B --> D[Existing classifiers] C --> E[Redact and cap] E --> F[Node last_error and attempt failure_message]Reviews (4) · Last reviewed commit: "Merge remote-tracking branch 'origin/mai..."