Skip to content

fix(runtime): record the failure Claude Code states instead of only "Claude run failed" - #1224

Merged
aviggiano merged 4 commits into
mainfrom
fix/1084-claude-stated-failure
Oct 1, 2026
Merged

aviggiano merged 4 commits into
mainfrom
fix/1084-claude-stated-failure

Conversation

@mrthankyou

@mrthankyou mrthankyou commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Refs #1084 (the remaining "opaque error" half; the retry half landed in #1171)

Problem

A failed Claude attempt is recorded only as Claude run failed. Smithers' ClaudeCodeAgent maps a failed result that has no error field to that generic string and drops the result text, which is where Claude Code states the cause, for example:

Failed to refresh OAuth token: another Claude Code process is refreshing it or exited mid-refresh. …

In #1084 that text survived only in the session transcript under CLAUDE_CONFIG_DIR.

Why not put the text in the error message

This follows the recommendation in the #1084 status update rather than cherry-picking 2c7fa0d1. Smithers classifies the thrown message:

  • BaseCliAgent runs its quota, config-invalid and session-loss classifiers on it;
  • the engine's auth regex (/invalid_authentication|401|…/, engine.js:8748) disables the agent for the run;
  • the scheduler treats ENOENT as terminal (failureClassification.js).

An expired-OAuth result in the message would therefore end the node on its first attempt. So the stated cause travels beside the message, in a field nothing classifies.

Change

  1. agents/claude.tsx: when a result line produces Smithers' generic Claude run failed, the adapter keeps Claude Code's result text and attaches it to the thrown error as details.agentStatedFailure. The message and code are unchanged.
  2. workflows/workflow.tsx: the agent-failure normalizer passes agentStatedFailure through. It gets the same secret redaction and 1,000-byte cap as failure_message. It is never a control input: the 402 promotion still reads only the message, so a quota- or 402-worded statement does not park the run.
  3. workflow-sync.ts: errorText() appends the statement. A node's last_error (shown by status and why) and the attempt ledger's failure_message read Claude run failed: <stated cause>.
  4. The boundary gate and docs/reference/agent-adapter-boundaries.md declare claude.tsx's new output-interpretation responsibility, linked to Concurrent agents race on OAuth token refresh; three immediate retries all re-race and kill the run with an opaque "Claude run failed" #1084. There is also a CHANGELOG entry. Existing projects pick up the adapter by re-running ultrafuzz init.

Not changed: retry, park, auth-disable and stall behaviour; no new flags, config or failure categories.

Known cosmetic effect: the existing secret redactor reads OAuth token: another as an assignment, so the recorded #1084 text shows OAuth token: <redacted> Claude Code process is refreshing it…. The cause is still recognisable, and the redactor is untouched.

Verification

  • New Bun adapter test: a fake claude CLI prints an is_error result and exits 1. The thrown message stays Claude run failed…, and details.agentStatedFailure carries the text for the OAuth-race case and for an expired-token 401 (the message never contains 401). A result with no text attaches nothing.
  • New normalizer test: the statement is redacted, and a 402/quota-worded statement doesn't promote the code. Empty or non-string values are dropped.
  • New sync test: a NodeFailed whose error carries the statement shows it in last_error and in the ledger's failure_message.
  • The details survive the engine: Smithers' errorToJson copies an error's own keys. The allowlisted details of normalized errors already reach NodeFailed payloads in real runs, e.g. the failureQuota details in a local metamorpho run's stream.ndjson.
  • Existing tests pass: the adapter-boundary gate, the Claude/DeepSeek adapter contracts, the init template test and the failed-node sync test. Lint, strict lint and Prettier are clean.
  • Local smoke run: smoke profile against Damn Vulnerable DeFi src/backdoor with ClaudeAgent/claude-opus-4-8 (subscription). 24/24 succeeded and the report was verified. No attempt failed, so this confirms no regression but doesn't exercise the new path.
  • Full local suite: the supporting suite's 46 local failures all come from the host environment. An ambient Codex config routes to a local provider, and the rest are macOS /var, ENAMETOOLONG, EILSEQ, forge ulimit and a git master default. The governance, lifecycle and audit-profile files pass with the ambient Codex config isolated. I haven't run the full runtime.test.ts locally.

Follow-up

DeepSeek runs Claude Code through its own adapter (deepseek.tsx) and has the same gap. That's in a separate stacked PR.

🤖 Generated with Claude Code

RetriggerConfidence Score: 5/5

This PR appears safe to merge; no new actionable issue was identified in its changes.

Summary

The PR preserves Claude Code’s stated cause alongside a generic thrown failure, then redacts and includes that cause in durable node and attempt records without changing the message used for scheduling decisions.

  • Adds adapter, normalization, and synchronization coverage for the failure path.
  • Updates the adapter-boundary declaration and documentation.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Claude failed result] --> B[Generic thrown error]
  A --> C[details.agentStatedFailure]
  B --> D[Existing classifiers]
  C --> E[Redact and cap]
  E --> F[Node last_error and attempt failure_message]
Loading

Reviews (4) · Last reviewed commit: "Merge remote-tracking branch 'origin/mai..."

…Claude run failed"

Smithers' ClaudeCodeAgent reports a failed result that has no `error`
field as "Claude run failed" and drops the `result` text, which is where
Claude Code states the cause, e.g. a contended OAuth refresh (#1084). The
cause survived only in the session transcript.

Smithers classifies the thrown message (quota park, auth disable, ENOENT,
session loss), so the text must not reach it. Instead:

- The Claude adapter keeps the result text of a generically failed
  result and attaches it to the thrown error as
  `details.agentStatedFailure`, leaving message and code unchanged.
- The workflow's agent-failure normalizer passes that one field through,
  redacted and capped like `failure_message`. It is never a control
  input: a quota- or 402-worded statement does not park the run.
- Run synchronization appends it to the node's `last_error` and the
  attempt ledger's `failure_message`, so `status` and `why` show
  "Claude run failed: <stated cause>".

The claude.tsx adapter now declares the output-interpretation
responsibility (#1084) in the boundary gate and its reference doc.

Refs #1084

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…d-failure

# Conflicts:
#	CHANGELOG.md
#	packages/runtime/test/generated-workflow-verifier.test.ts
@greptile-apps

This comment has been minimized.

…in the stated-failure sync test

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@mrthankyou

Copy link
Copy Markdown
Collaborator Author

Priority note for maintainers: this is not a stability fix, so merging it can wait.

This PR makes failures easier to diagnose; it doesn't change whether runs fail. It records the cause Claude Code states next to the generic Claude run failed, but retry, park, auth-disable and stall handling are unchanged, so a run that failed before still fails the same way.

The stability half of #1084 (immediate retries that re-raced the OAuth refresh until stall detection killed the run) was already fixed by #1171, merged 2026-09-29.

Please prioritise PRs that change run outcomes ahead of this one. The same applies to #1225, which is stacked on this PR and does the same for DeepSeekAgent.

Both branches are up to date with main. Their only CI failure is the @grpc/grpc-js advisory gate, which #1239 fixes.

@aviggiano

Copy link
Copy Markdown
Collaborator

Thanks @mrthankyou, this is a careful fix. Keeping Smithers' thrown message and code unchanged, and carrying Claude Code's result text beside them as details.agentStatedFailure, is the right design for #1084.

I traced every classifier in the pinned Smithers 0.35, and none of them reads the new field:

  • BaseCliAgent's quota, config-invalid and session-loss checks
  • the engine's auth regex, which reads message + responseText
  • isQuotaErrorPayload
  • computeErrorSignature
  • the 402 promotion

So retry, fallback and quota-park behaviour is unchanged, including the 60 s backoff from #1171. Redaction and the 1,000-byte cap apply in the normalizer and again when the text is written. CI is green. Merging.

We'll open a small follow-up PR before v0.1.3 that:

  • Corrects the CHANGELOG. status and why don't show the statement, because both relay Smithers' message-only summaries. It appears in ultrafuzz inspect <run-id> --json (last_error), the dashboard, attempts.jsonl failure_message and the public eval diagnostics. Existing projects must re-run ultrafuzz init, because run refuses a stale claude.ts.
  • Changes the join in errorText() to (agent stated: …). SmithersError appends See https://smithers.sh/reference/errors to the message, so today the text reads …/reference/errors: Failed to refresh OAuth token: …. The sync test will use that real message.
  • Extends the Bun adapter test with a reused agent instance, which covers the per-generation reset, and with a result that has an error field, which covers the generic-message guard.
  • Caps the result text kept in claudeResultText before redaction.
  • Drops the unused execution.agentCredentialEnv from the normalizer test.
  • Tidies the adapter-boundary row: the adapter also overrides generate. We'll add an upstream Smithers issue link if we file one.

Showing the statement in why itself stays tracked under #1084. No blockers. Thanks again!

@aviggiano
aviggiano merged commit 8a92ac0 into main Oct 1, 2026
17 checks passed
@aviggiano
aviggiano deleted the fix/1084-claude-stated-failure branch October 1, 2026 07:28
aviggiano pushed a commit that referenced this pull request Oct 1, 2026
…ike ClaudeAgent (#1225)

`DeepSeekAgent` runs Claude Code against DeepSeek's Anthropic-compatible endpoint through its own `ClaudeCodeAgent` subclass (`DeepSeekClaudeCodeAgent`). So it has the same gap #1224 fixes for `ClaudeAgent`: a failed result without an `error` field is reported only as `Claude run failed`, and Claude Code's `result` text, which states the cause (e.g. an auth error from the DeepSeek route), is dropped.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants