fix(runtime): agent adapters never hang or reroute on telemetry and environment quirks - #1173
Merged
Merged
Conversation
agent-adapter-boundaries.test.ts pinned every adapter template's SHA-256 twice (a structural policy and a separate responsibility policy) plus line and syntax-node ceilings. Those assertions check no behaviour: a comment edit or the one-line missing `import path` fix in deepseek.tsx failed the gate with "changed from its reviewed source fingerprint", so every adapter fix needed two hash edits. Merge the two tables into one policy per source (purpose, declared responsibilities, upstream links) and delete the fingerprints, the ceilings, and the tests that only exercised them. The real boundary checks stay: every source needs a policy, only registered adapters may own responsibilities, non-adapter helpers may carry no orchestration signals, and a statically detected responsibility still fails until it is declared. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…for undefined names deepseek.tsx called path.join without importing path, so a render without ULTRAFUZZ_CONFIG_PATH threw "path is not defined" from the factory. Nothing caught it: runtime templates are copied into projects and run by Bun, tsconfig only includes src/**/*.ts, and ESLint turns no-undef off for every .ts/.tsx file. Enable no-undef for packages/runtime/src/templates/**/*.tsx and declare the __ULTRAFUZZ_*__ placeholders the compiler substitutes as readonly globals. Against main the rule reports exactly this bug (deepseek.tsx 200:59 'path' is not defined) and nothing else. Co-Authored-By: Claude Opus 5.5 <[email protected]>
Comment on lines
840
to
+843
| const sources = readSourceTree(path.join(packageRoot, "src/templates/smithers/agents")); | ||
|
|
||
| assert.deepEqual( | ||
| Object.keys(sourcePolicies).sort(), | ||
| Object.keys(adapterPolicies).sort(), |
There was a problem hiding this comment.
Opaque adapter changes go unchecked
Removing the source fingerprints leaves this gate checking only responsibilities its static detector recognizes. The deleted test showed that an adapter could change output handling through a helper without producing such a signal. An equivalent future edit can now pass without an explicit responsibility review. Keeping a review tripwire for changes the detector cannot recognize would preserve that coverage without requiring two fingerprints.
Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/runtime/test/agent-adapter-boundaries.test.ts
Line: 840-843
Comment:
**Opaque adapter changes go unchecked**
Removing the source fingerprints leaves this gate checking only responsibilities its static detector recognizes. The deleted test showed that an adapter could change output handling through a helper without producing such a signal. An equivalent future edit can now pass without an explicit responsibility review. Keeping a review tripwire for changes the detector cannot recognize would preserve that coverage without requiring two fingerprints.
---
For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!
DeepSeekAgent overrode the Claude Code output interpreter with a strict usage parser that required DeepSeek's OpenAI-compatible field names (prompt_cache_miss_tokens/prompt_cache_hit_tokens) and rejected input_tokens and cache_read_input_tokens as "legacy aliases". Claude Code always reports Anthropic field names, so every finished DeepSeek task threw from onStdoutLine. That hook runs inside the child process 'data' listener, so the throw escaped the listener instead of failing the invocation, and the task sat until the idle or node timeout. Delete the usage overrides (createOutputInterpreter, generate, stream and their parsing/attachment helpers). Smithers' ClaudeCodeAgent already reads the Anthropic names. Failed attempts no longer carry adapter-attached usage, which matches ClaudeAgent. The old tests fed a result line with DeepSeek field names that Claude Code never prints, so they passed while real runs hung. The replacement runs a Claude-Code-shaped result line through generate(): on main it fails with "DeepSeek result usage contains unsupported legacy alias input_tokens" thrown from the stdout listener; now it resolves with the reported usage. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…ging tasks
The Kimi and Pi adapters parse CLI output inside onStdoutLine and
onStderrLine, which Smithers calls from the child process 'data'
listeners. A throw there escapes the listener instead of failing the
invocation, so the task waits for its idle or node timeout. Kimi threw
on any stdout/stderr line starting with "{" that was not strict JSON (a
Node util.inspect dump such as "{ code: 'ECONNRESET' }", a duplicate
key, a line over 1 MiB) and on a malformed resume hint; Pi threw when a
provider's totalTokens did not equal its components or its cost
breakdown was missing or inconsistent.
Kimi also read session recovery and wire usage from files in onExit
and at buildCommand, where any torn, oversized or replaced record
failed the invocation. A wire torn by a killed attempt made every later
resume of that session fail at buildCommand.
Make these reads total: an ambiguous resume hint carries no session,
an unreadable wire or baseline leaves the invocation's usage absent,
and an inconsistent Pi usage event stays uncounted. Usage is still
never partially counted or fabricated.
Tests that asserted these throws are flipped. New end-to-end tests run
the adapters on malformed output; against main they fail with the
throws quoted above (e.g. "Kimi output JSON is invalid", "Pi assistant
usage totalTokens does not equal its token component sum", "Kimi wire
has a torn or unterminated final record").
Co-Authored-By: Claude Opus 5.5 <[email protected]>
…under the target workflowControlChildEnvironment blanks every child value that contains a controller root. It derived those roots from ULTRAFUZZ_WORKFLOW_ PERSISTED_PATH and from "/modules/", "/controls/", "/dependencies/" and "/.smithers/workflows/" markers in controller variables. A native continuation (`ultrafuzz resume`) persists the target's own .smithers/workflows/*.tsx, so the whole project root became a "controller root" and adapter-supplied paths under it were set to "": OpenCode's run-scoped XDG_* and OPENCODE_DB fell back to the operator's real directories and Kimi's API-key home disappeared. Only Pi had been special-cased (#1035). Derive the roots only from the snapshot names the process anchor advertises (ULTRAFUZZ_SNAPSHOT_PERSISTED_ROOT/PROCESS_ROOT/SOURCE_ROOT), which every sealed launch sets and a native continuation does not, and drop the Pi special case. The marker heuristics also matched ordinary install paths (a checkout under any ".../modules/..." directory made its parent a root). ULTRAFUZZ_BUN_MODULE_CONFINEMENT, previously blanked only through those heuristics, is now blanked by name like the other controller-only variables. The continuation test asserted the old over-blanking; it now checks that inherited and adapter-supplied homes under the target survive while paths under an advertised snapshot are still blanked. A new OpenCode test fails on main with XDG_CONFIG_HOME === "". Co-Authored-By: Claude Opus 5.5 <[email protected]>
…implementation The provider route ID an adapter re-verifies on every invocation was computed by two hand-maintained copies (environment.tsx for the adapters, data-governance.ts at plan time) and both digested input that does not select a destination: - a Codex config.toml with any provider line was hashed whole, so the CLI's own rewrites ([marketplaces.*] last_updated, [projects.*] trust_level, #908) failed every later task; - proxy variables counted for every agent, and AWS_/GOOGLE_/AZURE_/ FOUNDRY_ variables counted for Claude even without a cloud platform flag (FOUNDRY_PROFILE is forge's profile variable), so resuming from another shell failed every task with "provider route changed after disclosure acknowledgement". providerRouteDestination in data-governance.ts is now the only implementation. Plan-time modelDestination calls it, and the generated environment.tsx loads it from the runtime module the rendered workflow already imports (ULTRAFUZZ_RUNTIME_MODULE, with the same snapshot fallback as workflow.tsx). It digests the agent's endpoint and platform variables, with Claude cloud prefixes counted only while a matching CLAUDE_CODE_USE_* flag is set; the Codex provider that model_provider (or a profile) selects, limited to its id, base_url, wire_api and env_key and parsed with smol-toml; Claude settings helpers and routing env entries; and, as before, the whole Kimi config. A config file the reader cannot parse is digested by its bytes, as before. Proxies no longer count. effectiveRouteEnvironment keeps its wider set for cloud credential forwarding. Route IDs that included the dropped inputs change, so affected policies must be re-acknowledged and affected paused runs re-planned. New tests fail on main with "provider route changed" for a proxy change, a Codex rewrite of unrelated sections, and changed ambient Claude cloud settings. Co-Authored-By: Claude Opus 5.5 <[email protected]>
Planning accepts only the byte-exact packaged .smithers/agents closure (controller-source.ts), but init treated those files as project-owned: without --force it preserved stale or customized copies, emitted review warnings, and upgraded adapters only during a recognized 0.32 manifest migration. Every adapter fix therefore left existing projects unplannable until the operator ran `init --force`, which also resets ultrafuzz.toml, the topology and the prompts. Every init now rewrites the files in STOCK_CONTROLLER_SOURCE_TEMPLATES, the same list planning checks, and the preserve/review/0.32-adapter upgrade code and its INIT_AGENT_* diagnostics are deleted. Project-owned files are preserved as before. A symlink, hard link or special file at an adapter path now fails init with INIT_PATH_UNSAFE, as `init --force` already did, instead of being preserved with a warning; the replacement open adds O_NONBLOCK so a FIFO fails rather than blocking init. controller-source now tells the operator to rerun `ultrafuzz init`. The tests that asserted preservation and review warnings are rewritten to assert the refresh and that linked paths are never followed or written through; against main they fail (6 of 6). Co-Authored-By: Claude Opus 5.5 <[email protected]>
Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano
force-pushed
the
claude/w13-adapters-fail-open
branch
from
September 29, 2026 00:06
c897f85 to
4dacfd7
Compare
…stry "legacy projects do not require newly added opt-in agent factories" ran a plain init and then asserted the hand-trimmed registry was still on disk. Plain init now rewrites the stock adapter closure, so that assertion locked in the removed preserve-on-init behaviour and failed on every run. The test now covers only what its name says: validate accepts a legacy registry that lacks an opt-in factory the config does not reference. Plain init restoring the closure is covered by the non-force init refresh tests. Co-Authored-By: Claude Opus 5.5 <[email protected]>
Plain init opened every stock adapter for writing, even one already byte-identical to its packaged template. An up-to-date adapter the caller cannot write (mode 0444, or left root-owned by an earlier sudo or container init) therefore failed plain init with INIT_PATH_UNSAFE, although planning accepts that file; origin/main preserved it and succeeded. init now reads each existing adapter through the same bounded single-link reader it uses elsewhere and skips it when the bytes already match. Missing, linked, special, or different files still go to the anchored writer, so the INIT_PATH_UNSAFE behaviour for symlinks, hard links, and FIFOs is unchanged. The anchored-open swap test now targets the write open of a stale adapter: its hook used to fire on the first open of environment.ts, which is now the read-only comparison, and an identical file is no longer rewritten at all. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…id not count Two gaps remained after the Pi parser stopped throwing in the stdout listener: - A response that overflowed a running total was half applied: the message count and fresh-input total were updated before the output sum threw, so that response's input was still counted. The next totals are now computed first and committed only after every check passes. - With nothing countable, the completed event kept Smithers' raw last-response usage object, from which the engine reads only `totalTokens`, so the invocation reported a lone total from one uncounted response. The adapter now removes that usage when its own aggregate is absent. Per-response dropping stays: earlier responses were already reported as cumulative usage snapshots, which the engine persists as replacements, so a later inconsistent response cannot make the invocation's usage unknown. The comment, test, and docs now say that such an invocation's usage is a lower bound, instead of claiming usage is never partially counted. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…flags read like Claude Code - On origin/main a Codex config.toml with a model_provider line was hashed whole, so changing the top-level openai_base_url (which redirects the built-in openai provider) changed the route ID. The selected-provider record dropped it, so after acknowledgement that redirect went unnoticed. The record now includes openai_base_url. - A CLAUDE_CODE_USE_* flag counted as set for any non-blank value, so CLAUDE_CODE_USE_BEDROCK=0 still pinned every AWS_ variable. Claude Code 2.1.284 reads these flags as set only for 1, true, yes, or on (any case, trimmed); the route digest now does the same. - The providerRouteDestination docstring no longer says every generated adapter re-verifies its invocation: Pi and OpenCode pass no route and are checked only at plan time. Co-Authored-By: Claude Opus 5.5 <[email protected]>
…drop a dead scan environment.tsx already imports @ultrafuzz/runtime at load time for the route digest, yet it still kept copies of ROUTE_ENV_PREFIXES, the sensitive-name pattern, isCredentialLikeEnvironmentVariableName, and routeOwnsCredentialLikeEnvironmentVariable, under a comment calling the file dependency-free. The copies were identical to the runtime exports the controller uses; the adapter now takes those two functions from the same import, so the prefix list and the credential pattern have one definition. Both exports already exist on origin/main, so this adds no new downgrade requirement beyond the one providerRouteDestination introduced. smithersExecutionControlFiles scanned each adapter for "ultrafuzz.toml" without ULTRAFUZZ_CONFIG_PATH right after assertControllerSourceDigest, which admits only the byte-exact packaged closure. Every packaged file that mentions ultrafuzz.toml also reads ULTRAFUZZ_CONFIG_PATH, so the scan could never fire; it is deleted along with its stale "init --force" advice. Co-Authored-By: Claude Opus 5.5 <[email protected]>
- docs/security.md listed the route digest as covering only route-bearing input and said every agent invocation re-verifies it. Any non-credential variable with the agent's provider prefix still counts (ANTHROPIC_LOG included), and Pi and OpenCode are checked only at plan time. The section now lists exactly what the digest covers and which adapters re-check it. - The breaking CHANGELOG entry said only IDs that included dropped inputs change. Every ID derived from a provider config file changes, because the config part is now a structured record rather than the file hash; it now says so, and that env-only IDs without the dropped inputs are unchanged. - "Upgrades no longer need init --force" held only for the adapter closure; the CHANGELOG and CLI reference now say that, and that an adapter already matching its template is left untouched. Co-Authored-By: Claude Opus 5.5 <[email protected]>
Every pull request in this batch inserts its entry at the same place in CHANGELOG.md, so each merge would conflict with the next. The entries are collected into one changelog update instead. Co-Authored-By: Claude Opus 5.5 <[email protected]>
This was referenced Sep 29, 2026
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Several agent-adapter quirks stop or stall campaigns for reasons unrelated to the audit itself:
input_tokens/cache_read_input_tokensas "legacy aliases". Claude Code always prints Anthropic names, so each result line threw fromonStdoutLine, and the task waited for its idle or node timeout.deepseek.tsxalso calledpath.joinwithout importingpath.{that was not strict JSON, for example a Nodeutil.inspectdump such as{ code: 'ECONNRESET' }.totalTokensdiffered from its components, or when its cost breakdown was missing.onExitand its buildCommand usage baseline threw on torn, oversized or replaced wire files. A wire torn by a killed attempt made every later resume of that session fail.config.toml, which the CLI rewrites itself ([marketplaces.*] last_updated,[projects.*] trust_level, Codex CLI config mutations change the acknowledged provider route mid-run, failing every sandbox agent #908).AWS_/GOOGLE_/AZURE_/FOUNDRY_variables even without Bedrock/Vertex/Foundry.FOUNDRY_PROFILEis forge's profile variable.AWS_PROFILE, failed every later task with "provider route changed after disclosure acknowledgement".ultrafuzz resume, the env scrub treated the whole project root as controller state. It set OpenCode's run-scopedXDG_*/OPENCODE_DBand Kimi's API-key home to"", so OpenCode fell back to the operator's real directories.init --force. Planning accepts only the byte-exact packaged.smithers/agentsclosure, but plaininitpreserved stale copies. The only advice wasinit --force, which also resetsultrafuzz.toml, the topology and the prompts.agent-adapter-boundaries.test.tspinned each template's SHA-256 twice, plus line and syntax-node ceilings. So the one-lineimport pathfix failed with "changed from its reviewed source fingerprint".Root cause
datalisteners, where a throw escapes the invocation instead of failing it.environment.tsxand one indata-governance.ts. Both fed whole files and ambient env into the digest.ULTRAFUZZ_WORKFLOW_PERSISTED_PATHand path-marker heuristics. In a native continuation that path is the target's own workflow.inittreated the adapter closure as project-owned, while planning treats it as packaged code.Change
DeepSeek:
createOutputInterpreter/generate/streamand their parsing and attachment helpers. Smithers'ClaudeCodeAgentalready reads the Anthropic names.pathimport.no-undefforpackages/runtime/src/templates/**/*.tsx, with the__ULTRAFUZZ_*__placeholders declared as readonly globals. Against main the rule reports onlydeepseek.tsx 200:59 'path' is not defined.Kimi and Pi fail open.
totalTokensfrom it.One route helper.
providerRouteDestination(agent, env, routeConfig?)indata-governance.tsis now the only digest. Plan-timemodelDestinationcalls it.environment.tsxloads it from the runtime module the rendered workflow already imports (ULTRAFUZZ_RUNTIME_MODULE, same fallback path asworkflow.tsx), together withisCredentialLikeEnvironmentVariableNameandrouteOwnsCredentialLikeEnvironmentVariable, whose adapter-side copies are deleted. What the digest covers:CLAUDE_CODE_USE_*flag is1,true,yesoron, in the process environment or insettings.jsonenv. These are the six flags Claude Code 2.1.284 checks, with the truthiness it applies to them;model_provider(aprofilemay select it), that provider'sbase_url/wire_api/env_key, and the top-levelopenai_base_url. The file is parsed withsmol-toml(new runtime dependency, already in the workspace);enventries under the same rules as the environment;ANTHROPIC_LOG(see below).effectiveRouteEnvironmentis unchanged because cloud credential forwarding still uses it.Env scrub:
ULTRAFUZZ_SNAPSHOT_{PERSISTED,PROCESS,SOURCE}_ROOT. The process anchor sets them for every sealed launch; a native continuation does not.ULTRAFUZZ_BUN_MODULE_CONFINEMENT, previously blanked only through the marker heuristics, is now blanked by name.init:
initrewrites each stock adapter (STOCK_CONTROLLER_SOURCE_TEMPLATES, the list planning checks) that differs from its packaged template. An adapter that already matches is left untouched, so an up-to-date read-only closure does not fail init.INIT_PATH_UNSAFE, the same asinit --forcedid. The replacement open addsO_NONBLOCK, so a FIFO fails instead of hanging init.controller-sourcemessages now sayrerun ultrafuzz init.ultrafuzz.tomltext scan insmithersExecutionControlFilesis deleted. It ran right afterassertControllerSourceDigest, which admits only the packaged closure, and every packaged file mentioningultrafuzz.tomlalso readsULTRAFUZZ_CONFIG_PATH.Boundary test: merge the two policy tables, and delete the fingerprints, the ceilings and the tests that only exercised them. Kept: every source needs a policy, only registered adapters own responsibilities, helpers carry no orchestration signals, and a detected responsibility fails until declared.
Docs:
docs/security.md: what a route ID covers, and which adapters re-check it.docs/config.md: DeepSeek, Kimi and Pi usage, and init.docs/reference/cli.md: init.docs/reference/agent-adapter-boundaries.md.Deliberately not built (and why)
usageevent at every countedmessage_end. The pinned-engine patch persists each one as a replacement snapshot for the attempt (SMITHERS_ENGINE_AGENT_USAGE_PROGRESS_PATCH). A later inconsistent response therefore cannot withdraw what was already recorded. Going by the patch source (not a test), marking the final aggregate unknown would leave the ledger at the last snapshot before the bad response. That is a looser lower bound than counting every consistent response.ANTHROPIC_LOG,ANTHROPIC_MODELandOPENAI_ORG_ID, so a resume from a shell where one of them differs still fails route verification. An allowlist of route-selecting names would fix that, but it fails open on any new routing variable a CLI adds, so it needs an owner decision. Follow-up.openai_base_urlwithout amodel_providerline is still not digested, as on main. This PR restores main's coverage when a provider is selected. Covering the unselected case would change acknowledged destinations for configs that plan today. Follow-up.~/.claude-style 0775 homes still fail. The OpenRouter "preserves opaque model IDs…" adapter test fails on main and on this branch here with "provider-home ancestors cannot be group/world writable".claude.tsx(w08's file) andopencode.tsx;Verification
Rows marked † were added in the review round. For those rows, I copied this branch's test files into detached worktrees at origin/main (b6dd1da) and at the previous PR head (4dacfd7), compiled them against that side's sources, and ran them; route rows were also probed with
modelDestinationfrom each side's build. The other rows are from the first round: the implementer ran them against origin/main (swapping main's template into the worktree, or copying the tests into a main worktree), and both reviewers reproduced the main/head results.generate()on a Claude-Code-shaped result lineDeepSeek result usage contains unsupported legacy alias input_tokensthrown from the stdout listenerKimi output JSON is invalid,Kimi wire record is invalid strict JSON)Pi assistant usage totalTokens does not equal its token component suminputTokens: 5, expected 4). With only that fixed, nothing countable reportstotalTokens: 3XDG_CONFIG_HOMEis""AWS_PROFILE/GOOGLE_CLOUD_PROJECT/FOUNDRY_PROFILE, plan/adapter paritymodel_provider = "openai"withopenai_base_urlchanged (route probe anddata-governance.test.ts)CLAUDE_CODE_USE_BEDROCK=0,AWS_REGIONchanged (route probe anddata-governance.test.ts)AWS_variable always counted)" Yes "still changes itsettings.json,AWS_REGIONchanged, then no flag anywhere (row corrected in review)model:anthropic(AWS_REGIONalways counted on main)INIT_PATH_UNSAFEinitrestores a customized adapterwarning: INIT_AGENT_ADAPTER_UPDATE_REQUIREDno-undefon main'sdeepseek.tsx(via--stdin)'path' is not definedAlso run on the final code:
data-governance(11),agent-adapter-boundaries(8) andcontroller-source: 20/20.runtime.test.tswhose names match init, registry, adapter, route, governance, disclosure, materialization, Modal, linked workflow or native continuation: 49 passed on the first run. The 50th, "init does not modify a regular file swapped after the anchored open", failed until its hook was retargeted at the write open of a stale adapter (see the init commit). The init group then passed 11/11 on rerun.node scripts/run-pr-smoke-tests.mjs(the PR smoke lane): 35 + 12 node tests and 4 Bun tests, all pass.^Bun adapter contract:suite, on code identical to the final commit except for comments: 48 pass, 1 skip, 3 fail.startRunsubmission, before any adapter code runs: the fake runner was killed with SIGTERM. The host load average was 13 to 24 from other agents' test runs. Two standalone reruns also timed out, at load 7 to 24.npx prettier --checkandnpx eslinton changed files,CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/main pnpm -w lint:strict:ci, runtime and CLI typecheck,pnpm -w knip,node scripts/docs-check.mjs: all pass.Not run:
Risk / compatibility
config.toml, Claudesettings.json, Kimiconfig.toml. The config part is now a structured record rather than the file hash. IDs that included a proxy or an unselected Claude cloud variable also change. Env-only IDs without those inputs keep their exact value; the hard-codedmodel:kimi-route-be51…test fixture still matches.ULTRAFUZZ_DATA_GOVERNANCE_POLICYand re-acknowledge.ultrafuzz init, thenresume, orresume --refresh-controller). Re-plan it rather than resume it.providerRouteDestinationfails the first route check with a TypeError. The two credential helpers it now also imports exist on origin/main.initnow rewrites any adapter that differs from its template on every run. Customized adapters were already rejected at plan. A symlinked adapter now fails init instead of producing a warning. An adapter that differs and cannot be written also fails init, withINIT_PATH_UNSAFE.packages/runtime/scripts/run-pr-smoke-tests.mjsnow lists the renamed DeepSeek tests.External static analysis(Super-Linter markdownlint MD013, 400-character lines) fails onCHANGELOG.md, as it does on fix(runtime): give agent retries a real wait, stop retrying deterministic failures, and label timeouts by code #1171, fix(runtime): name the runner error when a workflow fails with no failed node, and pin same-id recovery #1172 and fix(security): stop the publication secret gate rejecting public test keys and JWT-shaped identifiers #1165. About 43 existing entries already exceed the limit. This PR's two entries, one line each like their neighbours, exceed it too. The check gates the full release-validation lanes.claude.tsxanoutput-interpretationresponsibility (Concurrent agents race on OAuth token refresh; three immediate retries all re-race and kill the run with an opaque "Claude run failed" #1084) by editing both fingerprint tables inagent-adapter-boundaries.test.tsand the doc table's Baseline column. Whichever lands second should carry that responsibility and its issue link into this PR's singleadapterPoliciesentry forclaude.tsx, with no hash or ceiling. CHANGELOG anddocs/reference/cli.mdhunks may also conflict textually with parallel PRs; those conflicts are local.🤖 Generated with Claude Code
The PR appears safe to merge, with an open non-blocking concern about review coverage for adapter changes the boundary detector cannot recognize.
Fix with agent prompt
Summary
The PR refreshes stock adapters during plain
init, consolidates provider-route calculation, makes adapter telemetry failures non-fatal, and narrows continuation environment scrubbing. The changes since the previous review also revise synchronization, resume, and artifact bookkeeping. No distinct new actionable issue was established; the earlier adapter-boundary tripwire concern remains open.Diagram
%%{init: {'theme': 'neutral'}}%% flowchart LR Plan[Plan run] --> Route[Shared provider-route calculation] Route --> Ack[Disclosure acknowledgement] Ack --> Adapter[Agent adapter invocation] Adapter --> RouteReviews (5) · Last reviewed commit: "Merge origin/main into claude/w13-adapte..."