Skip to content

fix(runtime): agent adapters never hang or reroute on telemetry and environment quirks - #1173

Merged
aviggiano merged 16 commits into
mainfrom
claude/w13-adapters-fail-open
Sep 29, 2026
Merged

aviggiano merged 16 commits into
mainfrom
claude/w13-adapters-fail-open

Conversation

@aviggiano

@aviggiano aviggiano commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

Several agent-adapter quirks stop or stall campaigns for reasons unrelated to the audit itself:

  • DeepSeekAgent hangs every finished task. Its usage parser required DeepSeek's OpenAI-compatible field names and rejected input_tokens / cache_read_input_tokens as "legacy aliases". Claude Code always prints Anthropic names, so each result line threw from onStdoutLine, and the task waited for its idle or node timeout. deepseek.tsx also called path.join without importing path.
  • Kimi and Pi throw inside the child's output listeners.
    • Kimi threw on any stdout/stderr line starting with { that was not strict JSON, for example a Node util.inspect dump such as { code: 'ECONNRESET' }.
    • Pi threw when a provider's totalTokens differed from its components, or when its cost breakdown was missing.
    • Kimi's onExit and its buildCommand usage baseline threw on torn, oversized or replaced wire files. A wire torn by a killed attempt made every later resume of that session fail.
  • The provider route digest drifts on non-route input.
    • Codex configs with a custom provider hashed the whole config.toml, which the CLI rewrites itself ([marketplaces.*] last_updated, [projects.*] trust_level, Codex CLI config mutations change the acknowledged provider route mid-run, failing every sandbox agent #908).
    • The digest captured proxies for every agent.
    • For Claude it captured AWS_/GOOGLE_/AZURE_/FOUNDRY_ variables even without Bedrock/Vertex/Foundry. FOUNDRY_PROFILE is forge's profile variable.
    • Result: a CLI rewrite, or a resume from a shell with a different proxy or AWS_PROFILE, failed every later task with "provider route changed after disclosure acknowledgement".
  • Native continuations blank adapter-owned paths. On ultrafuzz resume, the env scrub treated the whole project root as controller state. It set OpenCode's run-scoped XDG_*/OPENCODE_DB and Kimi's API-key home to "", so OpenCode fell back to the operator's real directories.
  • Every adapter fix forces init --force. Planning accepts only the byte-exact packaged .smithers/agents closure, but plain init preserved stale copies. The only advice was init --force, which also resets ultrafuzz.toml, the topology and the prompts.
  • The boundary gate is a change detector. agent-adapter-boundaries.test.ts pinned each template's SHA-256 twice, plus line and syntax-node ceilings. So the one-line import path fix failed with "changed from its reviewed source fingerprint".

Root cause

  • The adapter parsers treated telemetry as authoritative input and threw on anything unexpected. Smithers calls these hooks from the child process data listeners, where a throw escapes the invocation instead of failing it.
  • The route digest had two hand-maintained copies, one in environment.tsx and one in data-governance.ts. Both fed whole files and ambient env into the digest.
  • The env scrub derived "controller roots" from ULTRAFUZZ_WORKFLOW_PERSISTED_PATH and path-marker heuristics. In a native continuation that path is the target's own workflow.
  • init treated the adapter closure as project-owned, while planning treats it as packaged code.

Change

  • DeepSeek:

    • Delete the usage overrides: createOutputInterpreter/generate/stream and their parsing and attachment helpers. Smithers' ClaudeCodeAgent already reads the Anthropic names.
    • Add the missing path import.
    • Enable ESLint no-undef for packages/runtime/src/templates/**/*.tsx, with the __ULTRAFUZZ_*__ placeholders declared as readonly globals. Against main the rule reports only deepseek.tsx 200:59 'path' is not defined.
  • Kimi and Pi fail open.

    • An ambiguous Kimi resume hint carries no session.
    • An unreadable Kimi wire or baseline leaves that invocation's usage absent.
    • A Pi response whose usage is inconsistent, invalid, or would overflow a running total is left out whole. That invocation's usage is then a lower bound.
    • When no Pi response is countable, the completed event carries no usage. Before, Smithers' raw last-response object stood in, and the engine read a lone totalTokens from it.
  • One route helper. providerRouteDestination(agent, env, routeConfig?) in data-governance.ts is now the only digest. Plan-time modelDestination calls it. environment.tsx loads it from the runtime module the rendered workflow already imports (ULTRAFUZZ_RUNTIME_MODULE, same fallback path as workflow.tsx), together with isCredentialLikeEnvironmentVariableName and routeOwnsCredentialLikeEnvironmentVariable, whose adapter-side copies are deleted. What the digest covers:

    • proxies are dropped;
    • Claude cloud prefixes count only while a matching CLAUDE_CODE_USE_* flag is 1, true, yes or on, in the process environment or in settings.json env. These are the six flags Claude Code 2.1.284 checks, with the truthiness it applies to them;
    • Codex contributes the selected model_provider (a profile may select it), that provider's base_url/wire_api/env_key, and the top-level openai_base_url. The file is parsed with smol-toml (new runtime dependency, already in the workspace);
    • Claude contributes its credential/process helper keys and its env entries under the same rules as the environment;
    • Kimi keeps its whole-file digest;
    • a file the reader cannot parse is digested by its exact bytes, as before;
    • any other non-credential variable with the agent's provider prefix still counts, including non-routing ones such as ANTHROPIC_LOG (see below).

    effectiveRouteEnvironment is unchanged because cloud credential forwarding still uses it.

  • Env scrub:

    • Aliasing roots come only from the advertised ULTRAFUZZ_SNAPSHOT_{PERSISTED,PROCESS,SOURCE}_ROOT. The process anchor sets them for every sealed launch; a native continuation does not.
    • The Pi special case (Detached resume drops external Pi CLI PATH before agent preflight #1035) is deleted.
    • ULTRAFUZZ_BUN_MODULE_CONFINEMENT, previously blanked only through the marker heuristics, is now blanked by name.
  • init:

    • Every init rewrites each stock adapter (STOCK_CONTROLLER_SOURCE_TEMPLATES, the list planning checks) that differs from its packaged template. An adapter that already matches is left untouched, so an up-to-date read-only closure does not fail init.
    • The preserve/review/0.32-adapter-upgrade code and its diagnostics are deleted.
    • A linked or special file at an adapter path now fails with INIT_PATH_UNSAFE, the same as init --force did. The replacement open adds O_NONBLOCK, so a FIFO fails instead of hanging init.
    • controller-source messages now say rerun ultrafuzz init.
    • The now-unreachable ultrafuzz.toml text scan in smithersExecutionControlFiles is deleted. It ran right after assertControllerSourceDigest, which admits only the packaged closure, and every packaged file mentioning ultrafuzz.toml also reads ULTRAFUZZ_CONFIG_PATH.
  • Boundary test: merge the two policy tables, and delete the fingerprints, the ceilings and the tests that only exercised them. Kept: every source needs a policy, only registered adapters own responsibilities, helpers carry no orchestration signals, and a detected responsibility fails until declared.

  • Docs:

    • docs/security.md: what a route ID covers, and which adapters re-check it.
    • docs/config.md: DeepSeek, Kimi and Pi usage, and init.
    • docs/reference/cli.md: init.
    • docs/reference/agent-adapter-boundaries.md.
    • CHANGELOG has one breaking and one other entry.

Deliberately not built (and why)

  • No legacy-digest compatibility path. Accepting old digests would keep both implementations alive. Affected runs are called out under Risk instead.
  • Pi usage is not all-or-nothing per invocation (Kimi's policy). Pi emits a cumulative usage event at every counted message_end. The pinned-engine patch persists each one as a replacement snapshot for the attempt (SMITHERS_ENGINE_AGENT_USAGE_PROGRESS_PATCH). A later inconsistent response therefore cannot withdraw what was already recorded. Going by the patch source (not a test), marking the final aggregate unknown would leave the ledger at the last snapshot before the bad response. That is a looser lower bound than counting every consistent response.
  • Prefix-matched non-routing variables still count. Examples are ANTHROPIC_LOG, ANTHROPIC_MODEL and OPENAI_ORG_ID, so a resume from a shell where one of them differs still fails route verification. An allowlist of route-selecting names would fix that, but it fails open on any new routing variable a CLI adds, so it needs an owner decision. Follow-up.
  • Codex openai_base_url without a model_provider line is still not digested, as on main. This PR restores main's coverage when a provider is selected. Covering the unselected case would change acknowledged destinations for configs that plan today. Follow-up.
  • Per-invocation route re-verification is kept, not removed. The spec asks to narrow and share the check; dropping it is an owner decision. Pi and OpenCode pass no route, so their destination is checked only at plan time, as before.
  • The 0700 provider-home requirement is untouched (owner decision). Follow-up: this host's ~/.claude-style 0775 homes still fail. The OpenRouter "preserves opaque model IDs…" adapter test fails on main and on this branch here with "provider-home ancestors cannot be group/world writable".
  • Out of scope (other items' files):
    • the false "edit this generated file" comments in claude.tsx (w08's file) and opencode.tsx;
    • Kimi's session-index seeding at buildCommand, which is still strict (session handling, not telemetry).

Verification

Rows marked † were added in the review round. For those rows, I copied this branch's test files into detached worktrees at origin/main (b6dd1da) and at the previous PR head (4dacfd7), compiled them against that side's sources, and ran them; route rows were also probed with modelDestination from each side's build. The other rows are from the first round: the implementer ran them against origin/main (swapping main's template into the worktree, or copying the tests into a main worktree), and both reviewers reproduced the main/head results.

Test origin/main 4dacfd7 this branch
DeepSeek generate() on a Claude-Code-shaped result line fails: DeepSeek result usage contains unsupported legacy alias input_tokens thrown from the stdout listener passes passes
Kimi ambiguous hints, unreadable-wire table, torn resumed wire, malformed output end-to-end fail (Kimi output JSON is invalid, Kimi wire record is invalid strict JSON) pass pass
† Pi inconsistent usage: inconsistent + consistent responses, an overflowing response, nothing countable fails: Pi assistant usage totalTokens does not equal its token component sum fails at the overflow case: its input is still counted (inputTokens: 5, expected 4). With only that fixed, nothing countable reports totalTokens: 3 passes
Continuation scrub, and OpenCode native continuation fail: XDG_CONFIG_HOME is "" pass pass
Codex proxy change, Codex CLI rewrite of unrelated sections, Claude ambient AWS_PROFILE/GOOGLE_CLOUD_PROJECT/FOUNDRY_PROFILE, plan/adapter parity 4 fail with "provider route changed…" or an assertion pass pass
† Codex model_provider = "openai" with openai_base_url changed (route probe and data-governance.test.ts) route ID changes (whole-file hash) route ID unchanged: assertion fails route ID changes
† CLAUDE_CODE_USE_BEDROCK=0, AWS_REGION changed (route probe and data-governance.test.ts) route ID changes (every AWS_ variable always counted) route ID changes: assertion fails unchanged; " Yes " still changes it
Bedrock enabled only in Claude settings.json, AWS_REGION changed, then no flag anywhere (row corrected in review) fails at the no-flag case: a route instead of model:anthropic (AWS_REGION always counted on main) passes passes
Init refresh, linked/dangling/FIFO adapter paths, Modal device split, startRun message 6 fail pass pass
† Plain init with an up-to-date adapter at mode 0444 passes (main never rewrote existing adapters) fails: INIT_PATH_UNSAFE passes; inode, mode and mtime unchanged
† "legacy projects do not require newly added opt-in agent factories" passes the original test fails: plain init restored the registry it asserted was preserved (reviewer 2). The rewritten test, without that assertion, passes passes
CLI: plain init restores a customized adapter fails: adapter preserved with warning: INIT_AGENT_ADAPTER_UPDATE_REQUIRED passes passes
ESLint no-undef on main's deepseek.tsx (via --stdin) reports 'path' is not defined clean clean

Also run on the final code:

  • Node, runtime:
    • data-governance (11), agent-adapter-boundaries (8) and controller-source: 20/20.
    • The 50 node tests in runtime.test.ts whose names match init, registry, adapter, route, governance, disclosure, materialization, Modal, linked workflow or native continuation: 49 passed on the first run. The 50th, "init does not modify a regular file swapped after the anchored open", failed until its hook was retargeted at the write open of a stale adapter (see the init commit). The init group then passed 11/11 on rerun.
  • Node, CLI: the 2 init tests.
  • node scripts/run-pr-smoke-tests.mjs (the PR smoke lane): 35 + 12 node tests and 4 Bun tests, all pass.
  • Bun, full ^Bun adapter contract: suite, on code identical to the final commit except for comments: 48 pass, 1 skip, 3 fail.
    • The OpenRouter "preserves opaque model IDs" failure is the host's 0700 provider-home issue above.
    • "native reset-node continuation reaches Pi preflight…" timed out at startRun submission, before any adapter code runs: the fake runner was killed with SIGTERM. The host load average was 13 to 24 from other agents' test runs. Two standalone reruns also timed out, at load 7 to 24.
    • Once the load fell to 3 to 5, it passed 4 of 4 standalone runs on the final commit (84 to 96 s). The same test at 4dacfd7 passed 4 of 4, alternating with those runs (80 to 87 s).
    • "generated Pi adapter binds OpenRouter…" ran immediately after that timed-out test and failed. It passes standalone, together with the other Pi test (2/2).
  • Bun, the same suite on the final commit, without the continuation test, at load average 15 to 24: 37 pass, 1 skip, 13 fail. Besides the 0700 test, the failures were OpenRouter/Codex deadline and timeout assertions (2 s recovery budgets, 5 s test timeouts, event-loop delay). One OpenRouter counter assertion was reported under the next test, a DeepSeek one.
    • A rerun of the OpenRouter, Codex and DeepSeek tests: 19 pass, 2 fail (the 0700 test, and "rethrows the last 429 without starting an attempt at its retry deadline").
    • That test also failed 3 of 3 runs at 4dacfd7 at the same load, and 2 of 3 on this branch.
  • npx prettier --check and npx eslint on changed files, CI=1 ESLINT_PLUGIN_DIFF_COMMIT=origin/main pnpm -w lint:strict:ci, runtime and CLI typecheck, pnpm -w knip, node scripts/docs-check.mjs: all pass.

Not run:

  • the complete runtime and CLI node:test suites, which are too large for this shared host;
  • any real provider CLI invocation.

Risk / compatibility

  • Route IDs change. Every route ID derived from a provider config file changes: Codex config.toml, Claude settings.json, Kimi config.toml. The config part is now a structured record rather than the file hash. IDs that included a proxy or an unselected Claude cloud variable also change. Env-only IDs without those inputs keep their exact value; the hard-coded model:kimi-route-be51… test fixture still matches.
    • Operators with such routes must update ULTRAFUZZ_DATA_GOVERNANCE_POLICY and re-acknowledge.
    • A paused run with such an ID fails route verification once its adapters are refreshed (ultrafuzz init, then resume, or resume --refresh-controller). Re-plan it rather than resume it.
  • The dropped inputs are no longer drift-checked mid-run. That is intentional for proxies; a TLS-intercepting proxy is not treated as a separate destination.
  • Generated adapters now import the runtime module at load time, exactly as the rendered workflow does. A process that cannot load that module could not load the workflow either. Pairing refreshed adapters with an older installed runtime (a downgrade) that lacks providerRouteDestination fails the first route check with a TypeError. The two credential helpers it now also imports exist on origin/main.
  • init now rewrites any adapter that differs from its template on every run. Customized adapters were already rejected at plan. A symlinked adapter now fails init instead of producing a warning. An adapter that differs and cannot be written also fails init, with INIT_PATH_UNSAFE.
  • CI:
  • Merge conflicts with fix(runtime): give agent retries a real wait, stop retrying deterministic failures, and label timeouts by code #1171. It gives claude.tsx an output-interpretation responsibility (Concurrent agents race on OAuth token refresh; three immediate retries all re-race and kill the run with an opaque "Claude run failed" #1084) by editing both fingerprint tables in agent-adapter-boundaries.test.ts and the doc table's Baseline column. Whichever lands second should carry that responsibility and its issue link into this PR's single adapterPolicies entry for claude.tsx, with no hash or ceiling. CHANGELOG and docs/reference/cli.md hunks may also conflict textually with parallel PRs; those conflicts are local.

🤖 Generated with Claude Code

RetriggerConfidence Score: 5/5

The PR appears safe to merge, with an open non-blocking concern about review coverage for adapter changes the boundary detector cannot recognize.

Fix All in Claude CodeFindings

  1. P2 Opaque adapter changes go unchecked ▶
Fix with agent prompt
### Issue 1
packages/runtime/test/agent-adapter-boundaries.test.ts:840-843
Removing the source fingerprints leaves this gate checking only responsibilities its static detector recognizes. The deleted test showed that an adapter could change output handling through a helper without producing such a signal. An equivalent future edit can now pass without an explicit responsibility review. Keeping a review tripwire for changes the detector cannot recognize would preserve that coverage without requiring two fingerprints.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Summary

The PR refreshes stock adapters during plain init, consolidates provider-route calculation, makes adapter telemetry failures non-fatal, and narrows continuation environment scrubbing. The changes since the previous review also revise synchronization, resume, and artifact bookkeeping. No distinct new actionable issue was established; the earlier adapter-boundary tripwire concern remains open.

Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  Plan[Plan run] --> Route[Shared provider-route calculation]
  Route --> Ack[Disclosure acknowledgement]
  Ack --> Adapter[Agent adapter invocation]
  Adapter --> Route
Loading

Reviews (5) · Last reviewed commit: "Merge origin/main into claude/w13-adapte..."

aviggiano and others added 2 commits September 28, 2026 23:34
agent-adapter-boundaries.test.ts pinned every adapter template's SHA-256
twice (a structural policy and a separate responsibility policy) plus
line and syntax-node ceilings. Those assertions check no behaviour: a
comment edit or the one-line missing `import path` fix in deepseek.tsx
failed the gate with "changed from its reviewed source fingerprint", so
every adapter fix needed two hash edits.

Merge the two tables into one policy per source (purpose, declared
responsibilities, upstream links) and delete the fingerprints, the
ceilings, and the tests that only exercised them. The real boundary
checks stay: every source needs a policy, only registered adapters may
own responsibilities, non-adapter helpers may carry no orchestration
signals, and a statically detected responsibility still fails until it
is declared.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…for undefined names

deepseek.tsx called path.join without importing path, so a render
without ULTRAFUZZ_CONFIG_PATH threw "path is not defined" from the
factory. Nothing caught it: runtime templates are copied into projects
and run by Bun, tsconfig only includes src/**/*.ts, and ESLint turns
no-undef off for every .ts/.tsx file.

Enable no-undef for packages/runtime/src/templates/**/*.tsx and declare
the __ULTRAFUZZ_*__ placeholders the compiler substitutes as readonly
globals. Against main the rule reports exactly this bug (deepseek.tsx
200:59 'path' is not defined) and nothing else.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano requested a review from a team as a code owner September 28, 2026 23:46
Comment thread packages/runtime/src/data-governance.ts Outdated
Comment on lines 840 to +843
const sources = readSourceTree(path.join(packageRoot, "src/templates/smithers/agents"));

assert.deepEqual(
Object.keys(sourcePolicies).sort(),
Object.keys(adapterPolicies).sort(),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Opaque adapter changes go unchecked

Removing the source fingerprints leaves this gate checking only responsibilities its static detector recognizes. The deleted test showed that an adapter could change output handling through a helper without producing such a signal. An equivalent future edit can now pass without an explicit responsibility review. Keeping a review tripwire for changes the detector cannot recognize would preserve that coverage without requiring two fingerprints.

Prompt To Fix With AI
This is a comment left during a code review.
Path: packages/runtime/test/agent-adapter-boundaries.test.ts
Line: 840-843

Comment:
**Opaque adapter changes go unchecked**

Removing the source fingerprints leaves this gate checking only responsibilities its static detector recognizes. The deleted test showed that an adapter could change output handling through a helper without producing such a signal. An equivalent future edit can now pass without an explicit responsibility review. Keeping a review tripwire for changes the detector cannot recognize would preserve that coverage without requiring two fingerprints.

---

For each issue above, determine whether it is valid and should be fixed. If so, fix it directly.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

Fix in Claude Code

aviggiano and others added 6 commits September 29, 2026 00:06
DeepSeekAgent overrode the Claude Code output interpreter with a strict
usage parser that required DeepSeek's OpenAI-compatible field names
(prompt_cache_miss_tokens/prompt_cache_hit_tokens) and rejected
input_tokens and cache_read_input_tokens as "legacy aliases". Claude
Code always reports Anthropic field names, so every finished DeepSeek
task threw from onStdoutLine. That hook runs inside the child process
'data' listener, so the throw escaped the listener instead of failing
the invocation, and the task sat until the idle or node timeout.

Delete the usage overrides (createOutputInterpreter, generate, stream
and their parsing/attachment helpers). Smithers' ClaudeCodeAgent
already reads the Anthropic names. Failed attempts no longer carry
adapter-attached usage, which matches ClaudeAgent.

The old tests fed a result line with DeepSeek field names that Claude
Code never prints, so they passed while real runs hung. The
replacement runs a Claude-Code-shaped result line through generate():
on main it fails with "DeepSeek result usage contains unsupported legacy
alias input_tokens" thrown from the stdout listener; now it resolves
with the reported usage.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…ging tasks

The Kimi and Pi adapters parse CLI output inside onStdoutLine and
onStderrLine, which Smithers calls from the child process 'data'
listeners. A throw there escapes the listener instead of failing the
invocation, so the task waits for its idle or node timeout. Kimi threw
on any stdout/stderr line starting with "{" that was not strict JSON (a
Node util.inspect dump such as "{ code: 'ECONNRESET' }", a duplicate
key, a line over 1 MiB) and on a malformed resume hint; Pi threw when a
provider's totalTokens did not equal its components or its cost
breakdown was missing or inconsistent.

Kimi also read session recovery and wire usage from files in onExit
and at buildCommand, where any torn, oversized or replaced record
failed the invocation. A wire torn by a killed attempt made every later
resume of that session fail at buildCommand.

Make these reads total: an ambiguous resume hint carries no session,
an unreadable wire or baseline leaves the invocation's usage absent,
and an inconsistent Pi usage event stays uncounted. Usage is still
never partially counted or fabricated.

Tests that asserted these throws are flipped. New end-to-end tests run
the adapters on malformed output; against main they fail with the
throws quoted above (e.g. "Kimi output JSON is invalid", "Pi assistant
usage totalTokens does not equal its token component sum", "Kimi wire
has a torn or unterminated final record").

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…under the target

workflowControlChildEnvironment blanks every child value that contains
a controller root. It derived those roots from ULTRAFUZZ_WORKFLOW_
PERSISTED_PATH and from "/modules/", "/controls/", "/dependencies/" and
"/.smithers/workflows/" markers in controller variables. A native
continuation (`ultrafuzz resume`) persists the target's own
.smithers/workflows/*.tsx, so the whole project root became a
"controller root" and adapter-supplied paths under it were set to "":
OpenCode's run-scoped XDG_* and OPENCODE_DB fell back to the operator's
real directories and Kimi's API-key home disappeared. Only Pi had been
special-cased (#1035).

Derive the roots only from the snapshot names the process anchor
advertises (ULTRAFUZZ_SNAPSHOT_PERSISTED_ROOT/PROCESS_ROOT/SOURCE_ROOT),
which every sealed launch sets and a native continuation does not, and
drop the Pi special case. The marker heuristics also matched ordinary
install paths (a checkout under any ".../modules/..." directory made its
parent a root). ULTRAFUZZ_BUN_MODULE_CONFINEMENT, previously blanked only
through those heuristics, is now blanked by name like the other
controller-only variables.

The continuation test asserted the old over-blanking; it now checks that
inherited and adapter-supplied homes under the target survive while
paths under an advertised snapshot are still blanked. A new OpenCode
test fails on main with XDG_CONFIG_HOME === "".

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…implementation

The provider route ID an adapter re-verifies on every invocation was
computed by two hand-maintained copies (environment.tsx for the
adapters, data-governance.ts at plan time) and both digested input that
does not select a destination:

- a Codex config.toml with any provider line was hashed whole, so the
  CLI's own rewrites ([marketplaces.*] last_updated, [projects.*]
  trust_level, #908) failed every later task;
- proxy variables counted for every agent, and AWS_/GOOGLE_/AZURE_/
  FOUNDRY_ variables counted for Claude even without a cloud platform
  flag (FOUNDRY_PROFILE is forge's profile variable), so resuming from
  another shell failed every task with "provider route changed after
  disclosure acknowledgement".

providerRouteDestination in data-governance.ts is now the only
implementation. Plan-time modelDestination calls it, and the generated
environment.tsx loads it from the runtime module the rendered workflow
already imports (ULTRAFUZZ_RUNTIME_MODULE, with the same snapshot
fallback as workflow.tsx). It digests the agent's endpoint and platform
variables, with Claude cloud prefixes counted only while a matching
CLAUDE_CODE_USE_* flag is set; the Codex provider that model_provider
(or a profile) selects, limited to its id, base_url, wire_api and
env_key and parsed with smol-toml; Claude settings helpers and routing
env entries; and, as before, the whole Kimi config. A config file the
reader cannot parse is digested by its bytes, as before. Proxies no
longer count. effectiveRouteEnvironment keeps its wider set for cloud
credential forwarding.

Route IDs that included the dropped inputs change, so affected
policies must be re-acknowledged and affected paused runs re-planned.
New tests fail on main with "provider route changed" for a proxy change,
a Codex rewrite of unrelated sections, and changed ambient Claude cloud
settings.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Planning accepts only the byte-exact packaged .smithers/agents closure
(controller-source.ts), but init treated those files as project-owned:
without --force it preserved stale or customized copies, emitted
review warnings, and upgraded adapters only during a recognized 0.32
manifest migration. Every adapter fix therefore left existing projects
unplannable until the operator ran `init --force`, which also resets
ultrafuzz.toml, the topology and the prompts.

Every init now rewrites the files in STOCK_CONTROLLER_SOURCE_TEMPLATES,
the same list planning checks, and the preserve/review/0.32-adapter
upgrade code and its INIT_AGENT_* diagnostics are deleted. Project-owned
files are preserved as before. A symlink, hard link or special file at
an adapter path now fails init with INIT_PATH_UNSAFE, as `init --force`
already did, instead of being preserved with a warning; the replacement
open adds O_NONBLOCK so a FIFO fails rather than blocking init.
controller-source now tells the operator to rerun `ultrafuzz init`.

The tests that asserted preservation and review warnings are rewritten
to assert the refresh and that linked paths are never followed or
written through; against main they fail (6 of 6).

Co-Authored-By: Claude Opus 5.5 <[email protected]>
aviggiano and others added 6 commits September 29, 2026 04:07
…stry

"legacy projects do not require newly added opt-in agent factories" ran a
plain init and then asserted the hand-trimmed registry was still on disk.
Plain init now rewrites the stock adapter closure, so that assertion locked
in the removed preserve-on-init behaviour and failed on every run. The test
now covers only what its name says: validate accepts a legacy registry that
lacks an opt-in factory the config does not reference. Plain init restoring
the closure is covered by the non-force init refresh tests.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Plain init opened every stock adapter for writing, even one already
byte-identical to its packaged template. An up-to-date adapter the caller
cannot write (mode 0444, or left root-owned by an earlier sudo or container
init) therefore failed plain init with INIT_PATH_UNSAFE, although planning
accepts that file; origin/main preserved it and succeeded.

init now reads each existing adapter through the same bounded single-link
reader it uses elsewhere and skips it when the bytes already match. Missing,
linked, special, or different files still go to the anchored writer, so the
INIT_PATH_UNSAFE behaviour for symlinks, hard links, and FIFOs is unchanged.

The anchored-open swap test now targets the write open of a stale adapter:
its hook used to fire on the first open of environment.ts, which is now the
read-only comparison, and an identical file is no longer rewritten at all.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…id not count

Two gaps remained after the Pi parser stopped throwing in the stdout
listener:

- A response that overflowed a running total was half applied: the message
  count and fresh-input total were updated before the output sum threw, so
  that response's input was still counted. The next totals are now computed
  first and committed only after every check passes.
- With nothing countable, the completed event kept Smithers' raw
  last-response usage object, from which the engine reads only `totalTokens`,
  so the invocation reported a lone total from one uncounted response. The
  adapter now removes that usage when its own aggregate is absent.

Per-response dropping stays: earlier responses were already reported as
cumulative usage snapshots, which the engine persists as replacements, so a
later inconsistent response cannot make the invocation's usage unknown. The
comment, test, and docs now say that such an invocation's usage is a lower
bound, instead of claiming usage is never partially counted.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…flags read like Claude Code

- On origin/main a Codex config.toml with a model_provider line was hashed
  whole, so changing the top-level openai_base_url (which redirects the
  built-in openai provider) changed the route ID. The selected-provider
  record dropped it, so after acknowledgement that redirect went unnoticed.
  The record now includes openai_base_url.
- A CLAUDE_CODE_USE_* flag counted as set for any non-blank value, so
  CLAUDE_CODE_USE_BEDROCK=0 still pinned every AWS_ variable. Claude Code
  2.1.284 reads these flags as set only for 1, true, yes, or on (any case,
  trimmed); the route digest now does the same.
- The providerRouteDestination docstring no longer says every generated
  adapter re-verifies its invocation: Pi and OpenCode pass no route and are
  checked only at plan time.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…drop a dead scan

environment.tsx already imports @ultrafuzz/runtime at load time for the route
digest, yet it still kept copies of ROUTE_ENV_PREFIXES, the sensitive-name
pattern, isCredentialLikeEnvironmentVariableName, and
routeOwnsCredentialLikeEnvironmentVariable, under a comment calling the file
dependency-free. The copies were identical to the runtime exports the
controller uses; the adapter now takes those two functions from the same
import, so the prefix list and the credential pattern have one definition.
Both exports already exist on origin/main, so this adds no new downgrade
requirement beyond the one providerRouteDestination introduced.

smithersExecutionControlFiles scanned each adapter for "ultrafuzz.toml"
without ULTRAFUZZ_CONFIG_PATH right after assertControllerSourceDigest, which
admits only the byte-exact packaged closure. Every packaged file that
mentions ultrafuzz.toml also reads ULTRAFUZZ_CONFIG_PATH, so the scan could
never fire; it is deleted along with its stale "init --force" advice.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
- docs/security.md listed the route digest as covering only route-bearing
  input and said every agent invocation re-verifies it. Any non-credential
  variable with the agent's provider prefix still counts (ANTHROPIC_LOG
  included), and Pi and OpenCode are checked only at plan time. The section
  now lists exactly what the digest covers and which adapters re-check it.
- The breaking CHANGELOG entry said only IDs that included dropped inputs
  change. Every ID derived from a provider config file changes, because the
  config part is now a structured record rather than the file hash; it now
  says so, and that env-only IDs without the dropped inputs are unchanged.
- "Upgrades no longer need init --force" held only for the adapter closure;
  the CHANGELOG and CLI reference now say that, and that an adapter already
  matching its template is left untouched.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Comment thread packages/runtime/src/data-governance.ts
Comment thread packages/runtime/src/data-governance.ts
Every pull request in this batch inserts its entry at the same place in
CHANGELOG.md, so each merge would conflict with the next. The entries are
collected into one changelog update instead.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant