Skip to content

docs: explain OpenRouter prompt-injection guardrail rejections - #1182

Merged
aviggiano merged 3 commits into
mainfrom
claude/w28-openrouter-guardrail-docs
Sep 29, 2026
Merged

aviggiano merged 3 commits into
mainfrom
claude/w28-openrouter-guardrail-docs

Conversation

@aviggiano

@aviggiano aviggiano commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

#1149 reports that OpenRouter rejects requests from the OpenRouter-backed OpenCode adapter with HTTP 403 Request blocked: prompt injection patterns detected. It blames transcript-like examples in Ultrafuzz's own prompt templates and proposes four things: a provider-specific prompt transform, snapshot pinning for that transform, a normalization record, and a live-provider fixture.

Root cause

The 403 comes from OpenRouter's opt-in prompt-injection guardrail. It fires when a guardrail that covers the operator's key uses the Block action. Ultrafuzz's prompt sources and the rendered prompts of a local run match none of OpenRouter's 33 documented exact patterns. That check does not cover OpenRouter's typoglycemia, misspelling and encoding layers, which its docs describe but do not publish as exact rules.

  • What OpenRouter documents (prompt injection, guardrails, errors, guardrail API):

    • Block rejects "the entire request ... with a 403 before it reaches the model". The error message is Request blocked: prompt injection patterns detected.
    • "Every incoming message is scanned", and base64/hex content is decoded and checked for injection keywords. The API's scan_scope for this detector "Defaults to all_messages" and can be set to user_only.
    • When several guardrails apply to one request, "the most restrictive action wins" (block > redact > flag).
  • Ultrafuzz's text does not match the exact patterns. I copied the 33 exact regexes from OpenRouter's documentation on 2026-09-28 and ran them over:

    • Ultrafuzz's prompt sources: .ultrafuzz/prompts, packages/prompts/src and packages/runtime/src/templates, 88 files.
    • The 8 rendered prompt snapshots and 5 dynamic templates from a local run.
    • Smithers 0.35.0's worktree-isolation notice.

    None of them matched.

  • Text Ultrafuzz does not write does match:

    • OpenCode 1.18.18's built-in prompts, checked at tag v1.18.18 in packages/opencode/src/session/prompt/*.txt. default.txt and trinity.txt each match role_delimiter_injection: the line assistant: [runs ls and sees foo.c, bar.c, baz.c] followed by a user: line. SystemPrompt.provider() picks default.txt for any model ID that is not Muse, GPT/o1/o3, gemini-, Claude, Trinity or Kimi, for example DeepSeek, Qwen or GLM. llm/request.ts puts it into the system prompt of every request, and Ultrafuzz never sees it.
    • Ordinary target files. Vyper (balances: HashMap[address, uint256] then user: address) and YAML (roles: [admin] then user: alice) both match role_delimiter_injection. I wrote these snippets myself; I did not find them in a real target.

So rewording Ultrafuzz's prompts cannot prevent this. It is an operator setting.

Correction to the analysis this work item started from. anthropic.txt is the prompt OpenCode uses for openrouter/anthropic/claude-opus-4.8, the example OpenCode profile in docs/config.md (the default config ships no OpenCode model profile). It does not match any documented exact regex: its lines that end in ] are followed by </example>, not by a role line. I do not know what tripped the reporter's request. The issue does not name the model.

Change

Docs plus a CHANGELOG entry. In docs/config.md:

  • A new ## OpenRouter guardrails section, placed after the Pi agent section because it covers OpenRouterAgent, PiAgent, and OpenCodeAgent with openrouter/ models. It:
    • quotes the 403 message and says where it comes from;
    • says the match can come from the harness's own system prompt or from target source and test output, since OpenRouter scans every message by default, using the verified OpenCode 1.18.18 example;
    • tells operators to set prompt-injection detection to Flag, or turn it off, on every guardrail that covers the key: the workspace default, member and API-key guardrails. The most restrictive action wins, so fixing only one of them is not enough. It also tells them not to use Redact, which forwards the request with each match replaced by [PROMPT_INJECTION];
    • says what relaxing costs. The workspace default covers every key in its workspace, and a member guardrail covers every key of that member. A workspace used only for the Ultrafuzz key keeps the workspace-default change to that key. In an organization account, only an organization admin can change guardrails;
    • says Ultrafuzz has no special handling for the rejection. The attempt fails like any other agent error and follows the [retry] policy. A retry on the same profile is rejected again when the match is in the task prompt or in the harness's system prompt.
  • The OpenCode agent section no longer says "The default root config includes an opt-in OpenCode profile". packages/config/defaults.toml and ultrafuzz.toml have [agents.OpenCodeAgent] but no [models.opencode], and git log -G '\[models\.opencode\]' over both files is empty. The text now says to add a model profile. The doctor paragraph's "keeping the shipped [models.opencode] profile" is left alone, because perf(runtime): halve launch fsync work, and doctor stops failing on unused agent profiles #1179 rewrites that paragraph and removes the claim.

I departed from the work-item spec in two places:

  1. Every guardrail, not just the key's. The spec said to set "the guardrail on the key" to Flag. Because OpenRouter applies the most restrictive action across all guardrails that cover a request, a workspace-default Block would stay in force. The doc names all three kinds.
  2. No promise that the fallback runs. The spec said to state "fresh-session retry, then the [retry] agents fallback". On current main that is not what happens when the 403 repeats identically. The agent Task's retry policy (smithers.ts:7917) sets no maxIdenticalFailures, so Smithers 0.35.0 falls back to DEFAULT_MAX_IDENTICAL_FAILURES = 3 (@smthrs/scheduler/src/errorSignature.js). It marks the node stalled on the third failure with the same error signature (makeWorkflowSession.js:878), before any later attempt or fallback profile runs. I read this in the pinned source; I did not run it. The new section points to the [retry] paragraph instead of restating it. That paragraph is itself wrong about identical failures on main, and fix(runtime): give agent retries a real wait, stop retrying deterministic failures, and label timeouts by code #1171 fixes it: it sets maxIdenticalFailures: 0 and states that "repeated identical failures do not end it before later attempts or fallback profiles run". The pointer is accurate only after fix(runtime): give agent retries a real wait, stop retrying deterministic failures, and label timeouts by code #1171 lands, so merge this PR after fix(runtime): give agent retries a real wait, stop retrying deterministic failures, and label timeouts by code #1171.

Deliberately not built (and why)

  • Prompt transform or delimiter rewriting, as the issue proposes: Ultrafuzz's prompts have no matches. The matching text is in OpenCode's system prompt and in tool history, which Ultrafuzz does not own. Rewriting target content would also corrupt code and findings evidence.
  • Snapshot pinning and a normalization record: more sealing and provenance machinery for a transform that should not exist.
  • Live-provider fixture: it would need a paid OpenRouter workspace with a Block guardrail in CI, and it would test OpenRouter's pattern set rather than Ultrafuzz.
  • Classifying the 403 as non-retryable AGENT_CONFIG_INVALID: that contradicts Add bounded error-agnostic agent retries with backoff and optional model fallback #572, which rules out matchers on provider error text. It would also stop the retry chain. Retrying can still succeed when the match came from content the failed session happened to read.
  • A committed test that pins OpenRouter's regex list: OpenRouter changes its patterns over time, so such a test would pin their docs, not Ultrafuzz behavior.
  • Recommending OpenRouter's phrase allowlist instead of Flag: entries are exact substrings, so they can exempt a known harness line but cannot anticipate target source or tool output. An allowlisted phrase also does not exempt a separate encoding or misspelling detector hit.

Verification

There is no behavior change, so there is no discriminating test. A test that pins doc prose would prove nothing.

Run for the first commit:

  • The regex scans described above. The script is in the scratchpad and is not committed.
  • Smithers 0.35.0 classification. In BaseCliAgent, classifyQuotaError and the 8 non-retryable patterns (loaded from the pinned source) match neither the Codex-shaped message (unexpected status 403 Forbidden: {"error":{...}}) nor the bare message. Both therefore become a generic AGENT_CLI_ERROR.
  • Ultrafuzz's failure normalizer. A scratch test through captureAgentFailure in packages/runtime/test/generated-workflow-verifier.test.ts, reverted afterwards. It covered both message shapes, with and without AGENT_CLI_ERROR, on both generate and preflight. Every case came out with no code or details and the provider message intact. In the same run, 402 gateway failures are promoted to Smithers quota parking controls and agent failure normalization preserves only validated Smithers recovery controls passed.
  • OpenCode 1.18.18 session/retry.ts. A 403 with OpenRouter's documented body is not retryable there, and its internal retries are capped at 5, so OpenCode does not loop on it.

Run for the review follow-up commit:

CI is red for a reason outside this PR's content.

Not verified:

  • I did not reproduce this against live OpenRouter; there is no key or Block-mode workspace here. How each harness (Codex, OpenCode, pi) surfaces the 403 comes from reading source, not from watching it happen.
  • OpenRouter does not document whether it scans the Responses API instructions field. That field matters because OpenRouterAgent runs Codex with wire_api = "responses".

Risk / compatibility

Closes #1149

🤖 Generated with Claude Code

RetriggerConfidence Score: 5/5

The PR appears safe to merge based on the reviewed changes.

Summary

The PR documents how OpenRouter prompt-injection guardrails can reject agent requests and clarifies that operators must add an OpenCode model profile.

  • Explains which guardrails operators need to check and the scope of relaxing them.
  • Describes the effect of a rejection on Ultrafuzz retries.

Reviews (3) · Last reviewed commit: "chore: move the changelog entry to the c..."

OpenRouterAgent, PiAgent, and OpenCodeAgent with an openrouter/ model
all use the key in OPENROUTER_API_KEY. When a guardrail that covers that
key sets prompt-injection detection to Block, OpenRouter rejects each
matching request with HTTP 403 "Request blocked: prompt injection
patterns detected" before it reaches a model.

#1149 blamed transcript-like examples in Ultrafuzz's prompts. OpenRouter's
documented exact regexes match none of Ultrafuzz's prompt sources or the
rendered prompts of a local run. They do match text Ultrafuzz does not
write: OpenCode 1.18.18's default system prompt, used for models without
a model-specific prompt (DeepSeek, Qwen, GLM), matches
role_delimiter_injection, and so does ordinary Vyper or YAML source.

Document the operator fix in docs/config.md. Set prompt-injection
detection to Flag, or turn it off, on every guardrail that covers the
key, because OpenRouter applies the most restrictive action across the
workspace default and member or key guardrails. Do not use Redact, which
forwards the request with each match replaced. Runtime behavior is
unchanged: the rejection is an ordinary agent failure under the [retry]
policy, and the doc points there instead of restating it.

Closes #1149

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano requested a review from a team as a code owner September 29, 2026 00:14
Comment thread docs/config.md
aviggiano and others added 2 commits September 29, 2026 03:01
Review follow-up for the OpenRouter guardrails section.

- The section said the OpenCode openrouter/ model came from "the shipped
  profile". No OpenCode model profile ships: packages/config/defaults.toml
  and ultrafuzz.toml carry [agents.OpenCodeAgent] but no [models.opencode].
  Drop the parenthetical, and correct the OpenCode agent section's "The
  default root config includes an opt-in OpenCode profile", the claim it
  repeated. The doctor paragraph's copy of the claim is left to #1179,
  which rewrites that paragraph.
- OpenRouter's scan_scope for the prompt-injection builtin defaults to
  all_messages and can be set to user_only, so say that every message is
  scanned by default rather than always.
- Say that the workspace default and member guardrails also cover other
  keys, that a workspace used only for the Ultrafuzz key confines the
  workspace-default change to that key, and that in an organization
  account only an organization admin can change guardrails.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Every pull request in this batch inserts its entry at the same place in
CHANGELOG.md, so each merge would conflict with the next. The entries are
collected into one changelog update instead.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@aviggiano
aviggiano merged commit 35b2a05 into main Sep 29, 2026
13 checks passed
@aviggiano
aviggiano deleted the claude/w28-openrouter-guardrail-docs branch September 29, 2026 06:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Normalize OpenRouter prompts to avoid false prompt-injection rejection

1 participant