Skip to content

feat: routing subsystem, Colab runtime backend, config audit - #35

Merged
araray merged 21 commits into
mainfrom
av/routing_impl
Oct 2, 2026
Merged

araray merged 21 commits into
mainfrom
av/routing_impl

Conversation

@araray

@araray araray commented Oct 1, 2026

Copy link
Copy Markdown
Owner

Implements the routing subsystem (T1–T10) and the Colab runtime backend
(R2–R6), plus the config audit that was asked for alongside them.

Supersedes #34 — that branch is merged in here, so the spec and its
implementation land together whichever order they are reviewed in.

Routing

Five layers that compose, each usable alone. Nothing is on by default:
with no [routing] section every call resolves exactly as it did before.

Config stops being an allow-list. A provider registered in PROVIDER_MAP,
with its key in the environment and its base URL already known, was still
unreachable. Any provider+model is now addressable by spec string:

await llm.chat("hi", target="xai:grok-4.1-20251117?effort=high")
await llm.chat("hi", target="ollama:llama3.3:70b")        # colons in model names
await llm.chat("hi", target="vllm:Qwen/Qwen3-30B#my-box")  # '#' pins an instance

Found on the way: xai, groq and together had no [providers.*] section
at all, so all three were unreachable despite being registered.

Pools, with a failure taxonomy rather than one retry rule. A 429 cools
down briefly; an empty wallet cools down for minutes and records a balance of
zero; a 5xx retries once in place before moving, since it is usually one bad
node and moving would discard a warm prompt cache; a bad key benches the
target for the process; a 400 does not fail over at all; a refusal does not
either, unless you opt in — retrying a refusal elsewhere is "shop until
someone says yes". A prompt that overflows the context window is treated as a
routing signal, not an error.

Seven selection strategies. Three of them needed a judgement call about
missing data, and in each case treating "unknown" as a number inverts the
strategy: an unpriced target ranks after every priced one; an unknown
balance ranks after a known one but ahead of a known-empty one;
lowest_latency explores an unmeasured target first.

Lanes and classifiers. A classifier names a lane, never a model. Nine
classifiers, ordered cheapest-first and then instructions-before-guesses —
both halves matter, because hint and heuristic are both free and running
the guess first would override what the caller asked for. Authority has four
levels, since a routing marker found in content may have arrived from a
retrieved document, where [[lane:deep]] would be a one-line prompt
injection.

Cascades (opt-in) and the privacy path, where the documentation says
plainly that redaction is a mitigation and routing is the guarantee.

Proxy mode: llmcore-bridge proxy gives an unmodified agent harness all of
the above through two environment variables. It refuses a non-loopback bind
without a bearer token, because the process holds every provider credential in
the config.

Colab runtime backend

Sizing is read-only and free and prints its own arithmetic. Every unknown
rounds toward needing more: overestimating buys a bigger GPU, underestimating
OOMs on the VM after billing has started.

The backend enforces four rules in code — state before compute, fail closed,
never connect to a session that does not exist, deadlines at creation — and
llmcore-runtimes exists for the commands you need when something has gone
wrong.

Bugs found and fixed

  • Gemini had no pricing or context window, silently. The card alias map ran
    the wrong way: provider type gemini mapped to a gemini namespace, but
    cards are filed under google/. Nothing raised — the lookup returned None,
    None means unknown, and every caller handles unknown quietly. So
    lowest_cost could not price any Gemini target. There is now an audit test.
  • Anthropic rejected thinking.budget_tokens with a 400 on 4.6+ while
    pre-4.6 requires it. effort is now generation-agnostic.
  • The test suite wrote into the real ~/.llmcore/runtimes, leaving a
    phantom entry claiming a GPU VM was running — the exact false signal that
    directory exists to prevent.
  • Tests could load the repo's .env and reach real vendors.
  • A 404 was classified unknown and retried the same target; it has its own
    kind now, because a missing model is deterministic here but a peer serves a
    different model.
  • prompt_tokens came back 0 after a failover; the proxy called a method that
    does not exist; a three-group phone number was only partly redacted.

Where I was wrong, corrected in place

The spec claimed local classification adds "<50 ms p50 on CPU". Measured:
191 ms for 2 lanes, 246 for 5, 314 for 9 — about 5x out, which is why the
gate said measured, not assumed. The spec now says so, and the encoder is
documented as a batch/agent feature.

The same model's raw confidence is meaningless without reading it against
chance (0.20 across five lanes is exactly uniform), so confidence is now
chance-corrected.

Config audit

All 660 keys checked. All 48 [routing] and 14 [runtimes] keys are wired.
Two pre-existing keys promise behaviour that does not exist and are now marked
NOT ENFORCED rather than quietly removed:

  • llmcore.admin_api_key claimed to protect admin endpoints "such as live
    configuration reloading" — nothing reads it, so anyone who set it believing
    their reload endpoint was guarded was wrong.
  • context_management.minimum_history_messages.

What is not done

  • The Colab end-to-end gate is unmet. "One real model served end to end"
    means provisioning a real GPU VM and spending your compute units — not mine
    to authorize. Everything short of spending is tested.
  • The Anthropic fix is unverified against a live 4.6+ model: the account
    still answers "Your credit balance is too low".
  • cache gc lists but does not delete; the agent-lens migration guide is not
    written, since it would tell another project to depend on an unproven path.
  • No accuracy figure is claimed for any classifier. None is validated on
    real traffic, and any number would be invented. That is the largest honest
    gap in the subsystem.

Validation

Full suite: 6564 passed, 87 skipped. The only exclusions need a Postgres
server and fail the same way on main.

Live against real vendors: failover over a nonexistent model, lowest_cost
picking the cheaper of two vendors from card pricing, a 700k-char prompt
skipping a 128k window for a 1M one, autoprovisioning an unconfigured
provider, the proxy driven by the real openai SDK, and sizing against the
live Hugging Face Hub.

🤖 Generated with Claude Code

araray and others added 16 commits October 1, 2026 01:18
Two things: a config audit with one real fix, and the design document for the
routing programme, which is awaiting approval before any implementation.

The audit compared default_config.toml against the provider registry and the
config keys the code actually reads. One real gap: xai, groq and together are
registered in PROVIDER_MAP and have their env var and base_url in
_OPENAI_COMPATIBLE_DEFAULTS, but a provider is only instantiated when its
[providers.*] section exists — so all three raised "not configured" and were
unreachable. Sections added; coverage is now 23/23 and all three resolve with
the right base URLs. Three other apparent gaps (media.jobs.webhook_secret,
media.routing, observability.cost_tracking) were false positives in my detector:
two are documented as comments and one is a declared table.

That gap is also the root of the dynamic-target complaint, and the spec treats
it as a design flaw rather than a missing feature: config is acting as an
allow-list when it should be a preset layer. Adding the three sections only
moves the wall.

The spec separates what was asked as one feature into five composable layers,
because specifying them as one would do every job badly: Target (which
provider/model/params, addressable as a spec string, autoprovisioned when absent
from config), Pool (failover and prioritisation), Lane + Classifier (what kind
of request is this), Cascade (was the cheap answer good enough), and Transform
(what must not leave this machine). Each is independently useful.

It proposes Pool and Lane as two names rather than one "group", because the
request used one word for two ideas: a pool is a set of interchangeable targets
for failover, a lane is a named destination chosen by a classifier. Collapsing
them forces every group to be both.

Grounded in prior art rather than invented: LiteLLM's router for strategies,
cooldowns and error-specific retry; Arch-Router for the key insight that one
should route to named lanes with the lane-to-model binding kept separate, so
swapping models needs no retraining; RouteLLM for strong/weak routing;
FrugalGPT for the cascade. Real downloadable local classifiers exist and are
named, including a 350M CPU encoder that scores a prompt against free-text lanes
zero-shot, and a family of 307M task encoders covering domain, PII and
factuality.

Four contributions beyond transcription are flagged as such: the failure
taxonomy distinguishes 429 from insufficient-credit from context-length, and
treats ContextLengthError as a routing signal toward a larger window rather than
an error; content-policy refusals do not fail over by default, because retrying
a refusal elsewhere is a decision a user must opt into; session affinity is
specified because mid-conversation failover loses prompt caching and shifts
style, which the request did not raise; and the privacy path constrains the pool
to local targets rather than relying on redaction, with the documentation
required to say that redaction alone is not a guarantee.

Balance-based prioritisation is specified honestly: most vendors publish no
balance endpoint, so the protocol returns None for unknown, unknown never ranks
as zero, and an observed insufficient-credit error is what supplies the signal
where no endpoint exists.

No accuracy or saving numbers appear anywhere, because none have been measured
on this library's traffic. An evaluation harness ships with the classifier phase
instead.

Five open questions are listed for decision, including naming, the default pool
strategy, and the scope of the first PR.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
First slice of the routing subsystem (T1 of docs/ROUTING_SUBSYSTEM_SPEC.md).

`Target` parses the spec strings routing is addressed with --
`provider:model?params#instance`. The split is on the FIRST colon, because
model names legitimately contain colons and slashes (`ollama:llama3.3:70b`,
`vllm:Qwen/Qwen3-30B`); splitting on the last one would mangle them.
`Target.key` deliberately excludes params, since health and rate limits
belong to the endpoint, not to a requested effort level.

`ProviderManager.resolve_target()` makes config a set of presets rather
than an allow-list: an unconfigured `provider:model` is built on demand
from a discovered credential and marked ephemeral. An explicit `#instance`
pin is never autoprovisioned -- naming something that does not exist is an
error, not a request to build something adjacent.

`RoutingSettings` resolves config -> env -> request, and the request wins.
Policy lives here, not in the failure enum: `FailureKind` says what a
failure *is* (`failover_is_pointless`, `affects_health`, `retry_same`) and
settings say what to do about it (`on_refusal`). Billing is classified
before auth, so a valid key on an empty account gets a 10-minute cooldown
instead of being benched for the process.

Cascade is off by default -- it costs an extra call.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The implementation PR should carry the spec it implements, so the two land
together whichever order the PRs merge in.
T2 and T3 of the spec. A pool is a set of interchangeable targets; selection
is split into eligibility (facts) and ordering (strategy), and every member
is reported with the reason it was or was not used -- "why did it pick the
slow one?" is only answerable if the record says the fast one had twelve
seconds of 429 cooldown left.

All seven strategies land together because they share that scaffolding:
priority, round_robin, weighted, lowest_latency, lowest_cost, least_busy,
most_credits. Three of them needed a judgement call about missing data, and
in each case treating "unknown" as a number would have inverted the strategy:

* `lowest_cost` ranks unpriced targets *after* priced ones. Scoring an
  unknown price as zero would make a model with no card beat a model known
  to be free.
* `most_credits` uses three bands -- known funds, unknown, known-empty. The
  obvious descending sort puts a confirmed zero ahead of an unknown, which
  is backwards: an unknown account might have money.
* `lowest_latency` explores an unmeasured target first, since it cannot
  prefer low latency without a measurement and one call buys one.

Self-hosted providers (ollama, vllm) are priced at zero rather than unknown.
Their cards carry no pricing block, and reading that absence as unknown would
rank your own GPU behind a paid API instead of ahead of it.

Order tiers (`?order=`) gate eligibility across strategies, which is how
"use my own GPU, and only pay a vendor if it is down" is expressed. Routing
params are stripped from the target before it reaches a provider, since
forwarding `weight` to a vendor is a 400.

Also here: a pre-call context-window check that skips a target whose card
already proves the prompt will not fit; `InMemoryRoutingState` with an
injectable clock and context-managed in-flight counters; and routing events
on the existing shared_events spine, guarded by `has_sinks()` so routing
costs nothing when nobody is listening.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
T5 and T6. A classifier names a *lane* and never a model, so swapping a
model is a config edit rather than a classifier change -- the one idea worth
borrowing wholesale from Arch-Router. That indirection is also what collapses
the original request's complexity tiers, speed tiers, domain routing and
privacy class into a single mechanism: "focus groups per complexity level"
are lanes whose names happen to be complexity levels, and nothing in
lanes.py knows that.

Chain ordering is cheapest-first, then instructions before guesses, both
declared per classifier rather than inferred from config order. The second
half of that rule is not a nicety: `hint` and `heuristic` are both free, so
cost alone cannot separate them, and running the guess first would silently
override what the caller asked for.

Authority has four levels rather than two, because a routing marker found in
*content* is not as trustworthy as an argument on the call. Markers are the
point of the feature -- in an agent harness the model's text is the only
channel that passes through, so that is how an agent routes itself -- but in
any RAG or tool-output path that text may have come from a retrieved
document, where `[[lane:deep]]` would be a one-line prompt injection, or
worse, a way out of the private lane. So `caller` > `policy` > `prompt` >
`inferred`, and a marker is honoured only when nobody more authoritative
spoke.

Two things measurement changed, both now corrected in the spec:

* The T6 gate claimed local classification adds "<50 ms p50 on CPU". Real
  figure on 8 CPU threads: 191 ms for 2 lanes, 246 ms for 5, 314 ms for 9.
  Wrong by ~5x, which is exactly why the gate said measured-not-assumed.
  The encoder is off by default, runs its forward pass in a worker thread so
  it cannot stall the event loop, and is now documented as a batch/agent
  feature rather than an interactive one.
* The encoder's raw top score is meaningless without reading it against
  chance -- 0.20 over five lanes is exactly uniform, i.e. no opinion, and
  taking it as 20% confidence would route on noise. Confidence is now
  reported chance-corrected, so one floor means the same thing at any lane
  count, and prompts matching no lane correctly abstain.

The heuristic lost its short-prompt rule for the same reason: length does not
predict complexity. "Write a 2000-word essay on X" is ten tokens, and so is
"tell me about the Azores in autumn". Length now only ever vetoes the trivial
lane; selecting it needs a recognisable simple-task verb.

`RoutingRequest.prompt` is optional, since `chat(messages=[...])` is a real
call shape with no prompt.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The piece where the five layers compose, plus the two layers they needed:
prompt transforms (T8) and response verifiers (T7).

Order matters and is fixed: classify, resolve, transform, attempt, cascade.
Transforms run *after* target selection, because a transform's whole value is
that it sees where the prompt is about to go -- the cost is that a
`constrain` action re-resolves against a different pool, which the manager
does explicitly and once.

On the privacy layer, the documentation says plainly what the feature is and
is not: **redaction is a mitigation, routing is the guarantee.** A detector
that misses one identifier has leaked it, and no detector catches
everything, so `constrain` changes the *destination* rather than the text.
Three places where the obvious implementation would have quietly undone
that, all now closed:

* `constrain` with no `pool` configured would degrade to "send it anyway",
  which is the one outcome the user was preventing. It blocks instead.
* a transform that raises would look like "nothing found". The chain fails
  closed by default.
* a constrain pool that does not exist would fall through to the remote
  target. It blocks.

Findings carry a 16-char hash, never the value, including in
`PromptBlockedError` -- a privacy feature whose exception text contained the
identifiers it found would be the leak it exists to prevent.

Verdicts are three-valued, and `None` means *could not judge*. Folding it
either way is the bug that makes cascades useless: read as a fail it
escalates every unjudgeable answer and inverts the saving, read as a pass it
silently disables the quality floor the moment the judge breaks. So it stays
a third value and `routing.cascade.on_unknown` decides.

Two bugs the end-to-end run found:

* `RoutingSettings.magic_pattern` defaulted to a regex with one *unnamed*
  group, and the manager handed it to a classifier that reads named groups
  `key` and `value` -- an IndexError on every classified request. The
  default is now None (use the built-in) and a custom pattern is validated
  at build time with a message that shows the expected shape.
* A three-group phone number was only partially redacted, leaving the last
  digits behind; and `192.168.1.42` matches both the ipv4 and phone
  patterns, so redacting both rewrote the same span twice. Hits are now
  de-overlapped, longest match first.

Also fixed, and more important than it looks: the test fixture built confy
`Config` objects with its default `load_dotenv_file=True`, which exports the
repo's `.env` into `os.environ`. That made a credential-discovery assertion
depend on the developer's machine, and it meant a test could reach a real
vendor and spend real money. Fixtures now pass `load_dotenv_file=False`.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
`[routing]` in default_config.toml, `llm.chat(target=/pool=/lane=/profile=/
complexity=/effort=/routing=)`, the `llm.routing` facade, and proxy mode on
the bridge (T9).

The proxy exists because of a constraint worth stating: an agent harness owns
its own API call, so llmcore cannot add a `lane=` argument to it -- but it
almost always lets you set a base URL and a model name. So those two strings
are the entire interface. `model="lane:deep"`, `"pool:main"`, a target spec,
a bare model name or `"auto"` all work, and a harness gets pools, failover,
classifiers, cascades, transforms and cost accounting without knowing llmcore
exists.

Two things the proxy refuses to hide. Usage reports the target that *actually*
served the request in an `llmcore` block, because under a pool the answering
model is genuinely not the one that was asked for and a harness logging spend
per model should not be lied to. And the bind refuses: this process holds
every provider credential in the config, so a non-loopback host without a
bearer token raises at construction rather than starting.

Live validation against real vendors found three real bugs:

* **Gemini had no pricing and no context window, silently.** The card alias
  map ran the wrong way -- provider type `gemini` was mapped *to* a `gemini`
  card namespace, but cards are filed under `google/`. Nothing raised: the
  lookup returned None, None means "unknown", and every caller handles
  unknown quietly. So `lowest_cost` could not price any Gemini target and the
  pre-call window check never fired, for one of the most used providers in the
  library. tests/routing/test_cards.py now audits every provider type against
  the packaged card tree so the next provider cannot reintroduce it.
* `_call_with_failover` passed `chat()`'s keyword names to
  `prepare_context()`, which has different ones -- so the first failover
  raised instead of failing over.
* A 404 was classified `unknown`, which retried the same target before moving
  on. A missing model is deterministic for that target but a peer serves a
  *different model*, so it is the one case where "permanent" and "try someone
  else" are both true. It now has its own `MODEL_NOT_FOUND` kind, checked
  before the 400 branch because several vendors answer 400 for it.

Also corrected a regression I introduced: routing made unsupported provider
kwargs drop with a warning everywhere, which silently broke the existing
contract that `chat("x", bogus=1)` raises. Dropping now applies only when a
*pool* chose the target -- where members genuinely have different parameter
surfaces and the caller cannot know which one serves -- and raises everywhere
else.

Verified live: failover over a nonexistent model, `lowest_cost` picking the
cheaper of two real vendors from card pricing, a 700k-char prompt skipping a
128k window for a 1M one, and an unconfigured provider reached by spec string.

Full suite: 6358 passed, 87 skipped. The only failures left need a Postgres
server (`PoolTimeout` after 30s) and fail the same way on main.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
`llmcore-bridge proxy` starts the OpenAI-compatible surface. A separate
subcommand rather than another transport on `serve`, because it is a
different thing: `serve` exposes llmcore's own contract to llmcore clients,
while this impersonates somebody else's API so an unmodified harness can be
pointed at it. A refused bind prints the reason instead of a traceback.

Validated against the real `openai` Python SDK pointed at the proxy:
`/v1/models` listing lanes and pools, pool routing with failover over a
broken member, lane routing through the model name, an explicit target spec,
streaming, and a bogus model surfacing as the SDK's own `NotFoundError`.

Three bugs the tests and that run found:

* The proxy called `get_last_interaction_info()`, which does not exist -- the
  real method is `get_last_interaction_context_info()`. My fake had defined
  the wrong name, so the unit tests passed while the usage block would have
  been empty against a real LLMCore.
* Usage was empty whenever a harness sent no session, since llmcore keys its
  per-turn introspection by session id. That defeats the endpoint's one
  promise -- to say which target actually answered. The proxy now uses a
  synthetic session, and `LLMCore.discard_transient_state()` drops it
  afterwards; without that, a proxy handling one synthetic session per
  request leaks a cache entry per call for the life of the process.
* `prompt_tokens` came back 0 after a failover. `chat()` derives it *after*
  `prepare_context()` returns, so copying the re-prepared details over the
  originals overwrote it with the unset value. Excluded and recomputed.

Error mapping now runs every exception through `classify_failure` rather than
only `ProviderError`, so a bare `TimeoutError` is a 504 instead of a 500 --
harnesses branch on these statuses, and getting them right is what makes the
proxy behave like the thing it is impersonating.

Routing and the proxy are ruff-clean; api.py's finding count went down.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
`docs/Routing_usage.md` is the how-to; the spec stays the why. Its last
section is "What this does not claim", which lists the limits plainly: no
accuracy figure for any classifier (none is validated on anyone's traffic, so
any number would be invented), cost estimates are for comparing targets rather
than for billing, PII redaction is not a leak guarantee, the local classifier
is not free in latency, routing state is per process, and `most_credits` only
works where a vendor exposes a balance.

Three skills in the bundled grimoire pack, so an agent can pick them up:
`skills/llmcore/proxy` (point a harness at llmcore), `skills/llmcore/routing`
(configure and debug it), `skills/llmcore/cost` (measure, then reduce).
`skilldoc_paths` added to the pack manifest; the pack now loads 21 spells,
3 runes, 7 promptlets and 3 skilldocs.

The spec is marked implemented, its five open questions are answered in a new
§13, and §13.1 records the six places the implementation forced a change on
the design -- because a spec that quietly diverges from its code is worse than
no spec. Three honest gaps are left open, the largest being that no classifier
has been validated against real traffic.

Every `provider:model` string in the new docs was checked against the bundled
model cards; all resolve, except `vllm:Qwen/Qwen3-30B`, which is correct since
vLLM is self-hosted and ships no cards by design.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
…t context

Read-only and free: nothing in `sizing.py` can start billable compute, which
is why sizing is a separate phase. Every `Plan` carries its own arithmetic,
because a sizer that answers "use an A100" and shows nothing is impossible to
argue with.

Every unknown rounds toward *needing more*, deliberately. Overestimating buys
a bigger GPU; underestimating OOMs on the VM after billing has started.

Three things the real Hub and the real CLI corrected in the draft design:

* **`A100-40` and `A100-80` are not SKUs you can ask for.** `colab new --gpu`
  accepts T4, L4, G4, H100, A100 and nothing else, so a plan naming A100-80
  could never have been provisioned. The ladder now uses the CLI's own names,
  A100 is sized at the conservative 40 GB (being handed an 80 GB card only
  adds headroom), and the draft spellings are kept as aliases so a config
  written against the spec still works. A test asserts every SKU's `--gpu`
  value is one the CLI accepts.
* **`Quantization` could not express GGUF precision.** Q4 and Q8 differ by 2x
  in weight bytes, which is routinely the difference between fitting a 24 GB
  card and not, so a single `GGUF` member forced the sizer to guess. Split
  into `GGUF_Q4/Q5/Q8`, plus `INT4`.
* **The KV fallback produced nonsense.** For a gated repo whose config.json
  is unreadable, the draft formula estimated 0.5 GB for a 70B model at 8k
  context -- about 5x under. It now scales from the parameter count at 8
  KB/token/B, calibrated against models where the real figures *are*
  available (3.2 / 4.6 / 7.5 KB measured for Qwen3-30B, Llama-3.3-70B,
  Qwen2.5-7B) and taking the high end on purpose.

The GQA term is read explicitly from `num_key_value_heads`, since falling back
to `num_attention_heads` when the key is present overestimates
Llama-3.3-70B's cache by 8x. The KV cache is priced at 16 bits even under
4-bit weights, because that is what vLLM actually does.

Context shrinks before a bigger GPU is chosen -- a bigger GPU costs money, a
shorter context costs nothing -- with a floor at 4k, since silently handing
someone a 1k window is worse than saying it does not fit. When nothing fits,
the sizer refuses *with a concrete alternative*, because a bare refusal sends
the caller to guess and the next guess is also a launch.

`vram_available_gb` reports the budget the fit decision actually compared
against, not the sticker VRAM, so a plan cannot look roomier than the sizer
believed.

Validated against the live Hub: Qwen2.5-7B, Qwen3-30B (dense and AWQ),
Llama-3.3-70B (gated, so exercising the fallback) and gemma-3-4b. 45 tests,
one of them marked `live` against the real Hub because its metadata shape is
the thing most likely to change underneath this.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The repo uses ASCII in strings; ruff flags the multiplication sign as
ambiguous. The arithmetic notes read the same either way.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Provisions a Colab GPU VM, serves an open-weights model on it, tunnels the
endpoint to localhost, and attaches it as an ordinary provider. Plus `bake`
and the Drive cache inventory (R5).

This is the module that spends money, so four rules are enforced in code
rather than left to the caller, and each has a test named after it:

1. **State before compute.** The handle is on disk before `colab new` runs.
   The dangerous window is a crash between assignment and bookkeeping -- a VM
   billing with nothing aware of it.
2. **Fail closed.** One error path, and it ends in `down`. If the release
   *also* fails, the log says MAY STILL BE BILLING in those words, because
   that is the one case a user has to act on.
3. **Never SSH into the void.** The session must appear in `colab sessions`
   before anything connects. Without the guard a failed assignment becomes a
   confusing SSH timeout instead of "no GPU available".
4. **Bounded at creation.** Idle *and* hard deadlines, since an idle reaper
   does not stop a runtime that is busy in a loop.

Everything that leaves the process goes through one seam, so there is a single
place where a command is logged, timed out and captured -- and so the tests
can drive the backend by asserting on the argv it would have run.

Checked against the real CLI (`google-colab-cli` 0.7.4) rather than the
spec's description of it, which was wrong in one place: `colab new --gpu`
accepts T4, L4, G4, H100, A100 and nothing else, so the spec's `A100-80` rung
could never have been provisioned. A test now asserts every SKU's `--gpu`
value is one the CLI accepts.

Secrets go on stdin or through `colab exec --env`, never in argv -- argv is
visible in the VM's process list and in llmcore's own debug logs. Two tests
assert the token appears exactly once across all argv and never inside the
generated bootstrap script.

The VM-side server is started with `setsid`: as a child of the kernel, a
kernel restart would kill the server silently while the VM kept billing.
Cached artefacts in Drive are gated on a sentinel written *after* the
artefact is complete, because a miss costs time while a corrupt hit costs a
debugging session on a billing VM. The bootstrap also waits for readiness on
the VM, so a startup failure is reported with the log that explains it rather
than as a local connection timeout.

Orphan detection is a safety feature: a Colab session llmcore has no record
of is reported with the `adopt` command that makes it killable. An adopted
runtime is marked DEGRADED and says it does not know what it is running --
claiming READY would be worse than admitting the gap.

`RuntimePhase.DETACHED` added, and it counts as billing. Detached means
llmcore stopped watching, not that the VM stopped; counting it as
not-billing would hide exactly the leak these rules exist to prevent.

`_parse_sessions` is deliberately forgiving and falls back to scraping names,
because it is how llmcore *finds things that are costing money* -- a parser
that raised on a cosmetic CLI change would break teardown precisely when it
is needed.

58 tests. **The spec's R3 gate -- one real model served end to end -- is not
met**, and deliberately: meeting it means provisioning a real GPU VM and
spending the account's compute units, which is not mine to authorize.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
**R4.** `reap()` existed but nothing called it, which made the idle and hard
deadlines documentation rather than limits. `up()` now starts a supervisor
that probes liveness and reaps on a timer -- started automatically, because a
caller who forgot would discover it from a bill.

Liveness marks a target DEGRADED after three consecutive failures, not one: a
single missed poll is usually the tunnel reconnecting. A failed probe never
tears a runtime down. A degraded runtime is still assigned and still billing,
so releasing it belongs to the deadlines or to the user -- letting a transient
network problem destroy an expensive VM would be worse than the problem.

The idle reaper needs to know when a runtime was last used and vLLM exposes
no last-request metric, so the count happens on llmcore's side: `attach()`
wraps the provider's `chat_completion`, which is exactly one call per real
request. Hooking `get_provider()` would miss a long stream; hooking the tunnel
would need a proxy. The wrap is idempotent, since a runtime may be re-attached
after a reconnect.

**R6.** `llmcore-runtimes` — estimate, up, status, down, logs, adopt, bake,
cache. A CLI because the commands that matter most are the ones needed when
something has gone wrong: `status` and `down` have to work from a shell, in a
hurry, possibly from a different process than the one that started the
runtime. `up` requires `--yes` and prints the burn rate first; `estimate` is
free and works while the subsystem is disabled, because deciding whether to
spend should not require enabling spend.

**Fixed: the test suite was writing into the user's real
`~/.llmcore/runtimes`.** That directory is the record of what is currently
costing money -- the spec calls it a safety mechanism rather than a cache, and
`llmcore-runtimes status` reads it. Running the tests left a phantom entry
claiming a READY L4 VM was running, which is exactly the false signal the
subsystem exists to prevent: someone checking whether they were being billed
would have been told yes, by their own test run. An autouse fixture now
redirects the default state directory per test.

`RuntimePhase.DETACHED`, `runtimes.defaults.supervise_seconds` and
`runtimes.colab.keepalive_seconds` added to the config, with the SKU ladder
corrected to the names the CLI accepts.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
`docs/Runtimes_usage.md` with a "What this does not claim" section that leads
with the unmet gate rather than burying it: nobody has run this against a real
Colab VM, because that spends real compute units.

The spec's phase table now marks R2 and R4 met, and R3/R5/R6 implemented with
their gates explicitly unmet and the reason given. `cache gc` is noted as not
implemented. The agent-lens migration guide is not written, because writing it
would mean telling another project to depend on an unproven path.

**Config audit** (the other half of the request): all 660 keys in
default_config.toml checked against the code. All 48 `[routing]` keys and all
14 `[runtimes]` keys are wired. The 116 unread `semantiscan.*` keys belong to
that package. Two keys promise behaviour that does not exist and are now
marked NOT ENFORCED in the file rather than silently removed:

* `llmcore.admin_api_key` claimed to protect "administrative endpoints such as
  live configuration reloading". Nothing reads it, and the bridge's
  ReloadConfig has no reference to it -- so a deployment that set this
  believing its reload endpoint was guarded was wrong. The comment now says
  so, and points at the transport-level auth that does work.
* `context_management.minimum_history_messages` -- truncation does not honour
  it.

Both kept rather than deleted so existing configs keep loading. A security
control that silently does nothing is worse than one that is absent: someone
has to be able to find out.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
`thinking.budget_tokens` is deprecated on Claude 4.6 and **rejected with a
400** on 5.x, while pre-4.6 models require it. llmcore forwarded whatever the
caller passed, so a reasonable request became an API error purely because of
which model served it. Recorded in PROVIDER_MODERNIZATION_PLAN.md and flagged
again by the routing spec -- pools make this routine, since members span
generations.

`effort` is now a first-class Anthropic parameter and means the same thing on
either generation: 4.6+ gets `{"type": "adaptive"}` plus
`output_config.effort`, earlier models get a token budget. A caller's
`budget_tokens` to a 5.x model is dropped with a warning naming the 400 it
would have caused, and a request for `adaptive` on an older model is converted
rather than failing.

Two judgement calls worth stating. An unparseable model id is assumed modern,
because the pre-4.6 family is the shrinking set and defaulting the other way
would break every new model by default. And `effort="none"` disables thinking
rather than requesting adaptive-at-low, which would quietly spend reasoning
tokens the caller explicitly asked not to spend.

37 tests. The live gate is unmet for a reason outside the code: the Anthropic
account available has no credit balance, so nothing could be sent to a real
4.6+ model to confirm the 400 is gone.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Copilot AI balanced review requested due to automatic review settings October 1, 2026 07:27

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Privacy transforms are bypassed or ineffective in production paths, and runtime supervision does not reliably enforce billing deadlines.

Review effort: Balanced
Findings: 7 High severity · 3 Medium severity · 1 Low severity

Open (11)
What changed in this PR

Adds a comprehensive routing subsystem, Colab runtime backend, operational CLIs, provider fixes, documentation, and tests.

Changes:

  • Introduces targets, pools, lanes, classifiers, transforms, cascades, and proxy mode.
  • Implements Colab runtime sizing, lifecycle supervision, and CLI operations.
  • Fixes model-card resolution and Anthropic thinking compatibility.
File Description
src/​llmcore/​api.py Integrates routing into chat calls.
src/​llmcore/​routing/​manager.py Orchestrates routing and cascades.
src/​llmcore/​routing/​* Adds routing models, settings, state, events, lanes, cards, protocols, classifiers, transforms, and verifiers.
src/​llmcore/​runtimes/​* Adds runtime lifecycle models, supervision, and CLI.
src/​llmcore/​providers/​manager.py Adds dynamic provider resolution.
src/​llmcore/​providers/​anthropic_provider.py Normalizes thinking across Claude generations.
src/​llmcore/​bridge/​cli.py Adds proxy CLI mode.
src/​llmcore/​exceptions.py Adds routing exceptions.
src/​llmcore/​grimoire_pack/​* Adds routing, proxy, and cost skill documentation.
tests/​routing/​* Covers routing configuration and card lookup.
tests/​runtimes/​* Covers runtime isolation and supervision.
tests/​providers/​test_anthropic_thinking.py Covers thinking normalization.
README.md Documents routing and Colab runtimes.
docs/​Runtimes_usage.md Adds the runtime usage guide.
docs/​COLAB_RUNTIME_SPEC.md Updates implementation status.
CHANGELOG.md Records the new features and fixes.
pyproject.toml Registers the runtime CLI and live-test marker.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread src/llmcore/api.py
Comment on lines +1863 to +1866
request = self._build_routing_request(
(getattr(chat_session, "id", None) and "") or "",
session_id=getattr(chat_session, "id", None),
)
Comment thread src/llmcore/bridge/cli.py
Comment on lines +316 to +320
for key, value in config.harness_env().items():
# Printed so the operator can paste it straight into the harness.
# The token is the one they configured, so echoing it here tells
# them nothing they did not already set.
logger.info(" %s=%s", key, value)
device=str(encoder.get("device") or "cpu"),
min_confidence=float(encoder.get("min_confidence", 0.35)),
max_chars=int(encoder.get("max_chars", DEFAULT_MAX_CHARS)),
trust_remote_code=bool(encoder.get("trust_remote_code", True)),
Comment on lines +10 to +15
The protocols here, in the order a request meets them:

1. :class:`RequestClassifier` — "what kind of request is this?"
2. :class:`PromptTransform` — "what must not leave this machine?"
3. :class:`BalanceProbe` — "how much is left on this account?"
4. :class:`ResponseVerifier` — "was the cheap answer good enough?"
Comment on lines +136 to +140
async def apply(
self, request: RoutingRequest, target: Target
) -> tuple[RoutingRequest, tuple[TransformResult, ...]]:
"""Run the chain and return the final request plus every result.

Comment on lines +126 to +127
finally:
await llm.close()
Comment on lines +774 to +778
# Re-plan against the constrained pool. This does not count as
# an attempt: nothing was called.
target = None
lane = None
continue
Comment on lines +993 to +1002
# Escalate. The rung we were already served by is skipped, so a cascade
# whose first rung is also the default pool does not pay twice for the
# same answer.
max_rungs = max(1, min(settings.cascade_max_rungs, len(rungs)))
used = result.rungs_used
for rung in rungs:
if used >= max_rungs + 1:
break
try:
escalated = await self.execute(
Comment on lines +553 to +557
seconds = float(
interval
if interval is not None
else self._get("runtimes.defaults.supervise_seconds", 60.0) or 60.0
)
Comment thread README.md
Comment on lines +324 to +327
A detector that misses one identifier has leaked it, so the guarantee is the
*route*: a prompt with personal data in it goes to a pool that never leaves the
machine. Redaction is stacked on top and documented as a mitigation rather than
a guarantee.
araray and others added 3 commits October 1, 2026 08:47
llmcore claims no accuracy for any classifier, because none has been
validated on real traffic and any number would be invented. This is the
machinery for replacing that gap with a measurement, and it is built around
one property a single accuracy figure destroys:

    The two error directions are not interchangeable.

Routing too cheap produces a bad answer. Routing too expensive only costs
money. The same 67% can be usable or unusable depending on which way it errs,
so every report separates them and labels the consequence. There is a test
asserting that two classifiers with identical 0% accuracy produce opposite
reports, because that is the whole reason the harness exists.

Two other things it refuses to do: abstentions are not scored as errors (a
classifier abstaining is passing the turn down the chain, which is the
designed behaviour -- scoring it as wrong would make the most honest
classifier look like the worst), and with no lane order supplied it reports no
direction rather than guessing one.

`llmcore-routing` -- `why`, `health`, `show`, `eval`. Plus
`docs/examples/lane_eval_starter.jsonl`, 29 hand-labelled cases across five
lanes, so replacing it with real traffic is a small edit rather than a blank
page. Its README is explicit that it is not a benchmark: invented prompts, one
person's labels, 29 cases.

**The harness earned its keep on its first run.** It found that the heuristic
routes

    "Here is my patient record: John Doe, DOB 1971-03-02, diagnosed with
     hypertension. Summarise it."

to the **trivial** lane, because "summarise" is a simple-task verb.

That is acceptable only because the privacy guarantee does not depend on the
classifier -- transforms run after target selection and change the
destination. So this is now asserted rather than assumed, in
`TestLayeringSurvivesAMisclassification`: one test proves the classifier
really does get it wrong (so the premise cannot rot silently), and another
proves the prompt still cannot reach a remote target.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The subsystem's largest honest gap -- no classifier validated on real traffic
-- is still open, but it is now *closable*, and the docs say how rather than
only admitting it.

The new section leads with what the reader must supply (their own labelled
prompts; nothing substitutes for it) and with the instruction to read the
too-cheap and too-expensive lines rather than the percentage. It also explains
the two scoring choices that would otherwise look like bugs: an abstention is
not an error, and with no lane order the report declines to guess a direction.

It ends with the worked example from the first real run, because it is more
convincing than the explanation: the heuristic sends a medical record to the
trivial lane, and the privacy path holds anyway because transforms change the
destination after selection.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The R3 gate is met. Qwen2.5-1.5B-Instruct served by vLLM on a real Colab T4,
reached through `llm.chat()` and through a pool containing the runtime, then
released -- `colab usage` 0.00/hr afterwards, total cost 0.6 compute units.

Every bug below passed the entire test suite, and each would have failed on a
VM that was already billing. That is the argument for one live run over any
amount of mocking:

* **The session parser could not find its own session.** `colab sessions`
  prints `[name] external-id | Hardware: T4 | ...`, not the box-drawn table
  the generic scraper assumed. The scraper *appeared* to cope -- it returned a
  row rather than raising -- with the whole `[name] id` chunk as the name, so
  the assignment guard never recognised the session it had just created and
  released a healthy T4 after 180 seconds. Being forgiving is not the same as
  being right.

* **`colab exec` returns 0 even when the code it ran raised.** The bootstrap
  script detected its own failure and exited saying so; the orchestration read
  rc=0, opened a tunnel to a server that had never started, and settled in to
  wait 45 minutes. Success is now the explicit `[llmcore] READY` marker, which
  the script prints only after the server answers on the VM, and the failure
  message leads with the line that explains it rather than the tail of the
  progress chatter.

* **A CUDA-mismatched torch companion killed the server on import.**
  Installing vLLM upgraded torch to 2.13.0+cu130 while Colab's preinstalled
  torchaudio 2.11.0+cu128 stayed put; transformers imports torchaudio
  unconditionally and it refuses to load against a different CUDA. Removing
  the mismatched companion beats chasing a matching multi-gigabyte build on a
  billing VM -- verified on the live VM before committing to it.

* **The environment cache used the wrong directory, twice.** First a
  hard-coded `/usr/lib/python3/dist-packages` (Colab installs into
  `/usr/local`), then `site.getsitepackages()[-1]`, which resolved to that
  same wrong path on a real VM -- so a future "cache hit" would have restored
  a tree with no recipe in it. Both now use `sysconfig.get_paths()["purelib"]`,
  the path pip itself resolves.

Also: `PIP_CACHE_DIR` pointed at the Google Drive mount, so every wheel
downloaded through FUSE to Drive. The tarball belongs on Drive; the cache is
scratch. And `--dtype half` is now actually passed on pre-Ampere cards -- the
SKU table said T4 has no bf16 and nothing acted on it, and vLLM refuses rather
than downcasting.

`runtimes.colab.auth` selects the CLI auth strategy, so a machine with
application-default credentials can skip the interactive code-paste flow.

Separately, a time-of-day flake: three `most_credits` tests recorded a balance
without an injected clock while asserting against a fixed one, so they passed
in the morning and failed in the afternoon. Verified clock-independent by
re-running with the fixture moment shifted 25 years into the past.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
@araray

araray commented Oct 1, 2026

Copy link
Copy Markdown
Owner Author

Update: Colab gate met, Anthropic still blocked, classifier harness shipped

The Colab gate is met

You authorised the spend, so I ran it. Qwen2.5-1.5B-Instruct on vLLM on a real
Colab T4, reached through llm.chat() and through a pool containing the
runtime:

[2.6s] 'Reply with exactly: COLAB'  ->  COLAB
[1.8s] 'what is grouped-query attention?'  ->  Grouped-query attention refers to...
plan: pool=local strategy=priority chosen=llmcore-e2e:Qwen/Qwen2.5-1.5B-Instruct
                                   ->  ROUTED

Released afterwards; colab usage confirms 0.00/hr, 0 assignments. Total
cost of validation: 0.6 of 198.78 compute units.

It found four bugs that the entire test suite had passed, and every one
would have failed on a VM that was already billing:

  1. The session parser could not recognise its own session. The CLI prints
    [name] external-id | Hardware: T4 | ..., not a table. The generic scraper
    appeared to cope — it returned a row rather than raising — with the whole
    [name] id chunk as the name, so the assignment guard released a healthy
    T4 after 180s. Being forgiving is not the same as being right.
  2. colab exec returns 0 even when the code it ran raised. The bootstrap
    detected its own failure and said so; the orchestration read rc=0, opened a
    tunnel to a server that never started, and waited 45 minutes. Success is now
    an explicit READY marker, not an exit code.
  3. A CUDA-mismatched torch companion killed the server on import — vLLM
    upgraded torch to cu130 while Colab's preinstalled torchaudio stayed on
    cu128. Verified the fix on the live VM before committing to it.
  4. The env cache used the wrong directory, twice — a future "cache hit"
    would have restored a tree with no recipe in it.

Plus: --dtype half is now actually passed on T4 (the SKU table said
pre-Ampere has no bf16 and nothing acted on it), and PIP_CACHE_DIR no longer
points at the Google Drive mount.

Still unproven, and the docs say so: a single unattended up() with all
four fixes. In the successful run the server was started by the corrected
bootstrap invoked by hand, after the orchestration's own attempt had failed.
Everything around it — assignment guard, tunnel, provider attach, chat, pool
routing, status, teardown — ran through the orchestration.

Anthropic: still blocked, re-checked

claude-sonnet-4-5 with budget_tokens returns "Your credit balance is too
low"
, so the generation-aware mapping remains unverified against a live 4.6+
model. The error did exercise the routing classifier correctly, though: a 400
whose body says "credit balance is too low" classifies as
insufficient_credit, not bad_request, so it fails over to a peer rather
than hard-failing.

The classifier: what I need from you

Labelled prompts from your own traffic — nothing substitutes for it.

{"prompt": "rename this variable", "expected": "trivial"}
{"prompt": "prove this lemma", "expected": "deep"}

~50 catches an obviously wrong chain, ~200 lets you compare classifiers. To
make that an edit rather than a blank page I have shipped llmcore-routing eval, a 29-case starter set, and support for unlabelled .txt traffic (which
measures coverage and latency without anyone labelling anything).

The report refuses to give one number, because the two error directions are not
interchangeable:

heuristic
  agreement  10/29 (34%)
  too cheap  11  <- these produce bad answers
  too dear   8   <- these only cost money

It earned its keep on its first run, finding that the heuristic routes
"Here is my patient record: John Doe, DOB 1971-03-02, hypertension. Summarise
it."
to the trivial lane — "summarise" is a simple-task verb. That is
survivable only because the privacy guarantee does not depend on the
classifier, so that is now asserted rather than assumed: one test proves the
classifier really does get it wrong (so the premise cannot rot silently),
another proves the prompt still cannot reach a remote target.

Full suite: 6594 passed, 87 skipped. Also fixed a time-of-day flake I had
introduced — three tests passed in the morning and failed in the afternoon.

🤖 Generated with Claude Code

araray and others added 2 commits October 1, 2026 20:53
…ll bug

Credits arrived, so the thinking/effort mapping was validated against the real
API -- and the measurement showed my first version was wrong.

**There are three boundaries, and they do not coincide.** One request per
model, 2026-10-01:

| generation | type=enabled + budget_tokens | type=adaptive | type=disabled |
|---|---|---|---|
| <= 4.5 | 200, the only option | 400 not supported | 200 |
| 4.6     | 200                  | 200               | 200 |
| >= 4.7 | 400 not supported    | 200               | 400 not supported |

The draft assumed a single cutoff at 4.6, which meant it discarded a caller's
explicit token budget on 4.6 -- a model that would have honoured it -- while
warning about a 400 that could not happen there. 4.6 is an overlap, and an
explicit budget now survives it.

Two further bugs, both of which 400'd live:

* **`{"type": "disabled"}` is rejected from 4.7**, and the API names its own
  replacement: *"send thinking: {"type": "between_tools"} instead"*. So
  `effort="none"` now sends whichever spelling the target accepts.
* **`max_tokens` must be *strictly* greater than `budget_tokens`** -- equal is
  a 400 -- and a budget under 1024 is rejected outright, so `effort="high"`
  with `max_tokens=512` was unsendable. The budget is clamped to fit, and
  where no valid budget exists thinking is turned off with a warning saying
  how to get it back. Sending a request that cannot succeed is worse than
  answering without reasoning and saying so.

**Also fixed, and more serious: every Anthropic call without a system message
failed on the SDK transport.** `system=system_prompt` was passed
unconditionally, so the SDK serialised `"system": null` and the API answered
400 "system: Input should be a valid array". The httpx path already omitted it
when absent. This is the class of bug the dual-transport work exists to
surface -- direct is the default, so the SDK fallback had never been exercised
against a real account.

Validated live: 12 of 12 cases across haiku-4-5, sonnet-4-6 and sonnet-5-5,
asserting both the response and the exact payload sent. Under $0.02 of the $5.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The routing spec's §6.4 and open question 3, and item 2 of the modernization
plan, all described a single generation boundary. The measurement found three.
Both documents now carry the measured table rather than the guess, and say
which way the guess was wrong -- 4.6 is an overlap that accepts every form, so
the single-cutoff version would have discarded a caller's explicit token
budget on a model that honours it.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
@araray

araray commented Oct 1, 2026

Copy link
Copy Markdown
Owner Author

Anthropic validated live — and the measurement corrected my fix

Thanks for the $5. It cost under $0.02 and was worth considerably more than
that, because the measurement showed my first version was wrong.

There are three generation boundaries, not one

I had assumed a single cutoff at 4.6. One request per model against the live
API says otherwise:

generation type=enabled + budget_tokens type=adaptive type=disabled
≤ 4.5 200 — the only option 400 not supported 200
4.6 200 200 200
≥ 4.7 400 not supported 200 400 — use between_tools

4.6 is an overlap that accepts every form. My single-cutoff version
discarded a caller's explicit budget_tokens there — on a model that would
have honoured it — while warning about a 400 that cannot happen. An explicit
budget now survives 4.6 and is only converted from 4.7.

Two more bugs, both of which 400'd live

  • {"type": "disabled"} is rejected from 4.7, and the API names its own
    replacement: "To turn thinking off on this model, send
    thinking: {"type": "between_tools"} instead."
    So effort="none" now
    sends whichever spelling the target accepts.
  • max_tokens must be strictly greater than budget_tokens — equal is a
    400 — and a budget under 1024 is rejected outright. So effort="high" with
    max_tokens=512 was unsendable. The budget is clamped to fit, and where no
    valid budget exists thinking is turned off with a warning saying how to get it
    back. Sending a request that cannot succeed is worse than answering without
    reasoning and saying so.

A more serious bug, found on the way

Every Anthropic call without a system message failed on the SDK transport.
system=system_prompt was passed unconditionally, so with no system message
the SDK serialised "system": null and the API answered 400 "system: Input
should be a valid array"
. The httpx path already omitted it when absent.

This is precisely the class of bug the dual-transport work exists to surface:
direct is the default, so the SDK fallback had never been exercised against a
real account until credits were available.

Validation

12 of 12 live cases across claude-haiku-4-5, claude-sonnet-4-6 and
claude-sonnet-5-5, asserting both the response and the exact payload sent:

PASS  sonnet-5-5 + budget_tokens=1024    sent thinking={"type":"adaptive"}
PASS  sonnet-4-6 + budget_tokens=2048    sent thinking={"type":"enabled","budget_tokens":2048}
PASS  haiku-4-5 + effort=max             sent thinking={"type":"enabled","budget_tokens":32768}
PASS  sonnet-5-5 + effort=none           sent thinking={"type":"between_tools"}
PASS  haiku-4-5 + effort=none            sent thinking={"type":"disabled"}
PASS  haiku-4-5 effort=high max=512      sent thinking={"type":"disabled"}
PASS  haiku-4-5 effort=high max=2048     sent thinking={"type":"enabled","budget_tokens":2047}

The spec's §6.4 and open question 3, and item 2 of the modernization plan, now
carry the measured table instead of the guess — including which way the guess
was wrong.

Full suite: 6617 passed, 87 skipped.

Remaining gaps, unchanged

  • A single unattended Colab up() with all four live-run fixes applied.
  • No classifier validated on real traffic — llmcore-routing eval is ready for
    your labelled prompts whenever you have some.

🤖 Generated with Claude Code

@araray araray self-assigned this Oct 2, 2026
@araray
araray merged commit 99e6cd6 into main Oct 2, 2026
3 checks passed
@araray
araray deleted the av/routing_impl branch October 2, 2026 00:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants