feat: routing subsystem, Colab runtime backend, config audit - #35
Conversation
Two things: a config audit with one real fix, and the design document for the routing programme, which is awaiting approval before any implementation. The audit compared default_config.toml against the provider registry and the config keys the code actually reads. One real gap: xai, groq and together are registered in PROVIDER_MAP and have their env var and base_url in _OPENAI_COMPATIBLE_DEFAULTS, but a provider is only instantiated when its [providers.*] section exists — so all three raised "not configured" and were unreachable. Sections added; coverage is now 23/23 and all three resolve with the right base URLs. Three other apparent gaps (media.jobs.webhook_secret, media.routing, observability.cost_tracking) were false positives in my detector: two are documented as comments and one is a declared table. That gap is also the root of the dynamic-target complaint, and the spec treats it as a design flaw rather than a missing feature: config is acting as an allow-list when it should be a preset layer. Adding the three sections only moves the wall. The spec separates what was asked as one feature into five composable layers, because specifying them as one would do every job badly: Target (which provider/model/params, addressable as a spec string, autoprovisioned when absent from config), Pool (failover and prioritisation), Lane + Classifier (what kind of request is this), Cascade (was the cheap answer good enough), and Transform (what must not leave this machine). Each is independently useful. It proposes Pool and Lane as two names rather than one "group", because the request used one word for two ideas: a pool is a set of interchangeable targets for failover, a lane is a named destination chosen by a classifier. Collapsing them forces every group to be both. Grounded in prior art rather than invented: LiteLLM's router for strategies, cooldowns and error-specific retry; Arch-Router for the key insight that one should route to named lanes with the lane-to-model binding kept separate, so swapping models needs no retraining; RouteLLM for strong/weak routing; FrugalGPT for the cascade. Real downloadable local classifiers exist and are named, including a 350M CPU encoder that scores a prompt against free-text lanes zero-shot, and a family of 307M task encoders covering domain, PII and factuality. Four contributions beyond transcription are flagged as such: the failure taxonomy distinguishes 429 from insufficient-credit from context-length, and treats ContextLengthError as a routing signal toward a larger window rather than an error; content-policy refusals do not fail over by default, because retrying a refusal elsewhere is a decision a user must opt into; session affinity is specified because mid-conversation failover loses prompt caching and shifts style, which the request did not raise; and the privacy path constrains the pool to local targets rather than relying on redaction, with the documentation required to say that redaction alone is not a guarantee. Balance-based prioritisation is specified honestly: most vendors publish no balance endpoint, so the protocol returns None for unknown, unknown never ranks as zero, and an observed insufficient-credit error is what supplies the signal where no endpoint exists. No accuracy or saving numbers appear anywhere, because none have been measured on this library's traffic. An evaluation harness ships with the classifier phase instead. Five open questions are listed for decision, including naming, the default pool strategy, and the scope of the first PR. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
First slice of the routing subsystem (T1 of docs/ROUTING_SUBSYSTEM_SPEC.md). `Target` parses the spec strings routing is addressed with -- `provider:model?params#instance`. The split is on the FIRST colon, because model names legitimately contain colons and slashes (`ollama:llama3.3:70b`, `vllm:Qwen/Qwen3-30B`); splitting on the last one would mangle them. `Target.key` deliberately excludes params, since health and rate limits belong to the endpoint, not to a requested effort level. `ProviderManager.resolve_target()` makes config a set of presets rather than an allow-list: an unconfigured `provider:model` is built on demand from a discovered credential and marked ephemeral. An explicit `#instance` pin is never autoprovisioned -- naming something that does not exist is an error, not a request to build something adjacent. `RoutingSettings` resolves config -> env -> request, and the request wins. Policy lives here, not in the failure enum: `FailureKind` says what a failure *is* (`failover_is_pointless`, `affects_health`, `retry_same`) and settings say what to do about it (`on_refusal`). Billing is classified before auth, so a valid key on an empty account gets a 10-minute cooldown instead of being benched for the process. Cascade is off by default -- it costs an extra call. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The implementation PR should carry the spec it implements, so the two land together whichever order the PRs merge in.
T2 and T3 of the spec. A pool is a set of interchangeable targets; selection is split into eligibility (facts) and ordering (strategy), and every member is reported with the reason it was or was not used -- "why did it pick the slow one?" is only answerable if the record says the fast one had twelve seconds of 429 cooldown left. All seven strategies land together because they share that scaffolding: priority, round_robin, weighted, lowest_latency, lowest_cost, least_busy, most_credits. Three of them needed a judgement call about missing data, and in each case treating "unknown" as a number would have inverted the strategy: * `lowest_cost` ranks unpriced targets *after* priced ones. Scoring an unknown price as zero would make a model with no card beat a model known to be free. * `most_credits` uses three bands -- known funds, unknown, known-empty. The obvious descending sort puts a confirmed zero ahead of an unknown, which is backwards: an unknown account might have money. * `lowest_latency` explores an unmeasured target first, since it cannot prefer low latency without a measurement and one call buys one. Self-hosted providers (ollama, vllm) are priced at zero rather than unknown. Their cards carry no pricing block, and reading that absence as unknown would rank your own GPU behind a paid API instead of ahead of it. Order tiers (`?order=`) gate eligibility across strategies, which is how "use my own GPU, and only pay a vendor if it is down" is expressed. Routing params are stripped from the target before it reaches a provider, since forwarding `weight` to a vendor is a 400. Also here: a pre-call context-window check that skips a target whose card already proves the prompt will not fit; `InMemoryRoutingState` with an injectable clock and context-managed in-flight counters; and routing events on the existing shared_events spine, guarded by `has_sinks()` so routing costs nothing when nobody is listening. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
T5 and T6. A classifier names a *lane* and never a model, so swapping a model is a config edit rather than a classifier change -- the one idea worth borrowing wholesale from Arch-Router. That indirection is also what collapses the original request's complexity tiers, speed tiers, domain routing and privacy class into a single mechanism: "focus groups per complexity level" are lanes whose names happen to be complexity levels, and nothing in lanes.py knows that. Chain ordering is cheapest-first, then instructions before guesses, both declared per classifier rather than inferred from config order. The second half of that rule is not a nicety: `hint` and `heuristic` are both free, so cost alone cannot separate them, and running the guess first would silently override what the caller asked for. Authority has four levels rather than two, because a routing marker found in *content* is not as trustworthy as an argument on the call. Markers are the point of the feature -- in an agent harness the model's text is the only channel that passes through, so that is how an agent routes itself -- but in any RAG or tool-output path that text may have come from a retrieved document, where `[[lane:deep]]` would be a one-line prompt injection, or worse, a way out of the private lane. So `caller` > `policy` > `prompt` > `inferred`, and a marker is honoured only when nobody more authoritative spoke. Two things measurement changed, both now corrected in the spec: * The T6 gate claimed local classification adds "<50 ms p50 on CPU". Real figure on 8 CPU threads: 191 ms for 2 lanes, 246 ms for 5, 314 ms for 9. Wrong by ~5x, which is exactly why the gate said measured-not-assumed. The encoder is off by default, runs its forward pass in a worker thread so it cannot stall the event loop, and is now documented as a batch/agent feature rather than an interactive one. * The encoder's raw top score is meaningless without reading it against chance -- 0.20 over five lanes is exactly uniform, i.e. no opinion, and taking it as 20% confidence would route on noise. Confidence is now reported chance-corrected, so one floor means the same thing at any lane count, and prompts matching no lane correctly abstain. The heuristic lost its short-prompt rule for the same reason: length does not predict complexity. "Write a 2000-word essay on X" is ten tokens, and so is "tell me about the Azores in autumn". Length now only ever vetoes the trivial lane; selecting it needs a recognisable simple-task verb. `RoutingRequest.prompt` is optional, since `chat(messages=[...])` is a real call shape with no prompt. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The piece where the five layers compose, plus the two layers they needed: prompt transforms (T8) and response verifiers (T7). Order matters and is fixed: classify, resolve, transform, attempt, cascade. Transforms run *after* target selection, because a transform's whole value is that it sees where the prompt is about to go -- the cost is that a `constrain` action re-resolves against a different pool, which the manager does explicitly and once. On the privacy layer, the documentation says plainly what the feature is and is not: **redaction is a mitigation, routing is the guarantee.** A detector that misses one identifier has leaked it, and no detector catches everything, so `constrain` changes the *destination* rather than the text. Three places where the obvious implementation would have quietly undone that, all now closed: * `constrain` with no `pool` configured would degrade to "send it anyway", which is the one outcome the user was preventing. It blocks instead. * a transform that raises would look like "nothing found". The chain fails closed by default. * a constrain pool that does not exist would fall through to the remote target. It blocks. Findings carry a 16-char hash, never the value, including in `PromptBlockedError` -- a privacy feature whose exception text contained the identifiers it found would be the leak it exists to prevent. Verdicts are three-valued, and `None` means *could not judge*. Folding it either way is the bug that makes cascades useless: read as a fail it escalates every unjudgeable answer and inverts the saving, read as a pass it silently disables the quality floor the moment the judge breaks. So it stays a third value and `routing.cascade.on_unknown` decides. Two bugs the end-to-end run found: * `RoutingSettings.magic_pattern` defaulted to a regex with one *unnamed* group, and the manager handed it to a classifier that reads named groups `key` and `value` -- an IndexError on every classified request. The default is now None (use the built-in) and a custom pattern is validated at build time with a message that shows the expected shape. * A three-group phone number was only partially redacted, leaving the last digits behind; and `192.168.1.42` matches both the ipv4 and phone patterns, so redacting both rewrote the same span twice. Hits are now de-overlapped, longest match first. Also fixed, and more important than it looks: the test fixture built confy `Config` objects with its default `load_dotenv_file=True`, which exports the repo's `.env` into `os.environ`. That made a credential-discovery assertion depend on the developer's machine, and it meant a test could reach a real vendor and spend real money. Fixtures now pass `load_dotenv_file=False`. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
`[routing]` in default_config.toml, `llm.chat(target=/pool=/lane=/profile=/
complexity=/effort=/routing=)`, the `llm.routing` facade, and proxy mode on
the bridge (T9).
The proxy exists because of a constraint worth stating: an agent harness owns
its own API call, so llmcore cannot add a `lane=` argument to it -- but it
almost always lets you set a base URL and a model name. So those two strings
are the entire interface. `model="lane:deep"`, `"pool:main"`, a target spec,
a bare model name or `"auto"` all work, and a harness gets pools, failover,
classifiers, cascades, transforms and cost accounting without knowing llmcore
exists.
Two things the proxy refuses to hide. Usage reports the target that *actually*
served the request in an `llmcore` block, because under a pool the answering
model is genuinely not the one that was asked for and a harness logging spend
per model should not be lied to. And the bind refuses: this process holds
every provider credential in the config, so a non-loopback host without a
bearer token raises at construction rather than starting.
Live validation against real vendors found three real bugs:
* **Gemini had no pricing and no context window, silently.** The card alias
map ran the wrong way -- provider type `gemini` was mapped *to* a `gemini`
card namespace, but cards are filed under `google/`. Nothing raised: the
lookup returned None, None means "unknown", and every caller handles
unknown quietly. So `lowest_cost` could not price any Gemini target and the
pre-call window check never fired, for one of the most used providers in the
library. tests/routing/test_cards.py now audits every provider type against
the packaged card tree so the next provider cannot reintroduce it.
* `_call_with_failover` passed `chat()`'s keyword names to
`prepare_context()`, which has different ones -- so the first failover
raised instead of failing over.
* A 404 was classified `unknown`, which retried the same target before moving
on. A missing model is deterministic for that target but a peer serves a
*different model*, so it is the one case where "permanent" and "try someone
else" are both true. It now has its own `MODEL_NOT_FOUND` kind, checked
before the 400 branch because several vendors answer 400 for it.
Also corrected a regression I introduced: routing made unsupported provider
kwargs drop with a warning everywhere, which silently broke the existing
contract that `chat("x", bogus=1)` raises. Dropping now applies only when a
*pool* chose the target -- where members genuinely have different parameter
surfaces and the caller cannot know which one serves -- and raises everywhere
else.
Verified live: failover over a nonexistent model, `lowest_cost` picking the
cheaper of two real vendors from card pricing, a 700k-char prompt skipping a
128k window for a 1M one, and an unconfigured provider reached by spec string.
Full suite: 6358 passed, 87 skipped. The only failures left need a Postgres
server (`PoolTimeout` after 30s) and fail the same way on main.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
`llmcore-bridge proxy` starts the OpenAI-compatible surface. A separate subcommand rather than another transport on `serve`, because it is a different thing: `serve` exposes llmcore's own contract to llmcore clients, while this impersonates somebody else's API so an unmodified harness can be pointed at it. A refused bind prints the reason instead of a traceback. Validated against the real `openai` Python SDK pointed at the proxy: `/v1/models` listing lanes and pools, pool routing with failover over a broken member, lane routing through the model name, an explicit target spec, streaming, and a bogus model surfacing as the SDK's own `NotFoundError`. Three bugs the tests and that run found: * The proxy called `get_last_interaction_info()`, which does not exist -- the real method is `get_last_interaction_context_info()`. My fake had defined the wrong name, so the unit tests passed while the usage block would have been empty against a real LLMCore. * Usage was empty whenever a harness sent no session, since llmcore keys its per-turn introspection by session id. That defeats the endpoint's one promise -- to say which target actually answered. The proxy now uses a synthetic session, and `LLMCore.discard_transient_state()` drops it afterwards; without that, a proxy handling one synthetic session per request leaks a cache entry per call for the life of the process. * `prompt_tokens` came back 0 after a failover. `chat()` derives it *after* `prepare_context()` returns, so copying the re-prepared details over the originals overwrote it with the unset value. Excluded and recomputed. Error mapping now runs every exception through `classify_failure` rather than only `ProviderError`, so a bare `TimeoutError` is a 504 instead of a 500 -- harnesses branch on these statuses, and getting them right is what makes the proxy behave like the thing it is impersonating. Routing and the proxy are ruff-clean; api.py's finding count went down. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
`docs/Routing_usage.md` is the how-to; the spec stays the why. Its last section is "What this does not claim", which lists the limits plainly: no accuracy figure for any classifier (none is validated on anyone's traffic, so any number would be invented), cost estimates are for comparing targets rather than for billing, PII redaction is not a leak guarantee, the local classifier is not free in latency, routing state is per process, and `most_credits` only works where a vendor exposes a balance. Three skills in the bundled grimoire pack, so an agent can pick them up: `skills/llmcore/proxy` (point a harness at llmcore), `skills/llmcore/routing` (configure and debug it), `skills/llmcore/cost` (measure, then reduce). `skilldoc_paths` added to the pack manifest; the pack now loads 21 spells, 3 runes, 7 promptlets and 3 skilldocs. The spec is marked implemented, its five open questions are answered in a new §13, and §13.1 records the six places the implementation forced a change on the design -- because a spec that quietly diverges from its code is worse than no spec. Three honest gaps are left open, the largest being that no classifier has been validated against real traffic. Every `provider:model` string in the new docs was checked against the bundled model cards; all resolve, except `vllm:Qwen/Qwen3-30B`, which is correct since vLLM is self-hosted and ships no cards by design. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
…t context Read-only and free: nothing in `sizing.py` can start billable compute, which is why sizing is a separate phase. Every `Plan` carries its own arithmetic, because a sizer that answers "use an A100" and shows nothing is impossible to argue with. Every unknown rounds toward *needing more*, deliberately. Overestimating buys a bigger GPU; underestimating OOMs on the VM after billing has started. Three things the real Hub and the real CLI corrected in the draft design: * **`A100-40` and `A100-80` are not SKUs you can ask for.** `colab new --gpu` accepts T4, L4, G4, H100, A100 and nothing else, so a plan naming A100-80 could never have been provisioned. The ladder now uses the CLI's own names, A100 is sized at the conservative 40 GB (being handed an 80 GB card only adds headroom), and the draft spellings are kept as aliases so a config written against the spec still works. A test asserts every SKU's `--gpu` value is one the CLI accepts. * **`Quantization` could not express GGUF precision.** Q4 and Q8 differ by 2x in weight bytes, which is routinely the difference between fitting a 24 GB card and not, so a single `GGUF` member forced the sizer to guess. Split into `GGUF_Q4/Q5/Q8`, plus `INT4`. * **The KV fallback produced nonsense.** For a gated repo whose config.json is unreadable, the draft formula estimated 0.5 GB for a 70B model at 8k context -- about 5x under. It now scales from the parameter count at 8 KB/token/B, calibrated against models where the real figures *are* available (3.2 / 4.6 / 7.5 KB measured for Qwen3-30B, Llama-3.3-70B, Qwen2.5-7B) and taking the high end on purpose. The GQA term is read explicitly from `num_key_value_heads`, since falling back to `num_attention_heads` when the key is present overestimates Llama-3.3-70B's cache by 8x. The KV cache is priced at 16 bits even under 4-bit weights, because that is what vLLM actually does. Context shrinks before a bigger GPU is chosen -- a bigger GPU costs money, a shorter context costs nothing -- with a floor at 4k, since silently handing someone a 1k window is worse than saying it does not fit. When nothing fits, the sizer refuses *with a concrete alternative*, because a bare refusal sends the caller to guess and the next guess is also a launch. `vram_available_gb` reports the budget the fit decision actually compared against, not the sticker VRAM, so a plan cannot look roomier than the sizer believed. Validated against the live Hub: Qwen2.5-7B, Qwen3-30B (dense and AWQ), Llama-3.3-70B (gated, so exercising the fallback) and gemma-3-4b. 45 tests, one of them marked `live` against the real Hub because its metadata shape is the thing most likely to change underneath this. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The repo uses ASCII in strings; ruff flags the multiplication sign as ambiguous. The arithmetic notes read the same either way. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Provisions a Colab GPU VM, serves an open-weights model on it, tunnels the endpoint to localhost, and attaches it as an ordinary provider. Plus `bake` and the Drive cache inventory (R5). This is the module that spends money, so four rules are enforced in code rather than left to the caller, and each has a test named after it: 1. **State before compute.** The handle is on disk before `colab new` runs. The dangerous window is a crash between assignment and bookkeeping -- a VM billing with nothing aware of it. 2. **Fail closed.** One error path, and it ends in `down`. If the release *also* fails, the log says MAY STILL BE BILLING in those words, because that is the one case a user has to act on. 3. **Never SSH into the void.** The session must appear in `colab sessions` before anything connects. Without the guard a failed assignment becomes a confusing SSH timeout instead of "no GPU available". 4. **Bounded at creation.** Idle *and* hard deadlines, since an idle reaper does not stop a runtime that is busy in a loop. Everything that leaves the process goes through one seam, so there is a single place where a command is logged, timed out and captured -- and so the tests can drive the backend by asserting on the argv it would have run. Checked against the real CLI (`google-colab-cli` 0.7.4) rather than the spec's description of it, which was wrong in one place: `colab new --gpu` accepts T4, L4, G4, H100, A100 and nothing else, so the spec's `A100-80` rung could never have been provisioned. A test now asserts every SKU's `--gpu` value is one the CLI accepts. Secrets go on stdin or through `colab exec --env`, never in argv -- argv is visible in the VM's process list and in llmcore's own debug logs. Two tests assert the token appears exactly once across all argv and never inside the generated bootstrap script. The VM-side server is started with `setsid`: as a child of the kernel, a kernel restart would kill the server silently while the VM kept billing. Cached artefacts in Drive are gated on a sentinel written *after* the artefact is complete, because a miss costs time while a corrupt hit costs a debugging session on a billing VM. The bootstrap also waits for readiness on the VM, so a startup failure is reported with the log that explains it rather than as a local connection timeout. Orphan detection is a safety feature: a Colab session llmcore has no record of is reported with the `adopt` command that makes it killable. An adopted runtime is marked DEGRADED and says it does not know what it is running -- claiming READY would be worse than admitting the gap. `RuntimePhase.DETACHED` added, and it counts as billing. Detached means llmcore stopped watching, not that the VM stopped; counting it as not-billing would hide exactly the leak these rules exist to prevent. `_parse_sessions` is deliberately forgiving and falls back to scraping names, because it is how llmcore *finds things that are costing money* -- a parser that raised on a cosmetic CLI change would break teardown precisely when it is needed. 58 tests. **The spec's R3 gate -- one real model served end to end -- is not met**, and deliberately: meeting it means provisioning a real GPU VM and spending the account's compute units, which is not mine to authorize. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
**R4.** `reap()` existed but nothing called it, which made the idle and hard deadlines documentation rather than limits. `up()` now starts a supervisor that probes liveness and reaps on a timer -- started automatically, because a caller who forgot would discover it from a bill. Liveness marks a target DEGRADED after three consecutive failures, not one: a single missed poll is usually the tunnel reconnecting. A failed probe never tears a runtime down. A degraded runtime is still assigned and still billing, so releasing it belongs to the deadlines or to the user -- letting a transient network problem destroy an expensive VM would be worse than the problem. The idle reaper needs to know when a runtime was last used and vLLM exposes no last-request metric, so the count happens on llmcore's side: `attach()` wraps the provider's `chat_completion`, which is exactly one call per real request. Hooking `get_provider()` would miss a long stream; hooking the tunnel would need a proxy. The wrap is idempotent, since a runtime may be re-attached after a reconnect. **R6.** `llmcore-runtimes` — estimate, up, status, down, logs, adopt, bake, cache. A CLI because the commands that matter most are the ones needed when something has gone wrong: `status` and `down` have to work from a shell, in a hurry, possibly from a different process than the one that started the runtime. `up` requires `--yes` and prints the burn rate first; `estimate` is free and works while the subsystem is disabled, because deciding whether to spend should not require enabling spend. **Fixed: the test suite was writing into the user's real `~/.llmcore/runtimes`.** That directory is the record of what is currently costing money -- the spec calls it a safety mechanism rather than a cache, and `llmcore-runtimes status` reads it. Running the tests left a phantom entry claiming a READY L4 VM was running, which is exactly the false signal the subsystem exists to prevent: someone checking whether they were being billed would have been told yes, by their own test run. An autouse fixture now redirects the default state directory per test. `RuntimePhase.DETACHED`, `runtimes.defaults.supervise_seconds` and `runtimes.colab.keepalive_seconds` added to the config, with the SKU ladder corrected to the names the CLI accepts. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
`docs/Runtimes_usage.md` with a "What this does not claim" section that leads with the unmet gate rather than burying it: nobody has run this against a real Colab VM, because that spends real compute units. The spec's phase table now marks R2 and R4 met, and R3/R5/R6 implemented with their gates explicitly unmet and the reason given. `cache gc` is noted as not implemented. The agent-lens migration guide is not written, because writing it would mean telling another project to depend on an unproven path. **Config audit** (the other half of the request): all 660 keys in default_config.toml checked against the code. All 48 `[routing]` keys and all 14 `[runtimes]` keys are wired. The 116 unread `semantiscan.*` keys belong to that package. Two keys promise behaviour that does not exist and are now marked NOT ENFORCED in the file rather than silently removed: * `llmcore.admin_api_key` claimed to protect "administrative endpoints such as live configuration reloading". Nothing reads it, and the bridge's ReloadConfig has no reference to it -- so a deployment that set this believing its reload endpoint was guarded was wrong. The comment now says so, and points at the transport-level auth that does work. * `context_management.minimum_history_messages` -- truncation does not honour it. Both kept rather than deleted so existing configs keep loading. A security control that silently does nothing is worse than one that is absent: someone has to be able to find out. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
`thinking.budget_tokens` is deprecated on Claude 4.6 and **rejected with a
400** on 5.x, while pre-4.6 models require it. llmcore forwarded whatever the
caller passed, so a reasonable request became an API error purely because of
which model served it. Recorded in PROVIDER_MODERNIZATION_PLAN.md and flagged
again by the routing spec -- pools make this routine, since members span
generations.
`effort` is now a first-class Anthropic parameter and means the same thing on
either generation: 4.6+ gets `{"type": "adaptive"}` plus
`output_config.effort`, earlier models get a token budget. A caller's
`budget_tokens` to a 5.x model is dropped with a warning naming the 400 it
would have caused, and a request for `adaptive` on an older model is converted
rather than failing.
Two judgement calls worth stating. An unparseable model id is assumed modern,
because the pre-4.6 family is the shrinking set and defaulting the other way
would break every new model by default. And `effort="none"` disables thinking
rather than requesting adaptive-at-low, which would quietly spend reasoning
tokens the caller explicitly asked not to spend.
37 tests. The live gate is unmet for a reason outside the code: the Anthropic
account available has no credit balance, so nothing could be sent to a real
4.6+ model to confirm the 400 is gone.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Privacy transforms are bypassed or ineffective in production paths, and runtime supervision does not reliably enforce billing deadlines.
Review effort: Balanced
Findings: 7
Open (11)
Routing transforms receive an empty prompt instead of user input · New Bearer token is exposed in application logs · New String false override incorrectly enables remote code execution · New External classifiers receive prompts before privacy transforms · New Configured transforms are bypassed on several chat paths · New Transformed routing request is not passed to chat provider · New CLI closes supervisor before runtime deadlines are enforced · New Re-planning incorrectly consumes a provider attempt · New Cascade repeats the already-used target before escalation · New Zero supervision interval incorrectly replaced with default · New Constrain does not guarantee local routing without detection · New
What changed in this PR
Adds a comprehensive routing subsystem, Colab runtime backend, operational CLIs, provider fixes, documentation, and tests.
Changes:
- Introduces targets, pools, lanes, classifiers, transforms, cascades, and proxy mode.
- Implements Colab runtime sizing, lifecycle supervision, and CLI operations.
- Fixes model-card resolution and Anthropic thinking compatibility.
| File | Description |
|---|---|
src/llmcore/api.py |
Integrates routing into chat calls. |
src/llmcore/routing/manager.py |
Orchestrates routing and cascades. |
src/llmcore/routing/* |
Adds routing models, settings, state, events, lanes, cards, protocols, classifiers, transforms, and verifiers. |
src/llmcore/runtimes/* |
Adds runtime lifecycle models, supervision, and CLI. |
src/llmcore/providers/manager.py |
Adds dynamic provider resolution. |
src/llmcore/providers/anthropic_provider.py |
Normalizes thinking across Claude generations. |
src/llmcore/bridge/cli.py |
Adds proxy CLI mode. |
src/llmcore/exceptions.py |
Adds routing exceptions. |
src/llmcore/grimoire_pack/* |
Adds routing, proxy, and cost skill documentation. |
tests/routing/* |
Covers routing configuration and card lookup. |
tests/runtimes/* |
Covers runtime isolation and supervision. |
tests/providers/test_anthropic_thinking.py |
Covers thinking normalization. |
README.md |
Documents routing and Colab runtimes. |
docs/Runtimes_usage.md |
Adds the runtime usage guide. |
docs/COLAB_RUNTIME_SPEC.md |
Updates implementation status. |
CHANGELOG.md |
Records the new features and fixes. |
pyproject.toml |
Registers the runtime CLI and live-test marker. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| request = self._build_routing_request( | ||
| (getattr(chat_session, "id", None) and "") or "", | ||
| session_id=getattr(chat_session, "id", None), | ||
| ) |
| for key, value in config.harness_env().items(): | ||
| # Printed so the operator can paste it straight into the harness. | ||
| # The token is the one they configured, so echoing it here tells | ||
| # them nothing they did not already set. | ||
| logger.info(" %s=%s", key, value) |
| device=str(encoder.get("device") or "cpu"), | ||
| min_confidence=float(encoder.get("min_confidence", 0.35)), | ||
| max_chars=int(encoder.get("max_chars", DEFAULT_MAX_CHARS)), | ||
| trust_remote_code=bool(encoder.get("trust_remote_code", True)), |
| The protocols here, in the order a request meets them: | ||
|
|
||
| 1. :class:`RequestClassifier` — "what kind of request is this?" | ||
| 2. :class:`PromptTransform` — "what must not leave this machine?" | ||
| 3. :class:`BalanceProbe` — "how much is left on this account?" | ||
| 4. :class:`ResponseVerifier` — "was the cheap answer good enough?" |
| async def apply( | ||
| self, request: RoutingRequest, target: Target | ||
| ) -> tuple[RoutingRequest, tuple[TransformResult, ...]]: | ||
| """Run the chain and return the final request plus every result. | ||
|
|
| finally: | ||
| await llm.close() |
| # Re-plan against the constrained pool. This does not count as | ||
| # an attempt: nothing was called. | ||
| target = None | ||
| lane = None | ||
| continue |
| # Escalate. The rung we were already served by is skipped, so a cascade | ||
| # whose first rung is also the default pool does not pay twice for the | ||
| # same answer. | ||
| max_rungs = max(1, min(settings.cascade_max_rungs, len(rungs))) | ||
| used = result.rungs_used | ||
| for rung in rungs: | ||
| if used >= max_rungs + 1: | ||
| break | ||
| try: | ||
| escalated = await self.execute( |
| seconds = float( | ||
| interval | ||
| if interval is not None | ||
| else self._get("runtimes.defaults.supervise_seconds", 60.0) or 60.0 | ||
| ) |
| A detector that misses one identifier has leaked it, so the guarantee is the | ||
| *route*: a prompt with personal data in it goes to a pool that never leaves the | ||
| machine. Redaction is stacked on top and documented as a mitigation rather than | ||
| a guarantee. |
llmcore claims no accuracy for any classifier, because none has been
validated on real traffic and any number would be invented. This is the
machinery for replacing that gap with a measurement, and it is built around
one property a single accuracy figure destroys:
The two error directions are not interchangeable.
Routing too cheap produces a bad answer. Routing too expensive only costs
money. The same 67% can be usable or unusable depending on which way it errs,
so every report separates them and labels the consequence. There is a test
asserting that two classifiers with identical 0% accuracy produce opposite
reports, because that is the whole reason the harness exists.
Two other things it refuses to do: abstentions are not scored as errors (a
classifier abstaining is passing the turn down the chain, which is the
designed behaviour -- scoring it as wrong would make the most honest
classifier look like the worst), and with no lane order supplied it reports no
direction rather than guessing one.
`llmcore-routing` -- `why`, `health`, `show`, `eval`. Plus
`docs/examples/lane_eval_starter.jsonl`, 29 hand-labelled cases across five
lanes, so replacing it with real traffic is a small edit rather than a blank
page. Its README is explicit that it is not a benchmark: invented prompts, one
person's labels, 29 cases.
**The harness earned its keep on its first run.** It found that the heuristic
routes
"Here is my patient record: John Doe, DOB 1971-03-02, diagnosed with
hypertension. Summarise it."
to the **trivial** lane, because "summarise" is a simple-task verb.
That is acceptable only because the privacy guarantee does not depend on the
classifier -- transforms run after target selection and change the
destination. So this is now asserted rather than assumed, in
`TestLayeringSurvivesAMisclassification`: one test proves the classifier
really does get it wrong (so the premise cannot rot silently), and another
proves the prompt still cannot reach a remote target.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The subsystem's largest honest gap -- no classifier validated on real traffic -- is still open, but it is now *closable*, and the docs say how rather than only admitting it. The new section leads with what the reader must supply (their own labelled prompts; nothing substitutes for it) and with the instruction to read the too-cheap and too-expensive lines rather than the percentage. It also explains the two scoring choices that would otherwise look like bugs: an abstention is not an error, and with no lane order the report declines to guess a direction. It ends with the worked example from the first real run, because it is more convincing than the explanation: the heuristic sends a medical record to the trivial lane, and the privacy path holds anyway because transforms change the destination after selection. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The R3 gate is met. Qwen2.5-1.5B-Instruct served by vLLM on a real Colab T4, reached through `llm.chat()` and through a pool containing the runtime, then released -- `colab usage` 0.00/hr afterwards, total cost 0.6 compute units. Every bug below passed the entire test suite, and each would have failed on a VM that was already billing. That is the argument for one live run over any amount of mocking: * **The session parser could not find its own session.** `colab sessions` prints `[name] external-id | Hardware: T4 | ...`, not the box-drawn table the generic scraper assumed. The scraper *appeared* to cope -- it returned a row rather than raising -- with the whole `[name] id` chunk as the name, so the assignment guard never recognised the session it had just created and released a healthy T4 after 180 seconds. Being forgiving is not the same as being right. * **`colab exec` returns 0 even when the code it ran raised.** The bootstrap script detected its own failure and exited saying so; the orchestration read rc=0, opened a tunnel to a server that had never started, and settled in to wait 45 minutes. Success is now the explicit `[llmcore] READY` marker, which the script prints only after the server answers on the VM, and the failure message leads with the line that explains it rather than the tail of the progress chatter. * **A CUDA-mismatched torch companion killed the server on import.** Installing vLLM upgraded torch to 2.13.0+cu130 while Colab's preinstalled torchaudio 2.11.0+cu128 stayed put; transformers imports torchaudio unconditionally and it refuses to load against a different CUDA. Removing the mismatched companion beats chasing a matching multi-gigabyte build on a billing VM -- verified on the live VM before committing to it. * **The environment cache used the wrong directory, twice.** First a hard-coded `/usr/lib/python3/dist-packages` (Colab installs into `/usr/local`), then `site.getsitepackages()[-1]`, which resolved to that same wrong path on a real VM -- so a future "cache hit" would have restored a tree with no recipe in it. Both now use `sysconfig.get_paths()["purelib"]`, the path pip itself resolves. Also: `PIP_CACHE_DIR` pointed at the Google Drive mount, so every wheel downloaded through FUSE to Drive. The tarball belongs on Drive; the cache is scratch. And `--dtype half` is now actually passed on pre-Ampere cards -- the SKU table said T4 has no bf16 and nothing acted on it, and vLLM refuses rather than downcasting. `runtimes.colab.auth` selects the CLI auth strategy, so a machine with application-default credentials can skip the interactive code-paste flow. Separately, a time-of-day flake: three `most_credits` tests recorded a balance without an injected clock while asserting against a fixed one, so they passed in the morning and failed in the afternoon. Verified clock-independent by re-running with the fixture moment shifted 25 years into the past. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Update: Colab gate met, Anthropic still blocked, classifier harness shippedThe Colab gate is metYou authorised the spend, so I ran it. Qwen2.5-1.5B-Instruct on vLLM on a real Released afterwards; It found four bugs that the entire test suite had passed, and every one
Plus: Still unproven, and the docs say so: a single unattended Anthropic: still blocked, re-checked
The classifier: what I need from youLabelled prompts from your own traffic — nothing substitutes for it. {"prompt": "rename this variable", "expected": "trivial"}
{"prompt": "prove this lemma", "expected": "deep"}~50 catches an obviously wrong chain, ~200 lets you compare classifiers. To The report refuses to give one number, because the two error directions are not It earned its keep on its first run, finding that the heuristic routes Full suite: 6594 passed, 87 skipped. Also fixed a time-of-day flake I had 🤖 Generated with Claude Code |
…ll bug
Credits arrived, so the thinking/effort mapping was validated against the real
API -- and the measurement showed my first version was wrong.
**There are three boundaries, and they do not coincide.** One request per
model, 2026-10-01:
| generation | type=enabled + budget_tokens | type=adaptive | type=disabled |
|---|---|---|---|
| <= 4.5 | 200, the only option | 400 not supported | 200 |
| 4.6 | 200 | 200 | 200 |
| >= 4.7 | 400 not supported | 200 | 400 not supported |
The draft assumed a single cutoff at 4.6, which meant it discarded a caller's
explicit token budget on 4.6 -- a model that would have honoured it -- while
warning about a 400 that could not happen there. 4.6 is an overlap, and an
explicit budget now survives it.
Two further bugs, both of which 400'd live:
* **`{"type": "disabled"}` is rejected from 4.7**, and the API names its own
replacement: *"send thinking: {"type": "between_tools"} instead"*. So
`effort="none"` now sends whichever spelling the target accepts.
* **`max_tokens` must be *strictly* greater than `budget_tokens`** -- equal is
a 400 -- and a budget under 1024 is rejected outright, so `effort="high"`
with `max_tokens=512` was unsendable. The budget is clamped to fit, and
where no valid budget exists thinking is turned off with a warning saying
how to get it back. Sending a request that cannot succeed is worse than
answering without reasoning and saying so.
**Also fixed, and more serious: every Anthropic call without a system message
failed on the SDK transport.** `system=system_prompt` was passed
unconditionally, so the SDK serialised `"system": null` and the API answered
400 "system: Input should be a valid array". The httpx path already omitted it
when absent. This is the class of bug the dual-transport work exists to
surface -- direct is the default, so the SDK fallback had never been exercised
against a real account.
Validated live: 12 of 12 cases across haiku-4-5, sonnet-4-6 and sonnet-5-5,
asserting both the response and the exact payload sent. Under $0.02 of the $5.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The routing spec's §6.4 and open question 3, and item 2 of the modernization plan, all described a single generation boundary. The measurement found three. Both documents now carry the measured table rather than the guess, and say which way the guess was wrong -- 4.6 is an overlap that accepts every form, so the single-cutoff version would have discarded a caller's explicit token budget on a model that honours it. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Anthropic validated live — and the measurement corrected my fixThanks for the $5. It cost under $0.02 and was worth considerably more than There are three generation boundaries, not oneI had assumed a single cutoff at 4.6. One request per model against the live
4.6 is an overlap that accepts every form. My single-cutoff version Two more bugs, both of which 400'd live
A more serious bug, found on the wayEvery Anthropic call without a system message failed on the SDK transport. This is precisely the class of bug the dual-transport work exists to surface: Validation12 of 12 live cases across The spec's §6.4 and open question 3, and item 2 of the modernization plan, now Full suite: 6617 passed, 87 skipped. Remaining gaps, unchanged
🤖 Generated with Claude Code |



Implements the routing subsystem (T1–T10) and the Colab runtime backend
(R2–R6), plus the config audit that was asked for alongside them.
Supersedes #34 — that branch is merged in here, so the spec and its
implementation land together whichever order they are reviewed in.
Routing
Five layers that compose, each usable alone. Nothing is on by default:
with no
[routing]section every call resolves exactly as it did before.Config stops being an allow-list. A provider registered in
PROVIDER_MAP,with its key in the environment and its base URL already known, was still
unreachable. Any provider+model is now addressable by spec string:
Found on the way:
xai,groqandtogetherhad no[providers.*]sectionat all, so all three were unreachable despite being registered.
Pools, with a failure taxonomy rather than one retry rule. A 429 cools
down briefly; an empty wallet cools down for minutes and records a balance of
zero; a 5xx retries once in place before moving, since it is usually one bad
node and moving would discard a warm prompt cache; a bad key benches the
target for the process; a 400 does not fail over at all; a refusal does not
either, unless you opt in — retrying a refusal elsewhere is "shop until
someone says yes". A prompt that overflows the context window is treated as a
routing signal, not an error.
Seven selection strategies. Three of them needed a judgement call about
missing data, and in each case treating "unknown" as a number inverts the
strategy: an unpriced target ranks after every priced one; an unknown
balance ranks after a known one but ahead of a known-empty one;
lowest_latencyexplores an unmeasured target first.Lanes and classifiers. A classifier names a lane, never a model. Nine
classifiers, ordered cheapest-first and then instructions-before-guesses —
both halves matter, because
hintandheuristicare both free and runningthe guess first would override what the caller asked for. Authority has four
levels, since a routing marker found in content may have arrived from a
retrieved document, where
[[lane:deep]]would be a one-line promptinjection.
Cascades (opt-in) and the privacy path, where the documentation says
plainly that redaction is a mitigation and routing is the guarantee.
Proxy mode:
llmcore-bridge proxygives an unmodified agent harness all ofthe above through two environment variables. It refuses a non-loopback bind
without a bearer token, because the process holds every provider credential in
the config.
Colab runtime backend
Sizing is read-only and free and prints its own arithmetic. Every unknown
rounds toward needing more: overestimating buys a bigger GPU, underestimating
OOMs on the VM after billing has started.
The backend enforces four rules in code — state before compute, fail closed,
never connect to a session that does not exist, deadlines at creation — and
llmcore-runtimesexists for the commands you need when something has gonewrong.
Bugs found and fixed
the wrong way: provider type
geminimapped to agemininamespace, butcards are filed under
google/. Nothing raised — the lookup returned None,None means unknown, and every caller handles unknown quietly. So
lowest_costcould not price any Gemini target. There is now an audit test.thinking.budget_tokenswith a 400 on 4.6+ whilepre-4.6 requires it.
effortis now generation-agnostic.~/.llmcore/runtimes, leaving aphantom entry claiming a GPU VM was running — the exact false signal that
directory exists to prevent.
.envand reach real vendors.unknownand retried the same target; it has its ownkind now, because a missing model is deterministic here but a peer serves a
different model.
prompt_tokenscame back 0 after a failover; the proxy called a method thatdoes not exist; a three-group phone number was only partly redacted.
Where I was wrong, corrected in place
The spec claimed local classification adds "<50 ms p50 on CPU". Measured:
191 ms for 2 lanes, 246 for 5, 314 for 9 — about 5x out, which is why the
gate said measured, not assumed. The spec now says so, and the encoder is
documented as a batch/agent feature.
The same model's raw confidence is meaningless without reading it against
chance (0.20 across five lanes is exactly uniform), so confidence is now
chance-corrected.
Config audit
All 660 keys checked. All 48
[routing]and 14[runtimes]keys are wired.Two pre-existing keys promise behaviour that does not exist and are now marked
NOT ENFORCED rather than quietly removed:
llmcore.admin_api_keyclaimed to protect admin endpoints "such as liveconfiguration reloading" — nothing reads it, so anyone who set it believing
their reload endpoint was guarded was wrong.
context_management.minimum_history_messages.What is not done
means provisioning a real GPU VM and spending your compute units — not mine
to authorize. Everything short of spending is tested.
still answers "Your credit balance is too low".
cache gclists but does not delete; the agent-lens migration guide is notwritten, since it would tell another project to depend on an unproven path.
real traffic, and any number would be invented. That is the largest honest
gap in the subsystem.
Validation
Full suite: 6564 passed, 87 skipped. The only exclusions need a Postgres
server and fail the same way on
main.Live against real vendors: failover over a nonexistent model,
lowest_costpicking the cheaper of two vendors from card pricing, a 700k-char prompt
skipping a 128k window for a 1M one, autoprovisioning an unconfigured
provider, the proxy driven by the real
openaiSDK, and sizing against thelive Hugging Face Hub.
🤖 Generated with Claude Code