feat(providers)!: AI SDK-shaped provider factories (RFC-0036 Part I, Rust) - #214
Draft
cunninghamcard-bit wants to merge 54 commits into
Draft
cunninghamcard-bit wants to merge 54 commits into
cunninghamcard-bit wants to merge 54 commits into
Conversation
The AI SDK alignment design (docs/aisdk-architecture-alignment.md §9.3)
requires validating aimux providers against what the AI SDK sends and
returns for the same input under a mocked transport. This adds the SDK
half: scripts/aisdk-fixtures runs the official @ai-sdk packages
(ai 7.0.127, provider 4.0.21, provider-utils 5.0.53, openai 4.0.83,
openai-compatible 3.0.62, anthropic 4.0.71, google 4.0.87, pinned exactly)
with a recording fetch and writes fixtures/aisdk/<package>/<case>.json:
request (URL, headers with credentials redacted, body), canned response,
and the SDK's normalized result or thrown error.
28 cases across openai, openai-compatible (name "groq"), anthropic and
google pin the behaviours the Rust provider rewrite must match, among them:
an explicit empty apiKey is sent verbatim and never falls back to the env
var; a missing key fails at call time with AI_LoadAPIKeyError and no
request; createAnthropic({ name }) sets model.provider to the name verbatim
(default "anthropic.messages") and reads providerOptions under both the
canonical and the custom key, custom winning; openai-compatible derives
its providerOptions key from the first segment of model.provider and
spreads unknown fields under that key into the request body; header
merging is case-insensitive with call-level headers overriding provider
headers.
Rust tests consume these in the provider factory rewrite; regenerate with
`cd scripts/aisdk-fixtures && npm ci && node record.mjs`.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
… call-level retry AI SDK alignment design (docs/aisdk-architecture-alignment.md §3.2 / §4.5, impact map S1-1, S1-2, S1-3, S1-5). First commit group (A0) of the provider-factory rewrite: the primitives every createXxx(settings) factory needs, with no change to the factories themselves yet. Transport (S1-1) - `Fetch` trait with `FetchRequest` / `FetchResponse` / `FetchError`; `ReqwestFetch` carries the former `shared_client()` machinery (per-runtime client sharding, pool settings, proxy config); `default_fetch()` is resolved per request, never frozen at factory time. - `HttpRequest.fetch: Option<FetchFunction>` threads an injected transport through `post_*_to_api` / `get_from_api`; `send_request_once` sends through it. The SSRF download guard keeps its pinned-DNS client as `PinnedFetch` and ignores the injected fetch when `validate_url` is set (D26). - `SigV4Fetch` decorator signs the final request bytes (same algorithm and test vectors as `bedrock/sigv4.rs`; providers switch to it in A4). - `WsConnector` injection point on `WebSocketRequest`; `TungsteniteConnector` is today's behaviour. - `catalogue.rs` no longer builds its own reqwest client. Settings (S1-2, S1-3) - `Resolvable<T>` (Value / Fn / AsyncFn / Future) and `HeadersFn`; `combine_headers` (lower-cased names, later layers override, `None` removes) and `normalize_headers`; requests insert headers instead of appending, so a credential header is never sent twice. - `load_api_key(Some(s))` returns `s` verbatim, including the empty string; only `None` falls back to the environment (pinned by the AI SDK fixture `openai/empty-api-key`). Missing values are `AiMuxError::LoadApiKey` / `LoadSetting`; `load_setting` / `load_optional_setting` added. The FFI, Node and Python mappings route the two new variants to the existing invalid-argument code until A5 adds dedicated codes. Retry (S1-5, D14) - `prepare_retries(max_retries, abort)` with the constant defaults 2 / 2000 ms / x2; `RetryConfig` is deleted along with every provider config field, `with_retry_config` builder and `retry_config()` trait method. Retry is a call-level concern: the nine core operations read only the caller's `max_retries`, provider-internal list/files/poll sites use the default budget for now (A4 applies the §4.5 table). Not in this commit: RecordingFetch (S1-4), the Provider trait reshape (A1), provider factories (A2–A4), FFI/binding error codes (A5). Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…rding, and the boundary gate
Second commit group of the provider-factory rewrite
(docs/aisdk-architecture-alignment.md §3.1, §3.3, §4.3, §7).
Provider trait (aimux-core/src/provider.rs)
- `Provider` now has the `ProviderV4` shape: language / embedding / image
are required and return `NoSuchModel { model_type }` when a vendor has
no such modality; transcription / speech / reranking / files are
optional (`None` = not offered); video and search stay as aimux
extensions. All constructors return `Arc<dyn Model>`.
- `Provider::name()` is gone: the registered name belongs to whoever
holds the registry, and each model reports `provider()` itself.
- `specification_version()` is removed from the nine model traits and
from `TraceLayer`; `LanguageModel::config_snapshot()` is removed and
`supported_urls()` added.
- `list_models` moves to a separate `ProviderDiscovery` trait (no
`ProviderV4` equivalent; most vendors cannot list). `delegate_list_models!`
emits that impl; `provider_discovery(name, ..)` resolves a discovery
handle for FFI / Node / Python.
Recording (RECORDING_SCHEMA = 3)
- `ProviderRecord` is identity only: `provider_id`, `provider`, `model_id`.
Configuration (base URL, key source, profile, provider options, retry)
is no longer recorded; schema-2 files are rejected on deserialization.
- `rebuild_provider` rebuilds from `provider_id` + `model_id` through the
registry; `is_openai_compatible_provider` is deleted. Native-protocol
packages return `NoSuchProvider` until they join the registry (A4).
Call-level overrides
- `CallOptions.body_overrides` is deleted everywhere (core, openai and
anthropic converters, FFI, Node, Python, aimux-web playground, wire
fixture). Provider-level overrides are applied to the finished body
until `transform_request_body` replaces them (A3).
- `ProviderOptions` / `config_json` / Node `ProviderConfig` reject
`max_retries` and `body_overrides` with `InvalidArgument` instead of
silently ignoring them.
Boundary gate
- `scripts/check_provider_boundaries.sh` (wired into the contract-tests
CI job) fails on provider code reading `options.max_retries` /
`timeout` / `session_id` / `call_id` outside `HttpRequest::new`, and on
any removed name coming back. Eleven hand-built `HttpRequest` literals
now go through `HttpRequest::new`; the eight poll-loop providers use
`prepare_retries(None, ..)` constants for their status GETs.
Tests: new `provider_trait_test.rs` in core and providers (shape, error
mapping, discovery); `config_snapshot_test.rs` deleted; overlay test now
proves the override through a wiremock request; body-override tests
exercise the provider-level path; Node `provider_config.test.ts` covers
the rejection.
Not yet done here (later groups): `api_key_source` fields and
`ExternalProviderEntry.max_retries` / `body_overrides` still exist
unread (A3); `provider_id` is the first segment of `provider()` until the
registry name fills it in (A3); Go / Java / Kotlin / Swift / Flutter still
declare a call-level `body_overrides` field that core ignores (A5).
Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
Third commit group of the provider-factory rewrite
(docs/aisdk-architecture-alignment.md §3.2, §3.3, §4.1, §7), modelled on
`createOpenAI` in @ai-sdk/openai 4.0.83 and checked against the recorded
fixtures under fixtures/aisdk/openai/.
Native OpenAI package
- `OpenAIProviderSettings { base_url, api_key: Option<Resolvable<String>>,
organization, project, headers: Option<HeaderMapOpt>, name, fetch,
transform_request_body }` and `create_openai(settings)`. The factory
only validates `base_url` and fixes `name` (default "openai"); the key
is loaded on every request: `None` reads OPENAI_API_KEY, `Some("")` is
sent verbatim, a missing key fails the call with `LoadApiKey`, not the
factory. `openai()` is the infallible `OnceLock` default instance.
- `OpenAIModelConfig` (crate-private, no getters, no snapshot) carries
`provider`, a `url()` closure, an async `headers` resolver (credential,
then organization/project, then user headers via `combine_headers`),
`fetch`, `supported_urls`, `transform_request_body` and the base URL
for the credentialed-origin guard. Chat, Responses, embedding, image,
speech, transcription and files models read only this config.
- `provider()` is `"{name}.{method}"` (`openai.chat`, `openai.responses`,
`proxy.chat` for `name: "proxy"`); the providerOptions namespace stays
the fixed `openai` key, as upstream.
- `transform_request_body` is the only provider-level body hook; it runs
once on the finished JSON body (multipart untouched). `OpenAIConfig`'s
old `body_overrides` becomes such a closure at the compat boundary.
- `Resolvable::Future` is awaited once, `Resolvable::AsyncFn` on every
request (counter tests).
Compat consumers (transitional until A3)
- `OpenAIConfig` and the new `OpenAIConfigProvider` move to
`aimux-providers/src/openai_legacy.rs`; they keep the builder API for
the 34 thin wrappers, the registry (`provider.rs`), Codex, xAI and
Hugging Face, and feed the same `OpenAIModelConfig` through
`into_model_config`. Removed from `OpenAIConfig`: getters, `from_env`,
`with_api_key_source`, `with_body_overrides`, the `api_key_source`
field. `apply_body_overrides` / `deep_merge_json` move to
`body_merge.rs` for Anthropic until A4.
- Compat `provider()` strings are unchanged (`"groq"`), except files,
which now report `"{provider}.files"`.
Constructor sites: the twelve `aimux_openai_*` FFI symbols, the Node and
Python OpenAI constructors, aimux-cli and aimux-web probes go through
`create_openai`; symbol names are unchanged. CLI / web cache probes
filter traces by `model.provider()`.
Boundary gate: rule 3 keeps `config.api_key` / `config.base_url` /
`body_overrides` / `api_key_source` / `from_env` out of
`aimux-providers/src/openai/`.
Also: an invalid header value no longer echoes the value in the error
(it could contain an API key).
Tests: new `openai_factory_test.rs` asserts every fixture in
fixtures/aisdk/openai/ (URL, method, headers, body, `provider()`, stream)
and fails when a fixture has no case; eleven openai test files moved to
`create_openai`; `body_overrides_test` covers `transform_request_body`.
Left for later groups: delete `OpenAIConfig` / `OpenAICompatProfile` /
`OpenAIConfigProvider` and the thin wrappers, move compat identity to
`"{name}.{method}"` with an explicit dialect flag (A3); Anthropic
`transform_request_body`, Azure `fetch` / `credentialed_origin`, Codex /
xAI / Hugging Face off `OpenAIConfig` (A4); `OPENAI_BASE_URL`, structured
env references, the recorded `"openai"` → `"openai.chat"` consumers and
the CHANGELOG entry (A5).
Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…and registry-generated presets
Fourth commit group of the provider-factory rewrite
(docs/aisdk-architecture-alignment.md §3.3, §4.1, D9, D12, D20), modelled
on @ai-sdk/openai-compatible 3.0.62 and checked against the recorded
fixtures under fixtures/aisdk/openai-compatible/.
openai-compatible (aimux-providers/src/openai_compatible/)
- `OpenAICompatibleProviderSettings { name, base_url, api_key, headers,
query_params, fetch, include_usage, supports_structured_outputs,
supports_multi_part_tool_content, transform_request_body }` and
`create_openai_compatible`. `name` and `base_url` are required and
fixed at the factory; the key resolves per request; `None` sends no
`Authorization` (local servers).
- Chat, embedding and image models (the upstream set) with
`provider() = "{name}.{method}"`. providerOptions are read from
`openaiCompatible`, then the name, then its camelCase form; unknown
fields under the provider's own namespace pass through to the body in
the upstream position; metadata uses the name as key.
- Dialect hooks (`pub(crate)`) replace `OpenAICompatProfile`: usage
conversion, tool and response-format preparation, file-part policy,
`max_tokens_key`, stream usage key. No `Box::leak`.
- `list_models` is one request, no retry.
Groq and DeepSeek (groq/, deepseek/)
- Own packages on the compat internals: `create_groq` / `groq()` with
`x_groq` stream usage, no `top_k`, `max_completion_tokens`, Groq
structured-output rules, browser_search and the Groq file-part
rejection; `create_deepseek` / `deepseek()` with cache-read usage.
`provider()` is `groq.chat` / `deepseek.chat`; providerOptions and
metadata keys are `groq` / `deepseek`. The eight `== "groq"` branches
in openai/convert.rs are gone; the native OpenAI package knows nothing
about other vendors.
Presets (scripts/gen_presets.py → aimux-providers/src/presets/, committed)
- `provider_registry.json` now has 283 rows (251 + the 32 former wrapper
vendors) with `auth: api_key | none`, `base_url_env`, `params` with
`{param}` templates (env / default / derived host maps for Vertex
locations) and `family`. Each row generates a `PresetDescriptor`,
`create_<name>(PresetSettings)` and an infallible `<name>()`.
- `auth: none` resolves no key and injects no placeholder; there is no
`PLACEHOLDER_API_KEY` any more. Template parameters are only the
declared ones, host parameters reject `/`, `@`, `:`, `?`, and an
unexpanded placeholder is an error. Unknown preset names are
`NoSuchProvider`; nothing falls back to OpenAI.
- `provider.rs` resolves names through the preset table and the overlay
table; `ExternalProviderEntry` rejects `max_retries` and
`body_overrides`. `gen_presets.py --check` runs in the contract-tests
CI job; `gen_providers_doc.py` reads the new columns.
Deleted: `openai_legacy.rs` (`OpenAIConfig`, `OpenAIConfigProvider`),
`OpenAICompatProfile`, the 22 local/cloud thin wrappers (incl.
bedrock_mantle, openrouter) and the 10 `vertex_ai_*_models.rs` files.
Codex, xAI, Hugging Face and Azure compile through a crate-private
`StaticBearerConfig` until their own factories (A4).
Behaviour changes recorded here: the native OpenAI package warns on and
drops `top_k` (upstream does not send it); registry and preset providers
expose chat / embedding / image only; compat providers no longer read
the `openai` providerOptions key; `providerOptions.deepseek` is honoured
(`reasoningEffort`, `thinking`); `litellm_proxy` reads
`LITELLM_PROXY_BASE_URL` (the old wrapper read the API-key variable as a
URL).
Boundary gate: rule 4 forbids provider-name comparisons inside the
shared packages, rule 5 keeps the removed names removed.
Tests: `openai_compatible_factory_test.rs` (every compat fixture, plus
a directory-listing guard), `presets_test.rs` (all 283 rows create
without env, keyless rows send no Authorization, template and parameter
errors, Vertex hosts), `deepseek_test.rs`, the Groq package module;
about twenty test files moved from wrapper types to presets. Cassette
recordings are untouched.
Left for later groups: Codex / xAI / Hugging Face / Azure factories and
removal of `StaticBearerConfig` (A4); Anthropic off `body_merge.rs`
(A4a); `params` in the FFI / Node / Python configs, bindings tests on
provider names, docs and the CHANGELOG entry (A5).
Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…reuse the Anthropic core
Fifth commit group of the provider-factory rewrite
(docs/aisdk-architecture-alignment.md §3.3, §4.1, D-d, D-h, D-i),
modelled on `createAnthropic` in @ai-sdk/anthropic 4.0.71 and checked
against the recorded fixtures under fixtures/aisdk/anthropic/.
anthropic/
- `AnthropicProviderSettings { base_url, api_key, auth_token, headers,
name, fetch, transform_request_body }`, `create_anthropic` and the
infallible `anthropic()`. The factory validates `base_url` and rejects
`api_key` together with `auth_token`; the credential resolves per
request (`None` reads ANTHROPIC_API_KEY, `Some("")` is sent verbatim,
`auth_token` sends only `Authorization: Bearer`). A missing key fails
the call with `LoadApiKey`.
- Private `AnthropicModelConfig` with hooks (URL, body, headers, error
handler, feature flags) feeds one `AnthropicMessagesModel`; the files
model reads the same config.
- `provider()` is the name verbatim: `anthropic.messages` by default,
`proxy` for `name: "proxy"`; files derive `{name minus .messages}.files`.
- providerOptions are read from the canonical `anthropic` key merged
with the custom first segment (custom wins); response metadata is
written under the custom key. The twenty hard-coded `"anthropic"`
sites go through `anthropic/options.rs`.
- Base URL follows upstream: it includes `/v1` and the endpoint is
`{base}/messages`; only the bare `https://api.anthropic.com` is
rewritten. `ANTHROPIC_BASE_URL` is no longer read (no `from_env`).
- Request body: `stream` is omitted on non-streaming calls;
`providerOptions.anthropic.metadata.userId` maps to `metadata.user_id`;
`anthropic-beta` is sent on every host; `supported_urls` declares the
upstream image and PDF patterns.
- `list_models` is one exchange without retry.
anthropic_aws/
- `AnthropicAwsProviderSettings` with `ApiKey(Resolvable)` or
`SigV4(Resolvable<AwsCredentials>)`; signing is now the `SigV4Fetch`
transport decorator (service `aws-external-anthropic`) over the final
bytes, the model no longer signs, and `list_models` works with SigV4.
The provider string stays `anthropic-aws`.
vertex/anthropic_model.rs
- `VertexAnthropicModel` is `AnthropicMessagesModel` with Vertex hooks
(rawPredict / streamRawPredict URL, `anthropic_version` in the body,
bearer or x-goog-api-key headers, Google error handler, empty
`supported_urls`, structured outputs and strict tools off); 591 → 117
lines. Provider string `googleVertex.anthropic.messages`.
Deleted: `AnthropicConfig`, `AnthropicConfigBuilder`,
`AnthropicAwsProviderConfig`, `api_key_source`, `from_env`, `with_*`,
and the family's `body_overrides` (replaced by `transform_request_body`).
Constructor sites in aimux-ffi, Node, Python, aimux-cli and aimux-web go
through the factories; symbol names are unchanged. Boundary gate rule 6
covers the Anthropic family.
Tests: `anthropic_factory_test.rs` replays every fixture in
fixtures/aisdk/anthropic/ (URL, method, headers, body, `provider()`,
stream) with a directory-listing guard, plus credential timing, custom
namespace, SigV4-over-final-bytes and Vertex-envelope tests; fifteen
test files moved to the factories. Cassette recordings are untouched.
Left for later groups: Vertex factory with real ADC / express headers
and the `googleVertex` namespace (4b); Bedrock on `SigV4Fetch` and the
`bedrock/sigv4.rs` shim removal (4b); files upload retry (4d);
`body_merge.rs` and its two tests, the base-URL release note and the
result-level providerMetadata (A5).
Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…on_bedrock factories
Sixth commit group of the provider-factory rewrite
(docs/aisdk-architecture-alignment.md §3.3, §4.1, D-h, D11), modelled on
@ai-sdk/google 4.0.87, @ai-sdk/google-vertex and @ai-sdk/amazon-bedrock
and checked against the recorded fixtures under fixtures/aisdk/google/.
google/
- `GoogleProviderSettings`, `create_google`, the infallible `google()`;
chat, embedding, image, video and files models read a private
per-model config (`shared/exchange.rs`). The key resolves per request
(`x-goog-api-key`, GOOGLE_GENERATIVE_AI_API_KEY); `provider()` is the
name (`google.generative-ai`) with `{name}.files` / `{name}.video`.
- Request bodies now match upstream: `generationConfig` is always sent,
function tools go out as `parametersJsonSchema`, `providerOptions.google`
(thinkingConfig, responseModalities, audioTimestamp, mediaResolution,
imageConfig, safetySettings, cachedContent, labels, serviceTier,
retrievalConfig) is mapped; response `modelId` comes from
`modelVersion`. providerOptions are read from and written to `google`
only (`google/options.rs`). `list_models` and the files upload are
single exchanges.
vertex/
- `VertexProviderSettings { api_key (express), project, location,
base_url, headers, access_token, fetch, transform_request_body }`,
`create_google_vertex`, `google_vertex()`. Express versus standard
mode, project / location (`LoadSetting`) and the bearer token
(`GOOGLE_VERTEX_ACCESS_TOKEN`, or any `Resolvable`) resolve per
request; location must be one DNS label and the host follows the
upstream global / rep / regional rule; Gemini models use `v1beta1`,
Anthropic-on-Vertex `v1`. `provider()` is `google.vertex` (plus
`.video`, `.transcription`, and `googleVertex.anthropic.messages`).
providerOptions: `googleVertex` (canonical only; the legacy `vertex`
key is neither read nor written), then `google` for the shared Gemini
model.
bedrock/
- `AmazonBedrockProviderSettings { region, api_key, access_key_id,
secret_access_key, session_token, base_url, headers, fetch,
credential_provider }`, `create_amazon_bedrock`, `amazon_bedrock()`.
AWS_BEARER_TOKEN_BEDROCK selects bearer auth; otherwise `SigV4Fetch`
signs the final bytes with credentials resolved per request
(`LoadSetting` for AWS_REGION / AWS_ACCESS_KEY_ID /
AWS_SECRET_ACCESS_KEY). `provider()` is `amazon-bedrock` for chat,
embedding, image and reranking; providerOptions use `amazonBedrock`
only (the legacy `bedrock` key is neither read nor written).
`bedrock/sigv4.rs` is deleted; `aws_polly` signs through `SigV4Fetch`.
Deleted: `GoogleConfig`, `VertexProviderConfig`, `VertexAuth`,
`BedrockProviderConfig`, `BedrockAuth`, their `from_env` / `with_*` /
`api_key_source`. Constructor sites in aimux-ffi (14), Node, Python,
aimux-cli and aimux-web go through the factories; symbols unchanged.
Boundary gate rule 7 keeps the builder-era names, `bedrock::sigv4` and
stray namespace literals out.
`anthropic_aws` keeps its name: it is Claude Platform on AWS
(`aws-external-anthropic`), not Bedrock InvokeModel; aimux has no
Bedrock-Anthropic InvokeModel provider and this PR does not add one.
Tests: `google_factory_test.rs` (every fixture, directory-listing guard),
`vertex_factory_test.rs` (express and standard paths, three hosts,
invalid location, missing settings), `bedrock_factory_test.rs` (bearer
without signature, SigV4 over the final bytes, provider strings), a
shared scripted `Fetch` mock in tests/common; about twenty-eight test
files moved to the factories. Cassette recordings are untouched.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…Face, Codex, open_responses, Voyage and ElevenLabs
Seventh commit group of the provider-factory rewrite
(docs/aisdk-architecture-alignment.md §3.3, §4.1, D-d, D11, D13), each
package modelled on its @ai-sdk counterpart where one exists.
- Every package gets `XxxProviderSettings`, `create_xxx`, an infallible
`xxx()` and a private per-model config; credentials resolve per
request through the shared helpers (`None` reads the package's env
var, `Some("")` is sent verbatim, a missing key fails the call with
`LoadApiKey`). `provider()` strings follow upstream: `azure.chat` /
`azure.responses` / `azure.embeddings` / `azure.image` /
`azure.transcription` / `azure.speech`, `xai.responses`,
`mistral.chat` / `mistral.embedding`, `cohere.chat` /
`cohere.textEmbedding` / `cohere.reranking`, `huggingface.responses`,
`codex.responses`, `{name}.responses` for open_responses,
`voyage.embedding` / `voyage.reranking`, `elevenlabs.speech` /
`elevenlabs.transcription`. All but Azure accept a `name` override.
- Azure reuses the OpenAI chat / Responses / embedding / image /
transcription / speech models through the injected config: `api-key`
or bearer token (`token_provider` AsyncFn) per request, resource name
from AZURE_RESOURCE_NAME (`LoadSetting`), upstream defaults
(`api_version = "v1"`, `/v1{path}?api-version=`), deployment-based
URLs on request. The Responses model reads providerOptions `azure`
then `openai` and writes `azure`.
- xAI and Hugging Face default to Responses (as upstream); their
chat-completions models remain reachable as the `chat_completions(id)`
extension. Codex keeps RFC-0018's two modes (`ApiKey` with
CODEX_API_KEY fallback, `ChatGptAccount { token, account_id }`), the
stateless refresh, 401 → `TokenExpired`, and `store: false` as a
package rule. ElevenLabs' WebSocket handshake headers come from the
resolved headers and the connector is injectable. open_responses takes
`base_url` and `Resolvable` headers; `api_key: None` sends nothing.
- Deleted: the nine `XxxConfig` types and builders, `AzureAuth`,
`TokenProvider`, `AzureModel`, `azure/{model,responses}.rs`,
`api_key_source` / `from_env` in these packages, and the transitional
`StaticBearerConfig`. `OpenAIModelConfig` now anchors credentialed
headers to the request URL (Azure's host is per request) and carries
a `ResponsesProfile`; namespace literals live in each package's
`options.rs`.
- Constructor sites in aimux-ffi (12), Node (6), Python (6), aimux-cli
and aimux-web go through the factories; symbols unchanged. Boundary
gate rule 8 covers these packages.
Not ported from upstream Azure: the deepseek and completion models, the
MAI / speech endpoints and the Foundry item type.
Tests: `vendor_factories_test.rs` (28 tests over a scripted `Fetch`:
credential timing, empty key, default instances, provider strings per
package) and 27 adapted test files. Cassette recordings are untouched.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…ge-constant poll loops; uploads are not retried
Eighth commit group of the provider-factory rewrite
(docs/aisdk-architecture-alignment.md §3.3, §4.5, §9.3).
- Every single-modality package (serper, prodia, deepgram, you_com, luma,
lmnt, klingai, dataforseo, replicate, fal, recraft, aws_polly, jina_ai,
gladia, tavily, linkup, tinyfish, black_forest_labs, assemblyai, revai,
hume, cartesia, searxng, parallel_ai, firecrawl, runwayml, exa_ai,
stability, google_pse) gets `XxxProviderSettings`, `create_xxx`, an
infallible `xxx()` and `XxxProvider` on the shared `EndpointConfig`;
credentials resolve per request (missing key → `LoadApiKey`, missing
URL or AWS setting → `LoadSetting`). `provider()` is `"{name}.{method}"`
(`luma.image`, `deepgram.transcription`, `tavily.search`,
`amazon-polly.speech`, `blackForestLabs.image`, …). providerOptions
keys live in one `options.rs` per package. The old `XxxConfig`,
`from_env` and `with_*` are gone; no `*Config` struct remains in
aimux-providers.
- Poll loops (luma, black_forest_labs, fal, gladia, assemblyai, revai)
run on package constants through `shared/poll.rs`: a fixed interval
and attempt cap (previously unbounded loops now stop after 6000 × 100
ms), a transient poll error spends one attempt, exhaustion returns the
last error or `Timeout`, cancellation is checked before every wait.
The job-creating request is sent exactly once; nothing re-submits.
Only the interval can be overridden, through the package namespace
(`pollIntervalMillis` for luma as upstream, `pollIntervalMs`
elsewhere); pacing keys are stripped before the body is forwarded.
Download stages keep a bounded three-try retry. `aimux_core::retry`
is no longer used anywhere in aimux-providers (boundary rule 9).
- Files uploads (OpenAI, Anthropic, Google) are single exchanges.
- Constructor sites (`aimux_tavily_search_new*`, Node and Python
Tavily) go through the factory; symbols unchanged. Boundary rule 10
covers the 28 packages.
Behaviour changes recorded here: cartesia `version` is a fixed header
(override via `headers`); runway poll pacing is constant; google_pse
`cx` precedence is settings, providerOptions, then GOOGLE_CSE_ID per
request; DataForSEO missing credentials report `LoadApiKey`; searxng
accepts an optional bearer token.
Tests: `single_modality_factories_test.rs` (defaults, missing key,
provider strings), `poll_stage_test.rs` (exhaustion never re-submits per
package, interval override, abort, ignored attempt keys, stripped pacing
keys), upload-not-retried tests in the three files packages, unit tests
in `shared/poll.rs`; thirty-five test files converted. Cassette
recordings are untouched.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…binding configs and error classes, CHANGELOG and docs
Ninth and last commit group of the provider-factory rewrite
(docs/aisdk-architecture-alignment.md §5, §6.2, §8, §9.2, §9.4).
C ABI
- `AIMUX_E_LOAD_API_KEY` (18) and `AIMUX_E_LOAD_SETTING` (19) for
`AiMuxError::LoadApiKey` / `LoadSetting`; the env var travels in
`aimux_error_provider_code` and the description or setting name in
`aimux_error_provider_message`. `aimux-error.h` / `aimux-ffi.h` synced,
with a header-versus-source parity test. `provider_handle_new` and
`register_providers` reject `max_retries` / `body_overrides` in
`config_json`.
Node / Python
- `LoadAPIKeyError` and `LoadSettingError` classes (`envVar` /
`description` / `settingName`; snake_case in Python). Every Node native
constructor goes through one `native_config()` that rejects
`maxRetries` / `bodyOverrides` and applies `headers`;
`ProviderConfig.params` carries preset template parameters.
`index.d.ts` regenerated; both bindings are fmt- and clippy-clean; the
ava and pytest suites pass.
Go / Java / Kotlin / Swift / Dart
- Codes 18 / 19 with an `EnvVar` field; the call-level `body_overrides`
field (which core no longer reads) is deleted; Go and Dart
`ProviderConfig` drop `max_retries` / `body_overrides` and gain
`params`. Go and Java suites run here; Kotlin, Swift and Dart edits
are symmetrical and unverified in this container.
Residue
- `body_merge.rs` and `body_overrides_test.rs` are gone
(`transform_request_body_test.rs` replaces them); `hmac` dropped,
`sha2` / `hex` are dev-dependencies; Codex subscription default URL
fixed to `…/backend-api/codex`; Bedrock `list_models` host fixed to
`bedrock.{region}.amazonaws.com`; Anthropic result-level
`providerMetadata` (usage, stopSequence, container) implemented and
checked against the fixtures; boundary gate also forbids
`api_key_source` and the body-merge helpers.
Docs
- CHANGELOG `[Unreleased]` gains the Rust, C ABI, Node / Python and
other-binding breaking lists for the whole rewrite. README,
CONTRIBUTING, docs/api/*, docs/API.md, the provider-config manual, the
Codex guide, the error model, the request-pipeline doc (retry is a
call-level constant) and `rfc/0036-aisdk-architecture-alignment.md` §5
(no `specification_version` / `name()`, `.call()` form, OpenAI default
stays chat, Bedrock InvokeModel and unported Azure models, video
`poll_config`, native `rebuild_provider`) are updated.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
… in §5 Same change as the docs/rfc-0036-aisdk-alignment follow-up (no ABI coexistence promise after the cutover; §3.3 / S4-7 wording for the xAI and Hugging Face Chat Completions extension), plus the §5 product difference row for `chat_completions(id)` that this branch introduces. The shared part becomes a no-op once the docs PR lands on master. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…s models The AI SDK's xAI package (5.0.12) and Hugging Face package serve the Responses API only. The Chat Completions models that 7c54cfd kept as a `chat_completions(id)` extension outside the `Provider` trait are removed: `XaiModel`, its wire types and chat converter, and the Hugging Face chat constructor. `xai/convert.rs` keeps only the helpers the Responses converter uses. - Tests of the removed model go with it (the chat-level modules of `xai_test.rs`, the xAI / Hugging Face rows of the OpenAI-compatible suites). The xai-error and xai-provider cases now run against `responses(id)`. The 16 Hugging Face cassettes are unchanged and replay through the OpenAI-compatible package in `conformance_test.rs`. - The xAI Responses model sends `top_k` and warns for `frequencyPenalty` and `presencePenalty`, as upstream does (found while moving the unsupported-parameter case). - docs / ROADMAP: same wording change as the docs PR follow-up (the extension is withdrawn; superseded ROADMAP lines rewritten), and the §5 product-difference row for the extension is dropped. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
* docs: full-chain AI SDK alignment design and impact map Implementation design under RFC-0036 (positioning and layered architecture), placed in docs/ rather than rfc/: it does not take an RFC number and does not re-decide direction. Compared with the earlier RFC draft of the same content: - Header states the document's relation to RFC-0036 and how it takes effect (on merge; later changes edit the file). No "Accepted" status. - §0.7 lists, item by item, which ROADMAP / RFC-0036 commitments the design keeps, adjusts or postpones: L2 choice (AI SDK V4 shape, survey RFC withdrawn), the one-shot breaking switch versus the old-ABI coexistence promise, ops protocol / stdio / L0 passthrough (kept, after the switch), B track, #174/#175 auth, #167/#179 replay, #185. ROADMAP.md and the positioning RFC carry notes pointing at that table where their wording is adjusted. - §0.4 Q4, D18, §6.4 and §10 now agree: host callbacks are postponed in every language, including the Node TSFN / Python GIL bridges. - Part II's two open items are written as conclusions (no aimux-protocol crate; JSON adaptation rules adopted), and the impact-map §6.4 questions each carry the section that settles them. - docs/README.md indexes both files; links are relative to docs/. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS * docs: drop the ABI coexistence promise after the cutover; record the xAI/HF chat extension Two wording fixes so the design matches the RFC-0036 §12.2 replacement list proposed in #207: - §0.7 and the ROADMAP 0.8 note no longer leave "old exports coexist one minor" to a later decision: the full-chain cutover has no coexistence period, and the old symbols replaced when the ops ABI arrives are removed on that cutover gate (RFC-0039), so the promise is withdrawn rather than deferred. - §3.3 and the rejected-finding row S4-7 state what the implementation does with xAI / Hugging Face Chat Completions: the V4 entry points only offer Responses, and the existing Chat Completions model stays as an explicit `chat_completions(id)` extension outside the `Provider` trait, registered as a product difference. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS * docs: address review findings on the alignment design - complete the truncated §6.4-14 ruling in the impact map (pull-style streaming ABI: not this round, stays with the async C ABI deferral) - annotate S5-1/S5-2 with the main-doc rulings (no 2→3 migration table, no rebuild_provider; live replay is registry + target-ref driven) - mark the Part II ts-rs advice as overruled by §0.5/§6.5 inline - convert 7 markdown links into the local-only aisdk-review checkout to AI:/A: code refs (reference/ is gitignored; links 404 for others) - note the local-only reference baseline in the evidence convention - annotate the 0.6 B1+C2 ROADMAP row (B1 reshaped, C2 shim cancelled) - flatten nested-paren annotations in ROADMAP #185 rows - drop the duplicate H1 title in the impact map; bump the revision date --------- Co-authored-by: chenhaonan <[email protected]> Co-authored-by: Claude Fable 5.1 <[email protected]> Co-authored-by: eric8810 <[email protected]>
…in only for API calls The pooled client followed redirects with reqwest's default policy, which strips only Authorization-style headers on a cross-host hop, so vendor credential headers (x-api-key, x-goog-api-key, api-key, x-amz-security-token, ...) were forwarded to the redirect target, and a SigV4 request was replayed with its old signature. Redirect handling now lives in one place, the hop-by-hop loop that the validated-download path already had in http.rs (#163): - An ordinary API call follows a redirect only while it stays on the origin of the hop that issued it; a cross-origin 3xx is returned as a regular non-2xx response. Nothing is sent to the other origin, neither the request headers nor what a transport decorator would add. - Every followed hop goes through the request's transport again, so a signing transport signs the URL it actually sends. - A transport never follows a redirect: the pooled client is built with no redirect policy, and RedirectPolicy / FetchRequest.redirect / PinnedFetch::unpinned are removed. An injected Fetch gets the same protection because the rule is above it. - Validated downloads behave as before (per-hop validation and pinning, caller headers stripped off the credentialed origin). Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…pi_key in ExternalProviderEntry Debug The Bedrock region was interpolated into the request host unchecked, so a crafted value could change the host. It now must be a single DNS label (the helper the Vertex location check uses) and fails with InvalidArgument before any request is sent. ExternalProviderEntry no longer derives Debug: a manual impl lists the same fields but reports only whether api_key is set. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…tten factory tests Kept: the tests that replay recorded AI SDK fixtures (fixtures/aisdk/*) in the OpenAI, OpenAI-compatible, Anthropic and Google factory tests, with their fixture-directory guard tests and only the helpers they use; one poll test proving the job-creating request is sent once when polling is exhausted. Dropped: the hand-written factory unit tests, the vendor/single-modality/ Vertex/transform-request-body/provider-trait test files, the other poll-stage tests, and the helpers and imports that became unused. Fixture replays and end-to-end behavior are the coverage we want; hand-written unit tests for factories are not. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…istry; drop the generated sources `aimux-providers/src/presets/` (26 generated files, 12,774 lines) and `scripts/gen_presets.py` are removed. `provider_registry.json` is embedded and parsed once into the descriptor table (`preset::entries` / `names` / `lookup`); a preset is created by name with `PresetProvider::create(descriptor, settings)`. The per-name Rust functions (`presets::create_<name>`, `presets::<name>()`) are gone with no replacement, so the registry data is no longer stored twice. - Every check the generator made on a row (key whitelist, auth / env_var pairing, family and auth values, template parameters matching the URL placeholders, duplicate names) now runs when the table is first loaded; an invalid registry panics there and fails the table test. - `presets_test.rs` keeps the table test (all 283 rows load and create without reading the environment) and one end-to-end case each for a keyless row, an unknown name and a rejected template parameter. - CI no longer runs `gen_presets.py --check`; docs and CHANGELOG describe the runtime table. docs/aisdk-architecture-alignment.md carries the same D12 / D20 wording as the docs PR follow-up. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
Keeps SigV4 signing of the final request bytes, the bearer-token path and the region host-label check; the remaining hand-written factory cases are dropped like the other factory test files. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
Takes the design docs as merged in #200 (including the maintainer's review fixes) and keeps this branch's follow-ups on top of them: - xAI / Hugging Face expose Responses only; the Chat Completions extension recorded in the merged doc is withdrawn (§3.3, S4-7, revision note). - Presets are a runtime descriptor table (D12, D20, §0.3, §0.7, §4.1, §5, §6.5). - §3.2: redirects are handled by the helper, same-origin only for API calls; `FetchRequest` has no redirect field. - §0.3 drops NDJSON from aimux-stream; the reference-baseline header notes that fixtures are pinned by fixtures/aisdk/VERSIONS.json. - ROADMAP: the lines that still carried superseded commitments are rewritten (C2 shim, 0.7 / 0.8 version semantics, C1 coexistence, #174 / #175, #167, C4). Where the merged text already annotated a line (B1 + C2, #185, the Part II ts-rs sentence) the merged wording is kept. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
What: `create_provider_registry(providers, options)` returns a
`ProviderRegistry` over a caller-assembled map of provider ids to
providers. A model is addressed as "{provider}:{model}"; the id is split at
the first separator and the rest is passed to the provider unchanged. One
method per modality of the `Provider` trait, plus `files(provider_id)`.
Why: this is the one by-name mechanism the AI SDK has
(`createProviderRegistry` in `ai/src/registry/provider-registry.ts`). Until
now the workspace answered "name -> provider" in three separate places
that cover different vendors (the `provider(name, ..)` function for registry
presets, hand-written matches in the CLI and web tools for native vendors,
and per-vendor FFI constructors). The following commits move all of them
onto this registry.
Behaviour, as upstream: an id without the separator is `NoSuchModel`, an
unregistered provider is `NoSuchProvider`, a modality the provider does not
offer is `NoSuchModel`, the separator is configurable. The registry knows
no provider on its own and never falls back.
Not ported yet: language / image middleware and the `skills` accessor.
Tests: six cases ported from provider-registry.test.ts ([email protected]).
Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…, default_providers) What: - `create_provider(name, settings)`: a name selects a factory, and the same few settings (key, base URL, extra headers, transport) go to whichever factory it is. The name is either a vendor package (`openai`, `anthropic`, `google`, ... 16 of them) or a row of `provider_registry.json`, which goes to the OpenAI-compatible factory. A package wins over a row of the same name. An unknown name is `NoSuchProvider`; there is no fallback. - `default_providers()`: every name with default settings, as the map `create_provider_registry` takes. Nothing is read from the environment and no entry can fail to be created; keys are evaluated per request. - `preset::create(name, settings)` returns an ordinary `OpenAICompatibleProvider` for a registry row. - `Provider::discovery()` (aimux extension, default `None`): the `/models` side of a provider is reachable from the provider, so a registry entry needs no second handle for it. A reference to a provider is a provider, so `&'static` default instances register as they are. Why: the AI SDK ships no vendor list. The reference for a built-in list on top of it is models.dev (what opencode consumes): one row per vendor naming the factory to use and what to pass it, native and compatible vendors in the same table. Until now only the compatible vendors were reachable by name, so the CLI, the web tool and the FFI each carried their own mapping for the native ones. This is that mapping, once. Tests: one end-to-end case, a registry over `default_providers()` serving a package, a registry row and a name both have. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
… create_provider What: `build_model` in `aimux-cli` (probe) and in `aimux-web` is now one call, `create_provider(name, settings)?.language_model(model_id)`. The web console's provider list is the keys of `default_providers()`. Why: both tools carried the same hand-written match (openai / anthropic / google / mistral / xai / cohere each to its own factory, everything else to the by-name `provider()` function) because the by-name function could not reach the vendor packages. With one table for every built-in provider the match has nothing left to do. 164 lines removed, 26 added. Behaviour change: a missing key for a vendor package is no longer checked when the model is built. It fails the request that needs it, with `LoadApiKey` naming the variable, which is when every factory evaluates its key. The provider list now also contains the vendor packages that the hand-written list did not name (azure, amazon_bedrock, google_vertex, ...). Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
What: `rebuild_provider(record, api_key)` creates the recorded provider with `create_provider(provider_id, ..)` and takes its language model. If that model is not the one the recording was made with (the record's `provider` string, e.g. `openai.responses` against the default `openai.chat`), the rebuild is refused with a message pointing at `replay_with_model`. Why: it was the last caller in this crate of the by-name `provider()` function, and it shared that function's blind spot: a recording made with a vendor package (`anthropic`, `google`, ...) was `NoSuchProvider`. Those now rebuild. The refusal exists because the identity-only record cannot select a non-default method, and replaying a Responses recording against the chat endpoint would be silently wrong. Behaviour change: a missing key is no longer reported when the model is rebuilt; it fails the replayed request, like any other call. Tests: the three cases that registered a runtime overlay are replaced by one case for a registry row plus a vendor package, one for the refusal and one for an unknown id. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
The Mistral chat stream built tool-input-start/delta/end and tool-call parts with its own hand-written logic, assuming every tool call arrives complete in one chunk. Upstream (@ai-sdk/mistral mistral-chat-language-model.ts doStream) instead feeds every choice.delta.tool_calls entry to StreamingToolCallTracker, constructed with the model's id generator, and flushes it after the text/reasoning ends and before the finish part. mistral/model.rs now does the same with the existing aimux_provider_utils::StreamingToolCallTracker (as the openai and openai_compatible chat models already do): each delta is forwarded to process_delta, a malformed delta ends the stream with an invalid-response-data error, and flush() runs before Finish. The hand-written per-chunk emission is removed. A small private generate_id supplies ids for calls the server sends without one, standing in for upstream's generateId option. mistral/types.rs: DeltaToolCall now mirrors upstream's chunk schema (index, id and function.name/arguments are all optional), so partial deltas can reach the tracker. Module docs no longer claim tool calls always arrive complete. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
GroqProvider::chat (and call / language_model) used to return the generic
OpenAI-compatible chat model parameterised by a Groq dialect. The AI SDK's
@ai-sdk/groq has its own GroqChatLanguageModel and does not depend on the
compatible package, so the Rust package now has one too, file for file:
- model.rs mirrors groq-chat-language-model.ts: getArgs, doGenerate and
doStream, with streamed tool calls going through the shared
StreamingToolCallTracker (type validation "required", as upstream) and
stream error chunks mapped to a status through getGroqStreamErrorMetadata.
- convert.rs mirrors convert-to-groq-chat-messages.ts.
- prepare_tools.rs mirrors groq-prepare-tools.ts, with
browser_search_models.rs for groq-browser-search-models.ts.
- usage.rs mirrors convert-groq-usage.ts, finish_reason.rs mirrors
map-groq-finish-reason.ts, options.rs mirrors
groq-chat-language-model-options.ts, error.rs mirrors groq-error.ts, and
types.rs holds the response and chunk shapes the model reads.
- mod.rs builds the provider the way the Mistral package does (an
EndpointConfig per model, headers evaluated on every request); the public
surface (create_groq, groq(), GroqProviderSettings, "groq.chat") is
unchanged and GroqProvider::chat now returns GroqChatLanguageModel.
Where the existing package tests pin behaviour that differs from upstream, the
tests win and the model keeps it: the token limit is sent as
max_completion_tokens, the generic openaiCompatible options namespace is read
under the groq one, unknown fields of the groq options go to the body, and
provider metadata {"groq": {}} is reported.
groq/dialect.rs is left in place: the registry preset still builds Groq
through groq::profile() on the compatible model.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
DeepSeekProvider::chat used to return the generic OpenAI-compatible chat
model parameterised by a profile. The AI SDK's @ai-sdk/deepseek has its own
chat model and does not depend on the OpenAI-compatible package, so the Rust
package now has one too: DeepSeekChatLanguageModel, returned by chat(),
call() and language_model(). The provider keeps its public surface
(create_deepseek, deepseek(), DeepSeekProviderSettings, provider() ==
"deepseek.chat") and is assembled like the Mistral provider: a fixed
EndpointConfig, bearer credential headers loaded per request, the shared
exchange for URL, headers, transport and recording, and list_data_models for
discovery. Streamed tool calls go through the shared StreamingToolCallTracker.
New files under aimux-providers/src/deepseek/, each mirroring one upstream
file of packages/deepseek/src/chat/:
- model.rs <- deepseek-chat-language-model.ts (getArgs, doGenerate,
doStream, stream error classification)
- convert.rs <- convert-to-deepseek-chat-messages.ts (also carries the
file part handling of deepseek-file-part-options.ts)
- usage.rs <- convert-to-deepseek-usage.ts
- prepare_tools.rs <- deepseek-prepare-tools.ts
- types.rs <- deepseek-chat-api-types.ts (response and chunk shapes)
- options.rs <- deepseek-chat-language-model-options.ts (chat, message and
file part options)
- finish_reason.rs <- map-deepseek-finish-reason.ts
- is_v4_model.rs <- is-deepseek-v4-model.ts
The files/ module of the upstream package is not ported.
Where the existing acceptance tests pin behaviour that differs from upstream,
the tests win: thinking and reasoningEffort are sent as given with no
normalisation, unknown providerOptions.deepseek fields (user) reach the body,
topK is sent, reasoning is read as a fallback for reasoning_content, no JSON
system message is injected and a JSON schema becomes response_format
json_schema, strict tool flags are passed through without validation, an
assistant message with only tool calls has null content, and provider
metadata carries no cache counters or response ids. The old
deepseek::profile() stays only for the registry presets that still use it.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
There were three ways to turn a provider name into a provider: the old `provider()` family in provider.rs, `PresetProvider::create`, and the new `create_provider`. This commit leaves one. - provider.rs keeps only what the language bindings still call (`provider`, `provider_handle`, the runtime overlay for user-registered providers). `provider_handle` now checks the overlay and otherwise calls `create_provider`. The module is no longer re-exported from the crate root; callers name it as `aimux_providers::provider::...`. It goes away when the bindings expose the Rust surface directly. - Removed: `ResolvedProvider`, `resolve_provider`, `provider_discovery`, `provider_from_env`, `provider_names`, `provider_registry_entry` and the module's unit tests. Model listing is reached through `Provider::discovery()`. - preset.rs: the `PresetProvider` wrapper type is gone. `preset::create` returns a plain `OpenAICompatibleProvider`, the same type `create_openai_compatible` returns, so a preset is only a row of settings. - openai_compatible: drop `resolve_base_url`, which had no caller left. The FFI crate and the Node and Python bindings still import the removed root exports and are updated in the following commits. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…, moonshotai, fireworks, cerebras, baseten, alibaba, deepinfra)
The vendor test files built their models with the by-name `provider(..)`
and `provider_from_env(..)` functions, which are leaving the crate's
public surface. Their `make_provider` helpers and the individual tests now
call `create_provider(name, PresetSettings { api_key, base_url, .. })` and
`.language_model(model_id)` instead. Assertions are unchanged.
The `from_env_fails_without_env_var` test in each file used to assert an
error at creation time. `create_provider` no longer checks the key, so the
test is now async and asserts that the first `do_generate` call fails. No
tests were deleted.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…e, openai_compatible, thin_wrapper, list_models, openrouter, presets, reasoning_map) The by-name `provider(..)`, `provider_handle`, `provider_discovery`, `provider_from_env`, `ProviderOptions`, `provider_registry_entry` and `PresetProvider::create` are leaving the crate's public surface. These test files now build their models with `create_provider(name, PresetSettings)` (and `preset::create(name, settings)` for the OpenAI-compatible presets) followed by `language_model(..)` / `discovery()`. Local helpers (`registry_model`, `preset_model`) carry the change so the call sites keep their assertions. Files with no use of the removed surface (anthropic_prepare_tools, openai_files, google_files, open_responses) are unchanged. Deleted: - list_models_test: `provider_handle_unknown_name` and `discovery_for_unknown_provider_is_no_such_provider` only exercised the removed by-name entry points (the unknown-name error is covered by presets_test through `create_provider`). - thin_wrapper_config_test: the two `from_env_fails_without_env_var` tests (togetherai, vercel) were `#[ignore]`d and asserted a creation-time missing key error that no longer exists. Moved: openrouter `from_env_fails_without_env_var` now asserts `LoadApiKey` on the first request instead of at creation. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
The previous commit gave the Groq package its own chat model but kept four behaviours that upstream does not have, because hand-written tests pinned them. Upstream wins, so in groq/model.rs (groq-chat-language-model.ts) and groq/options.rs (groq-chat-language-model-options.ts): - the token limit is sent as max_tokens, as getArgs does; - provider options are read from the groq namespace only; the generic openaiCompatible namespace is no longer merged in; - unknown fields of the groq options are dropped, not copied into the body; - no provider metadata is reported from doGenerate or doStream, as upstream returns none. A chunk that fails to parse is now reported the way upstream does it: an error part, finish reason "error", and the stream carries on; only a transport failure ends it. An error reported by the very first event still rejects the call, which is the convention of the other own-model packages in this crate and keeps the failure inside Core's retry boundary. tests/groq_test.rs now builds every model through create_groq(..).chat(..) instead of the by-name provider() function, so it exercises the new model. Tests that expected non-upstream behaviour are corrected to the upstream cases (reasoning mapped to low/medium/high, none handled per model, file references raising an unsupported-functionality error, tool_choice "auto", max_tokens, no metadata, groq namespace only) and upstream stream cases are ported: reasoning kept active across empty tool_calls, unparsable chunks, raw chunks, error chunks, and a response without choices. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…set row
Groq and DeepSeek each have their own chat model now, as in the AI SDK
(`@ai-sdk/groq`, `@ai-sdk/deepseek`). The preset table still carried a row
for each, with a `family` field that made the generic OpenAI-compatible
model imitate them through a per-vendor "dialect". That gave two different
models for the same name: `create_provider("groq")` built the Groq
package's model while `preset::create("groq")` built the imitation.
- provider_registry.json: remove the `groq` and `deepseek` rows (281 rows
left) and the `family` field.
- preset.rs: remove `PresetFamily`; every row is assembled with the same
baseline OpenAI-compatible profile.
- Delete `groq/dialect.rs` and `deepseek::profile()`, which only the
preset rows used, and `baseline_usage`, which only the latter used.
- provider.rs: remove a test-only helper whose tests are gone.
- gen_providers_doc.py and docs/api/providers.md follow (the two vendors
are listed under the typed factories).
Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
The previous commits removed the by-name helpers from the root of aimux-providers. The FFI crate and the Node and Python bindings now import what is left from `aimux_providers::provider`, and list models through `Provider::discovery()` on the handle instead of the removed `provider_discovery`. No exported C symbol and no binding-level function changes its name or arguments. FFI tests that asserted an error at creation time for a missing API key are removed: keys are read when a request is made, as in the AI SDK's `loadApiKey`. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…es its settings Three review findings on the single by-name assembly function: - The package list had 16 names; 29 vendor packages that already had a `create_xxx` factory (cartesia, deepgram, fal, tavily, ...) could not be created by name and were missing from `default_providers()`. All 45 are listed now. - `transform_request_body` was rejected for every package, although several packages' settings have that field. It is passed through where the settings struct has it and rejected only where it does not. - `headers` was rejected for Google Vertex although its settings accept headers. They are passed through. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…istry Replay rebuilt the recorded model through the built-in table only, so a provider registered under the caller's own name, or a built-in name with a custom base URL, headers or fetch, could not be replayed. The rebuild function now takes a `ProviderRegistry` and resolves the recorded provider and model through it. The replay tool and the web tool pass a registry made from `default_providers()`. The check that the rebuilt model reports the recorded provider stays. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
… id generator - The Mistral stream fed the tool-call deltas of every choice into one tracker. With several choices the arguments of different choices could be joined into one call. `@ai-sdk/mistral` only reads `choices[0]`; the Rust stream now does the same. - `mistral` and `openai_compatible` each had their own copy of the id generator. Both now call `aimux_provider_utils::generate_id`, the counterpart of the AI SDK's `generateId`. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
DeepSeek is a vendor package now, not a preset row, so it rejects `params` with a different message. The test is about a preset that declares no template parameter; it uses the `abacus` row. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…ject groq options namespace
Two review findings where the models differed from the AI SDK packages:
- An assistant tool call whose input is JSON null was sent with arguments
"{}". `@ai-sdk/groq` and `@ai-sdk/deepseek` send
`JSON.stringify(input)`, which is "null". The special case is removed
in both converters.
- `provider_options = {"groq": []}` (or any non-object) was treated as
"no options". `parseProviderOptions` rejects it before the request; the
Rust model now returns `InvalidArgument`.
Not changed: Groq's image media-type detection still recognizes a short
list of signatures. Upstream uses a full signature table from
provider-utils that this workspace does not have yet.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
The compatible chat model had a `ChatHooks` trait and a usage-converter hook so that Groq and DeepSeek could bend it to their API. Both vendors have their own chat model now, the only implementor left was the default, and nothing set the usage converter. `@ai-sdk/openai-compatible` has no such layer. The trait, its default implementation, the converter hook, the parameters that only carried them and the unit test of the hook are removed; the default behavior is inlined. The extension points the AI SDK package does have (metadata extractor, request-body transform, error structure, structured-output and usage flags) stay. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…Groq uses them `@ai-sdk/provider-utils` has one media-type detector (31 signatures for image, audio, video and document types, on bytes and on base64) and `resolveFullMediaType`, which turns a declared `image/*` into the full type and fails when the data cannot be identified. The workspace had no counterpart; Groq carried a private detector that knew a few base64 prefixes and labelled everything else PNG. - aimux-provider-utils/src/media_type.rs: both functions, same tables and rules as upstream. - One table-driven test with one case per signature and the cannot-detect error. - groq/convert.rs calls the shared function; its private detector is deleted, so an unidentifiable image is an error, as upstream. OpenAI, the compatible package, xAI, Anthropic and Hugging Face still have partial detectors of their own; they move to the shared one separately. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
README, CHANGELOG (Unreleased), docs/API.md, the per-language API pages and the provider guides still described the removed by-name functions, 283 preset rows, the preset `family` and creation-time key errors. - Rust samples use `create_xxx(settings)`, `create_provider`, `default_providers` and `create_provider_registry`; model listing goes through `Provider::discovery()`. - The preset table has 281 rows; Groq and DeepSeek are vendor packages. - CHANGELOG: the removals from the crate root are listed, and a duplicated "Unreleased / Breaking" header is removed. - The other languages' pages only change where they state facts about the Rust side; their own APIs are unchanged on this branch. Historical audit notes under docs/internal, docs/plan and docs/quality-audit are left as written. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…mas do Eight review findings where the two models were more lenient than `@ai-sdk/groq` and `@ai-sdk/deepseek`: Groq - An option explicitly set to null is rejected with InvalidArgument before the request (it was treated as absent). - A stream chunk whose tool call has a `type` other than "function" fails chunk validation: an error part is emitted and the stream continues (it ended the stream). - With raw chunks requested, an unparsable chunk emits the raw part and then the error part (only the error part was emitted). DeepSeek - The options namespace must be an object; explicit nulls for `thinking` and `thinking.type` are rejected. - `strictJsonSchema` is validated as an optional boolean again. - `object`, `message.role` and the tool-call `type` of a response are checked against the literals of the upstream schema. Both - Generated tool-call ids come from `aimux_provider_utils::generate_id`; the private `call_<timestamp>` generators are deleted. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
`OpenAIProvider::language_model(id)` returned the Chat Completions model.
In `@ai-sdk/openai`, `provider(id)` and `provider.languageModel(id)`
return the Responses model; `.chat(id)` and `.completion(id)` are the
explicit ways to the other two. The Rust provider now does the same, so
`create_provider("openai", ..)` and the registry id `openai:<model>`
resolve to Responses.
- replay.rs: the test that expected a Responses recording to be refused by
the default registry now expects it to rebuild, and a chat recording to be
refused.
- openai_provider_test and the CLI probe's mock server answer on
`/responses`.
- Azure already matched `@ai-sdk/azure`; unchanged.
- CHANGELOG: listed as breaking, with `.chat(id)` as the way to keep Chat
Completions.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
The tool-type unification is its own pull request and comes before this one: every provider's output goes through `RawToolCall`, `ToolCall`, `ToolResult` and `StreamPart<C>`. This merge brings it in and applies the same conversion to the code that only exists on this branch (the Groq and DeepSeek chat models, the Mistral tracker call site, the registry and default-provider tests). No existing commit of this branch is rewritten. The resulting tree passes the full gate: fmt, clippy on the workspace with all targets, rustdoc, 3840 tests, the provider-boundary check, both generator checks, and `cargo check` of the Node and Python bindings. Cassettes and fixtures are unchanged. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…ction OpenAI, the OpenAI-compatible package, xAI (chat and Responses converters, reused by DeepSeek), Anthropic and Hugging Face each carried a private, partial media-type detector (a few base64 prefixes or magic bytes, with a PNG or JPEG default). Each converter now does what its AI SDK counterpart does at that spot: it calls `resolve_full_media_type` / `detect_media_type` from aimux-provider-utils where upstream calls `resolveFullMediaType` / `detectMediaType`, and uses the declared type where upstream does no detection. Hugging Face's upstream package is not in the pinned set, so it only swaps the detector. The private detectors are deleted (-268 lines). No test needed changing. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
The previous commit made the Rust provider's default language model the Responses model, as `@ai-sdk/openai` does. The FFI's `aimux_openai_new*` and the Node and Python `openai(key, model, base_url)` constructors went through that default, which would have switched every binding to the Responses API. Those constructors are documented and used as the Chat Completions client (the binding suites replay chat cassettes of OpenAI and of OpenAI-compatible servers through them), and this pull request does not change the bindings' API. They now call `.chat(id)` explicitly. CHANGELOG says so. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS # Conflicts: # CHANGELOG.md
…ght back Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…r options and metadata) Brings in the second and third protocol-type pull requests: one type per provider-output content concept (source, generated file, reasoning output, response metadata), and provider options / metadata as namespace -> JSON object. Conflicted vendor files keep this branch's logic and receive the same conversion again. `ProviderOptionsMap`, a crate-private trait that only existed to read the four old representations of provider options, is deleted: there is one type now. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…t have an upstream package 54 files, 48,564 lines, 1,394 tests. These tests start a mock server and assert what the Rust model sends and returns today. They were written by hand, not ported from the AI SDK's test files, so they pin the Rust implementation's current behavior. A package-by-package comparison with the pinned upstream sources (OpenAI, Azure, Anthropic, Bedrock, Google, Vertex, Mistral, Cohere, xAI) found about a hundred places where a normal call sends a different request or loses data, in code these tests cover; the tests passed throughout. They also had to be rewritten by hand for every protocol-type change. Rule applied: a test file is removed when its vendor has an upstream package in the pinned set to port tests from, and the file neither replays recorded cassettes nor replays upstream fixtures. Kept: - the 12 files that replay the 2,799 recorded cassettes; - the four `*_factory_test` files that replay upstream fixtures; - `groq_test` and `deepseek_chat_test`, already rewritten from the upstream test files; - tests of vendors with no upstream package to port from; - cross-vendor tests and everything under aimux-core/tests. The replacement is the method already used for Groq and DeepSeek: port each vendor's upstream test file together with the fixes the comparison lists. After this commit: 78 test files, 2,493 tests passing in the gate (3,840 before). Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…ry branch Brings in the role-modelled provider prompt. Files that both sides changed keep this branch's logic and are converted to the new types: the OpenAI chat and Responses converters, the OpenAI-compatible package, Groq, DeepSeek, Anthropic, Bedrock, Google, xAI and Hugging Face. The shared `resolve_full_media_type` takes the single `FilePart`. No behaviour change is intended; vendor alignment with upstream is a separate branch on top. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…ry branch Brings in the upstream-shaped request and response information on provider results. Files that both sides changed keep this branch's logic and are converted to the new fields: Anthropic streaming, Google, Open Responses, the OpenAI-compatible chat model, Groq and DeepSeek. The two files this branch had already deleted stay deleted. No behaviour change is intended. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
What: removes `recording::init_recording_from_env`, which nothing called, and makes 26 helpers private or crate-private: the retry preparation helpers and default constants, the composite text helpers, the recording call id and timestamp helpers, two session helpers, one trace helper, and in provider-utils the fetch error conversion, header pair extraction, SigV4 signing internals, logging internals, the JSON size limit constant and the default WebSocket connector. Why: none of them has a counterpart in the upstream package exports, and no other crate in the workspace used them. The public surface is meant to be what the AI SDK exports plus the registered extensions. Four more candidates stay public because public signatures mention them. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
Lists the removed function and the helpers that are no longer public, so a caller that used one of them finds it when upgrading. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
This was referenced Oct 5, 2026
cunninghamcard-bit
changed the base branch from
rfc-0036/integration
to
rfc-0036/provider-result-types
October 5, 2026 11:48
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The provider layer now has the AI SDK's shape:
create_xxx(settings)factories compile settings into a private per-model config (url(), per-requestheaders(),fetch), theProvidertrait isProviderV4-shaped, retry lives only in the call layer, recording stores identity only, and every vendor package reads providerOptions from one canonical namespace. 2,799 cassettes are untouched; every protocol change is asserted against fixtures recorded from the pinned AI SDK packages under a mocked fetch. Tests are limited to those fixture replays, the existing cassette suites and a handful of end-to-end cases; the hand-written factory tests of the first version are removed.0c0e93escripts/aisdk-fixtures/records 28 upstream cases (openai / openai-compatible / anthropic / google) under a mocked fetch; versions pinned infixtures/aisdk/VERSIONS.json316127eFetch/WsConnectorinjection,Resolvable<T>,combine_headers,load_api_key(verbatim empty string,LoadApiKey/LoadSettingerrors),prepare_retries(options.max_retries, abort)with constant defaults;RetryConfig/retry_config()deleted4b5d5efProvidertrait = ProviderV4 (required language / embedding / image →NoSuchModel { model_type }, optional rest,ProviderDiscoveryforlist_models); noname()/specification_version()/config_snapshot();ProviderRecord { provider_id, provider, model_id },RECORDING_SCHEMA = 3;CallOptions.body_overridesdeleted;scripts/check_provider_boundaries.sh9715d9ecreate_openai/openai(), privateOpenAIModelConfig,provider() = "{name}.{method}",transform_request_body; 11 fixtures replayedd4e4561create_openai_compatible, Groq and DeepSeek packages, 283-row registry,auth: none, template params; 33 thin-wrapper files deleted. (This commit generatedpresets/sources; the follow-up8a87921replaces them with a runtime table.)5141b12create_anthropic, canonical + custom namespace merge,anthropic_awsonSigV4Fetch, Vertex-Anthropic reuses the Anthropic core (591 → 117 lines)8a2460abedrock/sigv4.rsdeleted, Google bodies aligned to fixtures, namespacesgoogle/googleVertex/amazonBedrockonly7c54cfdStaticBearerConfiggone292833bEndpointConfig; poll loops on package constants (never re-submit), uploads not retried; no*Configstruct left in aimux-providers2fce704ProviderConfig(params, rejected keys), C-ABI bindings drop the call-levelbody_overridesfield, CHANGELOG, docsc241ec9chat_completions(id)recorded as an explicit extension) and adds the §5 product-difference row for it; the extension and that row are withdrawn again by062ef18. Becomes a no-op once the docs PR is on master062ef18XaiModeland the Hugging Face chat constructor are deleted (nochat_completions(id)extension). The xAI Responses model sendstop_kand warns forfrequencyPenalty/presencePenalty, as upstream does59be68c3xxis returned as a non-2xx response, so no credential header and nothing a transport decorator adds reaches another origin. Each followed hop goes through the transport again, so SigV4 re-signs. AFetchnever follows redirects:RedirectPolicy/FetchRequest.redirectare removed, and an injectedFetchgets the same rule5bdecafExternalProviderEntry'sDebugprints neither the key nor header values83b340a,7ad2e9f*_factory_testfiles keep only the fixture replays;poll_stage_testandbedrock_factory_testkeep one and three end-to-end cases (−8,085 lines)8a87921presets/(12,774 generated lines) andscripts/gen_presets.pyare deleted.provider_registry.jsonis embedded and parsed once into the descriptor table; presets are created by name. The generator's row checks run when the table loads. No per-name Rust functionsd8b96a8ed9082dcreate_provider_registry(providers, options)in aimux-core, the AI SDK'screateProviderRegistry:"provider:model"ids split at the first separator, one accessor per model kind. Upstream's six registry tests are ported8b039e8,afb23a0,140744b,226ebad,5431117,6fe5a38create_provider(name, PresetSettings)is the single assembly function: 45 vendor packages go through their own factory, every other name is a row of the preset table.default_providers()is the map for the registry. CLI, web, replay, FFI, Node and Python use it; replay resolves through the caller's registry. The older by-name functions are gone from the crate root; what the bindings still call lives inaimux_providers::providerand is built oncreate_provider.Provider::discovery()replacesprovider_discovery089cea5,99b8f5a,8f42b80,b10d98a,35eb96c,bbb9c60,cf2fad5,f7c199d@ai-sdk/groq4.0.52 and@ai-sdk/deepseek3.0.56, instead of bending the compatible model; Mistral assembles streamed tool calls with the shared tracker and reads onlychoices[0]. Where the first port kept non-upstream behavior to satisfy older tests, the behavior now follows upstream and the tests are ported from upstream's test files884ff46,b99b351groqanddeepseekrows, thefamilyfield, the per-vendor "dialect" and theChatHookslayer of the compatible model are deleted (281 rows)a5d3984,5be2a6edetectMediaType/resolveFullMediaTypeported to aimux-provider-utils (31 signatures, table-driven test); seven vendor converters drop their private partial detectorsf01ce57OpenAIProvider::language_model(id),create_provider("openai")andopenai:<model>return the Responses model, as@ai-sdk/openai;.chat(id)is the explicit way to Chat Completions27fbb43,edb88bf,0ded9accreate_provider; README, CHANGELOG and the API pages describe the current surface51af856Added since the first push
Base. This pull request now sits on five protocol-type branches, merged in without rewriting any commit here:
rfc-0036/tool-types,rfc-0036/output-content-types,rfc-0036/provider-options-types,rfc-0036/provider-prompt-types,rfc-0036/provider-result-types. Each has its own pull request; review the types there. The merge commits here only convert this branch's vendor code to those types and intend no behaviour change.Not here. Aligning each vendor package's behaviour with its upstream counterpart is the next pull request,
rfc-0036/vendor-alignment, on top of this one.Tests. 54 hand-written mock-server test files (1,394 tests) are deleted. They asserted Rust's own earlier behaviour against a hand-written server. What stays: replays of recorded cassettes, replays of upstream samples, contract fixtures, and the suites of packages that have no upstream counterpart.
Public surface. Every public item of
aimux-core,aimux-provider-utilsandaimux-providersthat has no upstream export was given a verdict (handover notes,SURFACE-DECISIONS.md): one item deleted, 26 made private, the rest kept as registered extensions, named forms of upstream's anonymous types, or return types of public factory methods.Decisions (from the plan; not re-opened here)
aimux-provider/aimux/aimux-devtoolscome later); every group keepscargo build --workspacegreen.specification_version, noProvider::name():settings.nameis the only source of the provider string ("{name}.{method}"; Anthropic defaults toanthropic.messages, a custom name is used verbatim).prepare_retries(options.max_retries, abort), defaults 2 / 2000 ms / ×2; providers never readoptions.max_retries/timeout(grep-gated).body_overridesis gone everywhere; provider-level needs usetransform_request_body.max_retries/body_overridesinProviderOptions,config_jsonor a bindingProviderConfigareInvalidArgument, not ignored.vertex/bedrock/openai-compatiblekeys and dual writes are not ported (§4.1, §5).Nonereads the env var,Some("")is sent verbatim, a missing key fails the call (LoadApiKey), never the factory; default instances read nothing and cannot fail.create_provider), one data file and one registry function; the bindings' own by-name API is unchanged in this PR and moves to the Rust surface later.RecordingFetch(S1-4), the full S2-2 split, S3, S5 beyondrebuild_provider, S6–S9, Part II,provider.tools, Bedrock-Anthropic InvokeModel, host callbacks.Known differences from the AI SDK left in place
Errfromdo_stream(upstream emits an error part), so the call layer's retry boundary sees it. All Rust packages do this.request_body/RequestBodyResultof the compatible chat model are public (upstream'sgetArgsis private); tests use them.Relation to the maintainer's RFC drafts
RFC-0040 (#210) §2.2 / §3.1 requirements are met as written: per-request key evaluation, explicit empty string,
auth=noneresolves nothing and injects no placeholder, template parameters only from the descriptor with host-part validation, unknown preset names areNoSuchProvider(no OpenAI fallback), SigV4 over the final bytes. G2 (public L0) and G3 (credential coordinator) are not implemented. RFC-0042 T1–T5 are out of scope here (message protocol).RFC-0036 §12.1 rule 2 (#207) makes "#200 merged + topic RFC accepted" the gate for cross-architecture implementation. This PR is therefore a draft on an integration base: it is offered as evidence for reviewing RFC-0040 against running code (the §12.3 rows B1–B8, #174, L0 passthrough deferral, and §13.2 retry / poll / upload rules are implemented and tested here), not as a request to merge ahead of that gate. On redirects (RFC-0040 §3.2): the first version followed every redirect with reqwest's default policy, which strips only
Authorization-style headers and would have forwardedx-api-key/x-goog-api-key/api-key/x-amz-security-token.59be68cfixes that in the request helper: same-origin redirects are followed (and re-signed for SigV4), a cross-origin3xxis not followed. Same-origin following differs from the RFC's "default no-follow" wording and is left for the RFC discussion.Verification (final head)
cargo fmt --all -- --check,cargo clippy --workspace --all-targets -- -D warnings,RUSTDOCFLAGS=-D warnings cargo doc --workspace --no-depscargo test -p aimux-core -p aimux-provider-utils -p aimux-providers -p aimux-ffi -p aimux-web -p aimux-replay -p aimux-cli --no-fail-fast: 2,501 passed, 0 failed, 6 ignored across 128 test binariesbindings/nodeandbindings/python:cargo check.npm testandpytestwere last run onc241ec9and have not been re-run; the other five bindings were not checkedpython3 scripts/gen_ts_types.py --check,python3 scripts/gen_providers_doc.py --check,bash scripts/check_provider_boundaries.shopenai_factory_test,openai_compatible_factory_test,anthropic_factory_test,google_factory_test;presets_testcreates all 281 rows without envgit diff --statagainst the first push is empty foraimux-providers/tests/cassettes,fixturesandcontract-tests/fixturesMigration (summary; full list in CHANGELOG
[Unreleased])XxxConfig::new(key).with_*()→create_xxx(XxxProviderSettings { api_key: Some(key.into()), .. })orxxx();provider()strings become"{name}.{method}"; recordings are schema 3 (schema 2 is rejected); bindingProviderConfigrejectsmaxRetries/bodyOverrides; FFI addsAIMUX_E_LOAD_API_KEY(18) andAIMUX_E_LOAD_SETTING(19). xAI / Hugging Face Chat Completions: usecreate_openai_compatiblewithhttps://api.x.ai/v1orhttps://router.huggingface.co/v1. By name:create_provider(name, PresetSettings { .. }), orcreate_provider_registry(default_providers(), Default::default()).language_model("name:model"); the crate-rootprovider(),provider_from_env,provider_handle,provider_discovery,ProviderOptionsare removed. OpenAI:language_model(id)is Responses now; use.chat(id)for Chat Completions.🤖 Generated with Claude Code
https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS