Skip to content

fix(providers)!: align ten vendor packages with their upstream counterparts - #220

Draft
cunninghamcard-bit wants to merge 27 commits into
rfc-0036/provider-factoryfrom
rfc-0036/vendor-alignment
Draft

cunninghamcard-bit wants to merge 27 commits into
rfc-0036/provider-factoryfrom
rfc-0036/vendor-alignment

Conversation

@cunninghamcard-bit

@cunninghamcard-bit cunninghamcard-bit commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

What

Align vendor request conversion, response conversion, authentication and streaming with the pinned AI SDK sources. This pull request (#220), on rfc-0036/vendor-alignment, follows rfc-0036/provider-factory (#214) in the stack and uses the protocol types and factory API introduced below it. It ports previously missing vendor features and removes conflicting behavior; the remaining model families and public-surface decisions are listed below. The final branch diff against rfc-0036/provider-factory is 155 files changed, 19,410 insertions and 6,336 deletions.

Before and after

Area and upstream module Before After
OpenAI OpenAIEmbeddingModel Embeddings accepted base64 vector strings. Requests encoding_format: "float" and reads numeric vectors.
OpenAI OpenAIResponsesLanguageModel and convertToOpenAIResponsesInput Provider tool requests, events and history were incomplete. Adds typed provider-tool factories, request and event conversion, approval requests and compaction custom content. Approval IDs map back to the original tool-call IDs in generate and stream; custom-tool history retains async. The local shell environment accepts an omitted type tag, as in upstream shell.ts.
Azure provider and transcription modules Transcription used only the OpenAI route. Selects OpenAI or Azure Speech transcription, including the Speech endpoint, authentication headers and base-URL override.
Anthropic AnthropicLanguageModel and message conversion Document citations and container uploads were dropped; structured-output fallback and computer toolset mapping were incomplete. Returns document sources and container-upload custom content, supports the JSON-tool fallback, and maps computer toolset members to the caller's tool name with action and toolsetName. History retains the toolset mapping. Mid-stream errors are forwarded without a synthetic finish.
AmazonBedrockChatLanguageModel and message conversion Structured JSON handling and provider-defined Anthropic tools were missing; inline media could be rejected before detection. Implements native response format, JSON-tool and prompt-instruction paths, with generate/stream extraction and JSON-tool finish-reason and metadata handling. Adds provider-defined tools and model-family selection. Resolves inline media before format validation, percent-encodes model IDs, and emits metadata under amazonBedrock and bedrock. Response bodies are omitted as upstream does.
Google GoogleLanguageModel and message conversion Reasoning-file replay was rejected; null function-call arguments could become "null", and null partialArgs could discard a call. Preserves generated files, reasoning files and their replay metadata. Null args becomes {} without a parameter delta; null partialArgs permits the complete-call path. Provider response model IDs are left unset.
GoogleVertexProvider and GoogleVertexAnthropicProvider Authentication required an explicit or environment token, and Anthropic models also had an entry on the main Vertex provider. Adds application default credential resolution and googleAuthOptions, plus a separate Vertex Anthropic factory. Removes the main provider's anthropic_model entry. Vertex Anthropic uses googleVertex metadata and excludes Express mode.
Vertex language, image and transcription modules Partial function arguments and thought files were incomplete; Speech-to-Text requests lost authentication headers. Accumulates partial function arguments, preserves thought files, rejects provider references in user files, and carries credentials to Speech-to-Text. Adds Gemini transcription alongside Speech-to-Text. Image generation follows the upstream Gemini path; URL image input remains rejected by media-type resolution, matching the pinned upstream path.
Mistral MistralChatLanguageModel and message conversion Prompt and tool-result conversion used separate local helpers. Consolidates conversion, preserving the tool result's name and using the shared tool-output conversion.
Cohere CohereChatLanguageModel A stream ending without usage could panic. Returns a parse error for missing message-end usage.
xAI XaiResponsesLanguageModel Image-generation tool calls and results were dropped. Returns their calls and results in generate and stream.
Groq GroqChatLanguageModel Rejects a first SSE error before returning a stream. The final implementation still rejects it. The retained upstream-ported test now expects stream-start, error and finish, leaving an implementation/test mismatch to fix.

Where upstream expresses constraints only in TypeScript types, such as dotted IDs and kinds, schema types, non-null types, URL/Date values and numeric domains, it does not enforce those constraints at runtime either. Rust retains Value, String and u32 where those representations apply.

Tests

  • Keeps existing upstream-ported tests and adjusts them to the upstream behavior and current factory API. Restores the upstream expectation for the Groq first-error SSE sequence and the Bedrock end-to-end toolChoice: "auto" input and assertion. The OpenAI streaming selection asserts the upstream event sequence, and the structured-output selection uses the upstream model's capability path.
  • Adds selected upstream cases in the vendor upstream_e2e_test.rs files through the public factories with an injected transport. These cover selected generate, stream, tool, error and newly ported feature paths; they are not complete upstream suites.
  • Deletes hand-written tests or assertions that preserved removed Rust behavior, including Anthropic assertions requiring a synthetic finish after a mid-stream error. Remaining upstream ports and cassette replays are retained or updated for the new behavior.
  • Retains cassette regression coverage, including binary Bedrock body_base64 stream replay, and restores Google generated-file regression coverage from recordings. Deletes the OpenAI embedding replays whose recordings contain base64 vectors and the Bedrock conformance streaming replay whose text-encoded binary data cannot exercise the upstream decoder. Recording files are unchanged; rerecording is a maintainer decision.
  • Neither aimux-providers/tests/cassettes/ nor fixtures/ changes in this branch diff.

Later commits on this branch, after the review fixes: Groq and the openai-compatible chat model no longer read the first SSE event before returning the stream; an error event flows as stream-start, error, finish as upstream does. Bedrock gains the S3 supported-url entry, sends an empty tool-use input as {}, and its migrated image and embedding tests mount the percent-encoded model path. The two upstream-translated files bedrock_remaining_test.rs (38 cases) and anthropic_remaining_test.rs (21 cases), deleted lower in the stack, are back, adapted and no longer ignored; Anthropic gains the mid-conversation tool change handling, argument validation, leading system message handling, beta headers and the empty base-url error text those cases cover. Six Bedrock cases are left out: three whose upstream expectation changed, two environment cases without an upstream counterpart, one duplicated by the e2e selection.

Not in this pull request

  • Google speech, transcription and evaluation; xAI image, files, speech, transcription and video; Groq transcription; Mistral speech and transcription; OpenAI skills; and registry evaluationModel and customProvider are not ported. These require additional model families or interfaces, pending the maintainer's decision. xAI Responses image-generation tools do not provide a standalone image-model family.

  • Serde naming remains a maintainer decision: snake_case fields, PascalCase external tags, Option serialized as null, and generated Node fields that are required and nullable. This pull request does not settle the binding wire format.

  • The by-name entries in aimux_providers::provider, used by FFI, Node and Python, remain pending the maintainer's decision. Package factory changes do not settle those entries.

  • The call-layer ModelMessage tool-result shape and remaining tool-result/error behavior remain pending the maintainer's decision. The distinct call-layer stream enum introduced below this branch does not settle those remaining behavior choices.

  • The Google authentication port covers credential resolution and token exchange used by the Vertex factories. It does not reproduce every upstream auth-library transporter option, retry policy, environment probe or project-discovery path; those library behaviors require further work.

  • Groq first-error SSE handling remains unresolved in the final branch: do_stream still rejects before constructing the stream, although its upstream-ported test now expects stream-start, error and finish. The implementation must be fixed to satisfy that restored expectation.

  • The selected vendor tests do not port complete upstream suites. Groq options include runtime schema tests as well as a TypeScript-only inference case; options, usage, message conversion, tool preparation and generate/stream cases remain unported. Completing those ports requires retaining the upstream inputs and assertions.

  • The OpenAI chat and Responses models still read the first SSE event before returning the stream (aimux-providers/src/openai/model.rs, openai/responses/mod.rs); the same change as for Groq is owed there. No ported test covers it yet.

Not verified

This documentation review did not run compilation, tests or runtime gates. Node npm test and Python pytest were not run. Other language bindings edited earlier in the stack were not compiled locally. CI YAML was only syntax-checked, which does not establish CI execution. Live credential exchange, vendor requests and the unported upstream cases are not verified here.

Verification

Gate on the final commit 5dc85be6: cargo fmt --check, cargo clippy --workspace --all-targets -D warnings, cargo doc, tests of core, provider-utils, providers, ffi, web, replay and cli: 2,514 passed, 0 failed; Node and Python bindings cargo check; provider boundary script; gen_ts_types.py --check, gen_providers_doc.py --check. Cassettes and fixtures/aisdk unchanged relative to the base. 158 files, +21206 / −6464.

🤖 Generated with Claude Code

https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS

chenhaonan and others added 27 commits October 5, 2026 19:09
…parts

What: fixes the differences that a package-by-package comparison against
the pinned upstream sources found in OpenAI, Azure, Anthropic, Amazon
Bedrock, Google, Google Vertex, Mistral, Cohere, xAI and the
OpenAI-compatible package. The changes are in request building (option
names and validation, schema handling, cache and beta merging, tool
conversion), response mapping (usage, provider metadata keys, citations and
sources, reasoning) and stream handling (error parts, raw chunks, finish
data). Examples:

- Bedrock percent-encodes the model id in the URL, writes metadata under
  both `amazonBedrock` and the legacy `bedrock` key, and applies upstream's
  capability rules to sampling settings and reasoning.
- OpenAI embeddings request `encoding_format: "float"` and no longer decode
  base64 vectors; the Responses log-probability include follows upstream's
  condition.
- Google no longer invents `response.model_id`; the call layer supplies the
  default, as upstream does.
- Anthropic emits citation sources, forwards a mid-stream `error` event as
  an error part without a synthetic finish, and falls back to a JSON tool
  for structured output on models without native support.
- Vertex accumulates partial function-call arguments; Vertex Anthropic uses
  the `googleVertex` namespace and never Express mode; Vertex transcription
  sends its credentials to the Speech-to-Text host (they were dropped
  before, so every such request went out unauthenticated).
- xAI Responses returns image generation calls and results.
- Each package sends its own `user-agent` suffix. Mistral has no
  `transform_request_body` setting. The deprecated `text_embedding` and
  `text_embedding_model` aliases are not carried over.

Why: the direction is that Rust behaves as the AI SDK does, with upstream
as the reference. Of 106 severe findings, 94 are fixed here; the 12 left
need new model kinds, shared type changes or Google default credentials,
and are listed with reasons in the pull request description.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
What:
- The cassette replay harness and the Bedrock regression mocks compare
  request paths after percent-decoding, so an encoded model id matches a
  recording made either way.
- The upstream-sample tests ignore `user-agent` on both sides: the recorded
  value carries the upstream package version, ours carries this crate's.
- The Google upstream-sample test no longer expects the provider result to
  carry a model id.
- Bedrock regression tests accept the legacy `bedrock` metadata key next to
  `amazonBedrock` and read the reasoning signature from the delta that
  carries it. These tests replay values taken from recordings, so their
  assertions are adjusted rather than the tests removed.
- Removed tests that pinned behaviour Rust no longer has: top-level
  `cached_tokens` usage for two preset rows, a synthetic error finish after
  a mid-stream Anthropic error, a Responses stream that ends itself after
  an error event, two OpenAI embedding replays whose recordings hold base64
  vectors, and a Bedrock stream replay over recordings that stored the
  binary event stream as text.

Why: these tests asserted the old behaviour. No cassette or fixture is
changed; the recordings that can no longer be replayed are listed in the
pull request description.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…am tests

What: ten files, `<vendor>_upstream_e2e_test.rs`, for OpenAI, Azure,
Anthropic, Bedrock, Google, Vertex, Mistral, Cohere, xAI and the
OpenAI-compatible package: 67 tests. Each builds the provider through its
public factory with an in-process transport, calls the model, and asserts
the request that was sent and the parts that came back. Per model kind
there is at most one test each for generate, stream, tool call and error.
Every test names the upstream test it comes from.

Why: Azure and Vertex had no replay test at all, and the alignment needed
something that exercises it. This is a selection, not a case-by-case port.
The Vertex transcription test is the one that showed requests went out
without credentials. One test is ignored with its reason:
`CallOptions::tool_choice` is not optional, so `tool_choice: "auto"` is
always sent where upstream omits it.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
Request URLs, metadata keys, embedding encoding and removed provider
methods change for callers, so they are listed under Breaking.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
What: a Cohere message-end event without a usage object is a parse error
instead of a panic. Groq and DeepSeek return the raw response body again.
Mistral takes the tool message name from the tool result only. Bedrock has
one request builder. OpenAI and Azure lose the unused request-body
transform slot. Two end-to-end tests assert what the upstream tests
assert: the full OpenAI stream sequence, and native structured output
selected by the model id the upstream test uses.

Why: an independent review found a crash, two dropped fields, one finding
marked fixed that was not, a second code path, and two weakened tests.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
Why: bring the tool-result alias removal up the stack. The provider
sources keep this branch's versions; they already read the base URL
environment variables, skip empty keys and report upstream's provider
names, as the vendor alignment did those first.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
What: Bedrock structured JSON output for generate and stream, Bedrock
provider-defined Anthropic tools, a closed set of Bedrock model families.
Azure speech transcription with its own route and endpoint. Vertex and
Vertex Anthropic application default credentials with the upstream auth
options. A Gemini transcription model on Vertex. OpenAI Responses provider
tools with their typed factories, request mapping and response and stream
events. The Responses conversion returns an error where it could panic.
One test per new feature is selected from the upstream test files.

Why: these were listed as open differences from upstream; the owner's
decision is to close them in this pull request rather than carry them.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
What: brings up the tool, content and provider-option type changes and
converts this branch's code to them. OpenAI Responses emits tool approval
requests and compaction as custom content. Anthropic emits custom content
where upstream does. Google and Vertex return thought images and files as
reasoning files. A vendor omits tool_choice when the caller set none, and
the test that was ignored for this runs again.

Why: these open differences were waiting for the content variants that
the lower branches now provide.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
Why: carry the provider prompt changes up the stack: structured tool
result output with a required tool name, signatures in provider options,
and the reasoning-file, custom and tool-approval-response parts. Each
vendor converter consumes them as its upstream package does. Hand-written
tests that no longer compile are deleted.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
Why: carry the factory branch's additions up the stack: completion(id) for
OpenAI and the OpenAI-compatible package, Anthropic skills, registry
middleware, and the private request-body builder. This branch's
upstream-aligned vendor behaviour is kept. Vendor namespaces in the newly
ported code go through the package constants so the boundary check passes.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…back

What: the Bedrock stream test and the OpenAI factory test no longer expect
a default tool choice, which upstream does not send; a Google request-body
difference is fixed. In-file test modules and test functions that earlier
merge resolutions had restored are removed. Test files ported from
upstream test suites are kept.

Why: tool choice is optional now and the vendors omit it when unset, as
upstream does. Hand-written tests that pin Rust-only behaviour are
deleted in this project, and the branches below had already deleted these.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
Why: keep the stack linear after the factory branch dropped the
hand-written tests that earlier merges had restored. This branch's
upstream-aligned vendor code is kept in the conflicts.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…der docs

What: the data-loss regression test keeps this branch's assertions, which
follow the upstream-aligned Bedrock and Anthropic behaviour, instead of
the factory branch's. The Google end-to-end test no longer expects a
default tool config. docs/api/providers.md is regenerated.

Why: the last merge took the lower branch's copy of the regression test,
which asserts the pre-alignment behaviour, and the generated provider
list was stale after the new models were added.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…ours

What: brings back 30 test functions ported from upstream test suites
(DeepSeek, Groq, Open Responses). Anthropic document citations are
emitted as sources as far as the source type can carry them. Bedrock no
longer returns a response body upstream does not return. A DeepSeek
stream error event is reported as upstream reports it. The hand-written
provider trait test is deleted.

Why: an independent audit found that upstream-ported coverage had been
removed together with hand-written tests, and listed these differences.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…ntract fixtures

Why: carry up the stack the source union, the removal of the separate
thought-signature field and the deletion of the self-generated contract
fixtures. This branch's vendor code builds url and document sources and
carries signatures in provider metadata and options as upstream does.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…nges from below

Why: carry up the factory branch's settings alignment, image-model
middleware and restored upstream-ported tests, together with the raw
adapter removal from below. The settings take the factory branch's shape;
this branch's upstream-aligned vendor behaviour is kept; the test files
keep the union of upstream-ported tests from both sides.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
What: drops one assertion from the Bedrock generate test.

Why: upstream's Bedrock model returns id, timestamp, model id and headers
as response information and no body; the source already follows that.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
Why: the factory branch merged its split-out parts without changing any
file; this keeps the stack linear.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
What: Google and Vertex treat `partialArgs: null` and `args: null` as
upstream does, Vertex rejects user file references and loses its old
anthropic_model entry in favour of the Vertex Anthropic factory;
Anthropic maps computer toolset calls back to the caller tool with
action; OpenAI Responses maps approval ids to the original call id, keeps
the async option on custom tool history and accepts a shell environment
without a type tag; Bedrock detects inline media before validating it;
Groq turns a first error event into stream-start, error, finish. Removes
the GenerateResponse and StreamResponse aliases. The Bedrock e2e test
regains the upstream toolChoice input and assertion.

Why: review of this branch against the pinned vendor packages.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
Why: carry up the stack the fixes made after the third independent
review on the branches below; this branch's code is converted to them.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
Why: carry up the stack the fixes made after the third independent
review on the branches below; this branch's code is converted to them.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
… the encoded model id

What: Groq and the openai-compatible chat model no longer pre-read the
first SSE event; an error event flows as stream-start, error, finish as
upstream does. Bedrock gains the S3 supported-url entry and sends an
empty tool-use input as {}; its temperature warning matches upstream.
The migrated Bedrock image and embedding tests mount the percent-encoded
model path the model now requests.

Why: the gate showed sixteen failures after the merges; in each case
the pinned upstream behaviour decides.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…d remaining cases

What: brings back bedrock_remaining_test.rs (38 cases) and
anthropic_remaining_test.rs (21 cases), translated from the upstream
packages, adapted to the current API and no longer ignored; Anthropic
gains the mid-conversation tool change handling, argument validation,
leading system message handling and beta headers those cases cover, and
the empty base-url error text. Six Bedrock cases are left out: three
whose upstream expectation changed (merged into the hash case), two
environment cases with no upstream counterpart, one duplicated by the
e2e selection.

Why: the files were deleted lower in the stack as hand-written; they are
upstream translations and the project keeps those.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
Why: carry up the stack the fixes made after the third independent
review on the branches below; this branch's code is converted to them.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
What: the Azure speech path spells the package name through
options::NAMESPACE instead of a literal.

Why: the provider boundary check forbids the literal outside options.rs.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
Why: carry the formatting commit up the stack; no code change.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
Why: carry the ancestry up the stack; no code change.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant