Skip to content

refactor(core)!: one tool-call and one tool-result type per protocol layer - #215

Draft
cunninghamcard-bit wants to merge 9 commits into
rfc-0036/integrationfrom
rfc-0036/tool-types
Draft

cunninghamcard-bit wants to merge 9 commits into
rfc-0036/integrationfrom
rfc-0036/tool-types

Conversation

@cunninghamcard-bit

@cunninghamcard-bit cunninghamcard-bit commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

What

Unify tool calls and provider-executed tool results in aimux-core so providers, recording, replay and bindings use shared payload structs instead of repeating field lists. This branch is based on origin/rfc-0036/integration and precedes output-content-types in the public-surface alignment stack. It separates raw provider argument text from parsed call-layer input, makes provider tool choice optional, and removes the self-generated contract data rather than treating Rust's former serialization as the upstream contract.

Before and after

Type or behavior Before After and upstream reference
Provider tool calls Generation and stream variants repeated fields; stream input mixed raw text and parsed JSON. GenerateContent::ToolCall(RawToolCall) and StreamPart::ToolCall(RawToolCall) carry input: String, corresponding to LanguageModelV4ToolCall.
Call-layer tool calls Separate raw-call definition and conversion heuristics. ToolCall carries parsed Value, or the original argument string when parsing fails, with invalid and error; TypedToolCall is the upstream reference. TextStreamPart is currently an alias for StreamPart<ToolCall>.
Provider-executed results An unused standalone result struct and repeated enum payloads. GenerateContent::ToolResult and StreamPart::ToolResult share ToolResult, including required tool_name and optional dynamic and provider_metadata, corresponding to LanguageModelV4ToolResult.
Tool choice CallOptions always supplied a choice. CallOptions.tool_choice is optional; direct provider calls omit an unset choice. generate_text and stream_text supply auto when a nonempty tool list is supplied without a choice, following ai/src/prompt/prepare-tool-choice.ts.
Tool examples and arguments Examples were arbitrary JSON values; provider arguments could be nonobjects. FunctionTool.input_examples contains { input: object }; ProviderTool.args is a JSON object, following LanguageModelV4FunctionTool and LanguageModelV4ProviderTool.
Thought signatures Tool calls and prompt parts exposed a separate signature field. Signatures travel through the vendor namespace in provider metadata on output and provider options on input, where the upstream Google and Vertex modules read and write them.
Legacy argument decoding Generation accepted a nonstring input and serialized it back to text; stream conversion guessed its representation. The compatibility decoder and raw_tool_input heuristic are deleted. Provider input must be raw text, as in LanguageModelV4ToolCall.

JSONSchema7, dotted provider-tool IDs and NonNullable<JSONValue> constrain TypeScript types without upstream runtime enforcement at these type declarations. Rust keeps Value for schemas and results and String for IDs; this branch adds no corresponding runtime checks. This does not imply that vendor option parsers accept arbitrary values.

Tests

Delete contract-tests/, its self-generated JSON and runner, together with the 16 tests reading that data. These files recorded Rust's own former output rather than upstream payloads. CI, the local CI script and contributor and binding documentation stop invoking or recommending those readers; the generator drift checks remain.

Delete hand-written tests that pinned removed shapes and the invented-data tests in openai_output.rs and moa.rs. Restore the tool-choice cases ported from the upstream OpenAI, Groq, Mistral and Amazon Bedrock suites, adapting them to optional provider choice while preserving their upstream request expectations. Other retained tests follow the shared payload structs. Cassette replays are kept; aimux-providers/tests/cassettes and fixtures/ have no changes in this branch.

The diff also deletes older provider test suites, including suites with upstream-port annotations. The restored tool-choice cases do not establish complete upstream suite coverage; this body makes no claim that every deleted provider case was hand-written or restored.

Not in this pull request

  • Provider ToolChoice serialization still uses bare strings for auto, none and required, while LanguageModelV4ToolChoice uses tagged objects. This refactor changes optionality, not that existing wire representation.
  • Serde naming remains pending the maintainer's decision: snake_case fields, PascalCase external tags, Option fields serialized as null where omission is not configured, and generated Node types with required-nullable fields. Those binding contract choices are not resolved here.
  • The call-layer ModelMessage tool-result shape remains pending the maintainer's decision. TextStreamPart also still carries the provider result rather than the upstream call-layer input/output result and lacks a separate tool-error event for invalid non-provider-executed calls; call-layer result and error separation is outside this payload refactor.
  • The public TextStreamPart alias is replaced by a separate enum on provider-result-types, where response metadata is absorbed into the finish step. That later change is not part of this branch.
  • Function-tool provider options still use a broad JSON namespace map here; the shared namespace type is introduced on provider-options-types, which owns option and metadata alignment.
  • A repaired call that fails validation is still wrapped as a repair error, and invalid-call fallback still uses plain JSON parsing rather than the upstream secure parser. These existing parser differences are not fixed by consolidating the payload types.
  • Provider prompt unions are handled on provider-prompt-types; the remaining output content unions are handled on output-content-types, to keep this change focused on tool payloads.

Not verified

Go, Java, Kotlin, Swift and Flutter type and decoder changes were not compiled locally. Node npm test and Python pytest were not run. CI YAML was only syntax-checked; CI execution is not established. This documentation rewrite did not run compilation or tests, and retained upstream test coverage has not been exhaustively mapped.

Verification

Gate on the final commit 8a5acf97: cargo fmt --check, cargo clippy --workspace --all-targets -D warnings, cargo doc, tests of core, provider-utils, providers, ffi, web, replay and cli: 3,321 passed, 0 failed; Node and Python bindings cargo check; provider boundary script; gen_ts_types.py --check, gen_providers_doc.py --check. Cassettes and fixtures/aisdk unchanged relative to the base. 136 files, +1107 / −25133.

🤖 Generated with Claude Code

https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS

chenhaonan and others added 9 commits October 4, 2026 16:56
aimux-core defined "tool call" five times and "tool result" four times, with
`input` meaning raw text in some and parsed JSON in others, so
`raw_tool_input` had to guess which. Mirror the AI SDK: one type per layer.

- Provider output: `tool::RawToolCall` (moved from parse_tool_call), input is
  always the raw argument string. `GenerateContent::ToolCall(RawToolCall)`.
- Core output: `tool::ToolCall`, input always parsed. Unchanged.
- `tool::ToolResult` is now the union of the old variant fields and is the
  payload of both `GenerateContent::ToolResult` and `StreamPart::ToolResult`.
- `StreamPart<C = RawToolCall>` is generic over the tool-call payload;
  `TextStreamPart = StreamPart<ToolCall>` is what `stream_text` produces and
  later stages consume. `map_tool_call` re-emits the parsed call.
- ts-rs exports the generic with `concrete(C = ToolCall)`, so the TS shape of
  the stream part is unchanged; bindings regenerated.
- Deleted `raw_tool_input`, `deserialize_tool_call_input` (and its legacy-shape
  test) and every Value::String round trip of tool input.

Wire format is unchanged except that a provider-layer `StreamPart::ToolCall`
input is now always a JSON string. aimux-providers, aimux-ffi, bindings and
tools are intentionally not updated here.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
Rebuilding the raw call from an invalid `ToolCall` whose error did not
record the argument text used `input.to_string()`, which turns the stored
text `{broken` into the JSON string literal "{broken" and makes malformed
arguments look valid on the next parse. `invalid_tool_call` stores
unparsable text as a JSON string by definition, so that string is returned
as is; a structured input is serialized. The AI SDK keeps the original
`toolCall.input` here (parse-tool-call.ts).

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…types

Follows the core change: a provider emits `RawToolCall` (raw argument
text) in `GenerateContent::ToolCall` and `StreamPart::ToolCall`, and
`tool::ToolResult` for provider-executed results. Streams after
`stream_text` are `TextStreamPart`.

- Vendor packages: variants built with the shared structs; the places that
  wrapped the argument text in `Value::String` pass the text itself.
- Tests: provider-layer inputs are compared as strings; tests reading a
  `stream_text` stream match the parsed `ToolCall`.
- FFI, Node, Python: the forwarded stream is typed `TextStreamPart`.
  Exported symbols and JSON output are unchanged.
- Generated TypeScript types regenerated.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
With one struct per layer, the remaining places that passed or rebuilt a
tool call or result field by field take the struct:

- `response_messages::tool_result` takes a `ToolResult` instead of seven
  parameters; the two callers in generate.rs no longer destructure one to
  pass the same fields on.
- parse_tool_call.rs: the three places that copied every field between
  `RawToolCall` and `ToolCall` share one private conversion each way.
- replay.rs passes the recorded `RawToolCall` through.

No behavior or public API change.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…shape note

- CHANGELOG (Unreleased, Breaking): the new types, where each is used, and
  the three wire-format changes.
- docs/api/gaps.md §9 said a pre-refactor `GenerateResult` with a parsed
  `input` keeps loading through a compatibility deserializer. That
  deserializer is removed; the section now says so and describes the
  provider stream and `TextStreamPart`.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…uments

What: CallOptions.tool_choice is optional and each vendor omits it when
unset, as its upstream package does. FunctionTool.input_examples is a list
of { input: object } and ProviderTool.args is a JSON object. Hand-written
tests that pinned the old shapes are deleted rather than rewritten.

Why: an independent field-by-field comparison with the pinned upstream tool
types found these three differences. The project keeps cassette replays,
upstream samples and contract fixtures as its tests; those still pass.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…read them

What: removes contract-tests/ (wire-format.json, tool-call-repair.json and
the Node runner), the Rust contract test and host-side repair test, and
the tests in the Go, Java, Kotlin, Swift, Flutter, Python and Node
bindings that read those files. CI, the local CI script and the docs no
longer mention them; the CI job keeps the two generator drift checks.

Why: the fixtures recorded Rust's own serialized output from before the
AI SDK alignment, not real data. Three records pinned shapes that conflict
with the AI SDK types (a separate thought-signature field, a source with
neither url nor title, the earlier raw result JSON). The owner's decision
is to keep only tests that run on real data and delete this set.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
What: removes the thought_signature field from the provider tool call and
from the user-level tool call. The value travels in provider metadata on
output and provider options on input, under the vendor namespace, where
each upstream vendor package reads and writes it. The language bindings'
types and decoders follow.

Why: upstream's tool-call types have no such field. The Go, Java, Kotlin,
Swift and Flutter changes are not compiled here; no toolchain is
available locally.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
… tool-choice tests

What: generate_text and stream_text supply `auto` when tools are given and
no choice is set, as upstream prepare-tool-choice does; the provider layer
still omits an unset choice. Restores the fifteen tool-choice tests ported
from upstream (OpenAI, Groq, Mistral, Bedrock) that the earlier commit
deleted along with the hand-written ones, adapted to the optional choice.
Deletes the invented-data tests inside openai_output.rs and moa.rs.

Why: an independent review found the call-layer default missing and
ported tests deleted; ported tests are kept, hand-written ones are not.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant