Unified LLM service access layer — one API to access 325 AI providers
Shared reference — parameter tables, result shapes, factory functions, and the feature coverage matrix — lives in the API overview.
pip install arcships-aimuxfrom aimux import openai, generate_text
model = openai("sk-...", "gpt-4o")
result = generate_text(model, "What is Rust?")
print(result["text"])from aimux import provider, generate_text
# Key from the provider's env var (GROQ_API_KEY etc.):
model = provider("groq", None, "llama-3.3-70b")
# Explicit key + base URL override:
model = provider("groq", "sk-...", "llama-3.3-70b", "https://relay.example/v1")
# Full ProviderOptions via config dict:
model = provider("groq", "sk-...", "llama-3.3-70b",
config={"headers": {"X-Custom": "1"}, "max_retries": 0})
result = generate_text(model, "Hello")provider(name, api_key, model_id, base_url=None, config=None) covers all
251 built-in OpenAI-compatible providers. config takes the full
ProviderOptions shape (base_url / headers / organization / project /
max_retries / body_overrides); the base_url parameter wins over
config["base_url"]. openai / anthropic / deepseek factories
remain (deepseek is now registry-backed).
Scope:
provider(name)covers only the 251 registry OpenAI-compatible providers; Anthropic/Google/multimodal/local → typed factories (anthropic(api_key, model)); custom endpoints →base_urlparam. Full list: providers.md.
Non-streaming text generation; returns the complete result.
from aimux import openai, generate_text
model = openai("sk-...", "gpt-4o")
result = generate_text(model, "Explain Rust ownership.", {
"max_output_tokens": 100,
"temperature": 0.7,
})
print(result["text"])
print(result["usage"])
print(result["finish_reason"])Parameters, return value, and the
raw.contentvariants are documented in the API overview.
Returns generated content as a stream, output chunk by chunk.
from aimux import openai, stream_text
model = openai("sk-...", "gpt-4o")
for part in stream_text(model, "Write a haiku about Rust."):
if "TextDelta" in part:
print(part["TextDelta"]["delta"], end="")
if "Finish" in part:
print("\n[done]")Stream part variants are documented in the API overview.
Tool definitions are language-agnostic data descriptions (JSON Schema) that require no macros.
# Python — the same data shape via the options dict
tools = [{
"type": "function",
"name": "get_weather",
"description": "Get current weather",
"input_schema": {
"type": "object",
"properties": {"location": {"type": "string", "description": "City name"}},
"required": ["location"]
}
}]
result = generate_text(model, "What's the weather in Tokyo?", {"tools": tools})
if len(result["tool_calls"]) > 0:
call = result["tool_calls"][0]
print(call["tool_name"]) # get_weather
print(call["input"]) # {"location": "Tokyo"}A tool call the model got wrong never fails generation: it comes back with
invalid: true and a typed error. Pass repair_tool_call to get one shot at
fixing it (RFC-0035, the equivalent of the AI SDK repairToolCall):
def repair(ctx):
# ctx = {tool_call, error, input_schema, tools, messages, instructions};
# ctx["tool_call"]["input"] is the provider's raw argument text.
args = json.loads(ctx["tool_call"]["input"])
if "location" not in args:
return None # None = leave the call invalid
return dict(ctx["tool_call"], input=json.dumps({"city": args["location"]}))
result = generate_text(model, "What's the weather in Tokyo?",
{"tools": tools, "repair_tool_call": repair})
print(result["tool_calls"][0]["input"]) # {"city": "Tokyo"}The function is called once per invalid call, after generation, and its result
is re-validated against the tool's schema: a call that is still wrong — and one
whose repair raises — comes back invalid with a ToolCallRepair error carrying
the original one. Returning None leaves the call untouched. It runs on the
Python side, outside any native call, so it may itself call generate_text to
ask a model for better arguments. Both tool_calls and response_messages are
patched, so the next turn replays the repaired arguments.
The same option exists on the typed wrapper's GenerateTextOptions, where it
takes a ToolCallRepairContext and returns a RawToolCall. It applies to
generate_text, generate_object, consume_stream_text and stream_text
(which yields the repaired ToolCall part; tool-input deltas are the
provider's own text and pass through untouched), and to
generate_text_as_openai: a chat.completion carries no invalid marker, so
it repairs the native result and converts it
(Model.generate_text_result_as_openai). stream_text_as_openai does not
reflect repair — its argument deltas are the provider's text, as in the AI SDK.
Only an exception raised by the repair function itself means "failed". An
error in the binding's own decoding or encoding — including a typed hook that
returns something other than a RawToolCall or None — propagates as itself.
Pass tool_choice through the options dict:
opts = {
"tools": tools,
"tool_choice": "auto" # "auto" | "none" | "required" | {"type": "tool", "toolName": "get_weather"}
}prompt accepts a message array to implement multi-turn conversation; roles support system / user / assistant / tool:
# Python — system + user multi-turn
result = generate_text(model, [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is Rust?"},
])Converts text into a vector representation.
from aimux import openai_embedding
import json
embedder = openai_embedding("sk-...", "text-embedding-3-small")
# embed() takes a JSON string, returns a JSON string
result = json.loads(embedder.embed(json.dumps(["hello", "world"])))
print(len(result["embeddings"])) # 2
print(len(result["embeddings"][0])) # 1536Converts text into speech audio.
from aimux import openai_speech
import json, base64
speaker = openai_speech("sk-...", "tts-1")
result = json.loads(speaker.generate(json.dumps({
"text": "Hello world!",
"voice": "alloy",
"output_format": "mp3",
})))
if "Base64" in result["audio"]:
audio_bytes = base64.b64decode(result["audio"]["Base64"])
with open("out.mp3", "wb") as f:
f.write(audio_bytes)Converts audio into text (non-streaming).
from aimux import openai_transcription
import base64, json
transcriber = openai_transcription("sk-...", "whisper-1")
audio_b64 = base64.b64encode(open("audio.mp3", "rb").read()).decode()
result = json.loads(transcriber.generate(audio_b64, "audio/mp3"))
print(result["text"])
print(result["segments"])from aimux import openai_image
import json, base64
imager = openai_image("sk-...", "dall-e-3")
result = json.loads(imager.generate(json.dumps({
"prompt": "A cute baby sea otter",
"n": 1,
"provider_options": {},
})))
if "Base64" in result["images"]:
with open("out.png", "wb") as f:
f.write(base64.b64decode(result["images"]["Base64"][0]))Video generation typically returns a URL (not binary).
from aimux import google_video
import json
videor = google_video("sk-...", "veo-3.0")
result = json.loads(videor.generate(json.dumps({
"prompt": "A cat playing piano",
"n": 1,
"provider_options": {},
})))
# result["videos"] is usually [{"Url": {"url": "...", "media_type": "..."}}]
if "Url" in result["videos"][0]:
print(result["videos"][0]["Url"]["url"])Reorders a document list by relevance.
from aimux import cohere_reranking
import json
reranker = cohere_reranking("sk-...", "rerank-v3.0")
# docs_json is the externally-tagged `RerankingDocuments` enum —
# {"Object": {"values": [...]}} for JSON documents, {"Text": {"values": [...]}} for strings
docs = {"Object": {"values": [
{"text": "Rust is a systems programming language."},
{"text": "Rust is a chemical element."},
]}}
result = json.loads(reranker.rerank(
"What is Rust?",
json.dumps(docs),
# opts_json is a whole RerankingCallOptions: only top_n and provider_options
# are read from it, but query and documents must be present to deserialize
json.dumps({"query": "What is Rust?", "documents": docs, "top_n": 3}),
))
# result["ranking"] sorted by relevance_score
for rank in result["ranking"]:
print(rank["index"], rank["relevance_score"])from aimux import tavily_search
import json
searcher = tavily_search("tvly-...")
result = json.loads(searcher.search("What is Rust?"))
print(result["results"][0]["title"]) # ordered result list
print(result["answer"]) # provider's summary, if anyUploads a file to the provider and returns a file ID.
from aimux import openai_files
import base64, json
files = openai_files("sk-...")
file_b64 = base64.b64encode(open("doc.pdf", "rb").read()).decode()
result = json.loads(files.upload_file(file_b64, "application/pdf"))
print(result["provider_reference"]) # {"openai": "file-xxx"}The aimux package has two layers:
| Layer | Source | Boundary |
|---|---|---|
| Native (PyO3) | bindings/python/src/lib.rs (aimux.abi3.so) |
JSON strings in / JSON strings out |
| Typed wrapper | bindings/python/python/aimux/wrapper.py |
pydantic models / Python dicts |
| Class | Factory functions | Methods |
|---|---|---|
Model |
openai / anthropic / deepseek |
generate_text(prompt_json, opts_json=None), stream_text(...) |
EmbeddingModel |
openai_embedding / cohere_embedding / google_embedding |
embed(values_json, opts_json=None) |
SpeechModel |
openai_speech |
generate(opts_json) |
TranscriptionModel |
openai_transcription |
generate(audio_base64, media_type, opts_json=None) |
ImageModel |
openai_image / google_image |
generate(opts_json) |
VideoModel |
google_video |
generate(opts_json) |
RerankingModel |
cohere_reranking |
rerank(query, docs_json, opts_json=None) |
SearchModel |
tavily_search(api_key, base_url=None) |
search(query, opts_json=None) |
Files |
openai_files(api_key, base_url=None) |
upload_file(data_base64, media_type, opts_json=None) |
StreamIterator |
returned by Model.stream_text |
__iter__ / __next__ of StreamPart JSON strings |
Three module-level native functions back tool-call repair (all pure and
synchronous, JSON in / JSON out): tool_call_repair_context(tool_call_json, prompt_json, opts_json=None), apply_tool_call_repair(tool_call_json, opts_json, reply_json) and apply_tool_call_repair_to_result(result_json, opts_json, tool_call_id, reply_json). The repair_tool_call option drives
them for you; call them directly only to build your own loop.
All factories accept an optional base_url and return instances synchronously
(no await). The typed wrapper adds five functions: generate_text (returns a
pydantic GenerateTextResult), stream_text (yields parsed StreamPart
dicts), and parse_stream_part (validates a dict into a typed StreamPart),
plus generate_text_as_openai / stream_text_as_openai, which return OpenAI
ChatCompletion / ChatCompletionChunk models.
The wrapper's types are pydantic models in bindings/python/python/aimux/wrapper.py:
from aimux.wrapper import (
# type aliases
Role, FinishReasonUnified, ReasoningEffort, MessageContent, ContentPart,
Tool, ToolChoice, ResponseFormat, StreamPart, GenerateContent,
FileData, FileBytes, Warning, AiMuxErrorValue,
# pydantic models
TokenUsage, Usage, FinishReason, ResponseMetadata, ToolCall,
ModelMessage, FunctionTool, ProviderTool, TextContentPart,
GenerateTextOptions, GenerateTextResult, GenerateResult,
RawToolCall, ToolCallRepairContext, RepairToolCall,
# functions
generate_text, stream_text, parse_stream_part,
)AiMuxError and recorder failures raise an exception hierarchy (explicit types —
OpenAI/Anthropic SDK style, same idea as Vercel AI SDK on JS):
Exception
└── AimuxError
├── APICallError # provider call/transport failure; status when observed
├── RetryError # the retry loop gave up; reason, errors (oldest first), last_error
├── JSONParseError / InvalidResponseDataError / ToolError
├── JSONParseError / InvalidResponseDataError
├── NoSuchToolError / InvalidToolInputError / ToolCallRepairError # tool-contract errors
├── InvalidArgumentError / InvalidPromptError
├── TokenExpiredError
├── UnsupportedFunctionalityError
├── NoSuchModelError / NoSuchProviderError
├── APITimeoutError
├── RequestAbortedError
└── OtherError
Exception # the recorder's own failure type — not an AimuxError
└── RecordingError # init_recording(): code "Init" | "OpenFile" | "Spawn"; recording_try_flush(): "WriterGone" | "FlushTimeout" | "Write"
AimuxError subclasses carry only their own payload. There is no common string
discriminator and no JSON companion.
RecordingError mirrors the core's separate RecordingError type: it carries
its own code, and is not caught by except AimuxError.
Failures of the pyo3 bridge itself are not aimux types — they surface the way
pyo3 does, as Python builtins, never disguised as an AimuxError:
| scenario | raises |
|---|---|
a JSON text you passed (prompt_json / opts_json / config_json / …) does not parse |
ValueError("prompt_json: invalid JSON: …") — the message names the argument |
| closed / ended session handed back to the binding | ValueError("… is closed") |
| binding could not serialize a result | RuntimeError("serialize result: …") |
| a bridge invariant broke | RuntimeError |
| argument type errors | pyo3's own TypeError |
| panic in native code | pyo3_runtime.PanicException |
JSON that parses but has a bad value is still InvalidArgumentError.
Payload attributes belong to the class that carries them and are absent on the
others. APICallError has status / retryable / retry_ms / url /
request_body_values / response_headers / provider_code /
provider_message / response_body / data; optional values use Python's
normal None. RetryError has reason ("maxRetriesExceeded" — every
permitted attempt failed with a retryable error — or "errorNotRetryable" —
a later attempt failed non-retryably), errors — the per-attempt history,
oldest first, each itself an exception from this hierarchy — and
last_error. TokenExpiredError has status == 401, NoSuchModelError has
model_id / model_type, and NoSuchProviderError has provider_id.
from aimux import (
generate_text,
AimuxError,
APICallError,
)
try:
generate_text(model, "hi")
except APICallError as e:
# classify on status:
if e.status == 429:
... # rate limited — e.retry_ms
elif e.status == 401:
... # auth failure
elif e.status == 404:
... # model not found
except AimuxError:
... # any AiMuxError failure(The pydantic wire type named AiMuxErrorValue in aimux.wrapper is only for
stream/payload shapes, not the raised exception.)
Key shapes (mirroring the shared JSON schema):
class GenerateTextResult(BaseModel):
text: str
tool_calls: list[ToolCall]
finish_reason: FinishReason
usage: Usage
warnings: list[Warning]
raw: GenerateResultStreamPart is a RootModel over the external-tagged union dict, e.g.
{"TextDelta": {"id": ..., "delta": ...}}. Iterate dicts with
if "TextDelta" in part: (as in Streaming Generation)
or validate them with parse_stream_part(part) for attribute access.