Skip to content

feat(providers): dual transport for every provider that has a vendor SDK (21/23) - #32

Merged
araray merged 2 commits into
mainfrom
av/transport_duality
Oct 1, 2026
Merged

araray merged 2 commits into
mainfrom
av/transport_duality

Conversation

@araray

@araray araray commented Oct 1, 2026

Copy link
Copy Markdown
Owner

You were right on both counts: TypeSafe does have an official SDK, and the four SDK-only providers should follow the pattern. All six are fixed here.

Coverage: 21 of 23 providers. The two that remain single-transport have no official vendor SDK — which matches your read on Kimi and DeepSeek.


First, the false claim

My audit asserted "TypeSafe publishes no Python SDK." typesafe-sdk 0.7.2 is official, and the evidence was in llmcore's own support matrix, which lists typesafe-sdk-python v0.7.2 as a vendor clone. I wrote that table and then contradicted it.

The remaining exemptions now cite what was actually checked rather than asserting. For Kimi and DeepSeek the PyPI names that turn up are not vendor packages:

Package Publisher
deepseek Deskpai.com
deepseek-sdk Sifat Hasan
kimi-sdk no stated author or repository

SDK fallbacks added (direct stays default)

  • typesafe — via a response shim so neither existing parser changes. One honest limitation: request ids are unavailable on that path, because the SDK doesn't surface response headers.
  • mistral — mistralai v3 for chat, streaming, models, embeddings. OCR/audio/classification/moderation/FIM stay direct, because the SDK models those with its own typed resources rather than OpenAI-shaped dicts. Logged per call, so backend="sdk" never silently does nothing for half the surface.

Direct REST paths added (SDK stays default, for stated reasons)

Provider Direct path Why the SDK keeps the default
anthropic /v1/messages + streaming SDK owns prompt-caching/beta headers and retry behaviour
gemini generateContent + SSE Vertex authenticates with Google ADC, so Vertex is forced onto the SDK
deepgram /v1/listen, /v1/speak Realtime surfaces are duplex WebSocket protocols
ollama /api/chat (NDJSON), /api/tags SDK handles host resolution and framing

All four still offer backend = "httpx", and the audit now requires each to have a direct client, a selector, and its own error mapper — because a dual transport that reports failures differently isn't really dual.

The discipline: one normalization path per provider, not per transport

  • Anthropic's stream normalizer consumes event dicts from either source.
  • Gemini's reads a wire shim presenting REST JSON the way the SDK's typed objects look — so the trickiest logic in that provider isn't duplicated.
  • Deepgram's batch parser reads a JSON attribute wrapper.

Six bugs found by live calls, not by reading

  1. Mistral SDK appends its own version prefix → passing llmcore's base_url verbatim gave /v1/v1/... and a "no Route match" 404.
  2. Mistral's two streaming paths returned different shapes — SDK gave an async generator, the existing httpx path returns a coroutine resolving to one. Matched to the existing contract rather than "improved", since callers depend on it.
  3. Gemini finish_reason arrives as a plain string but readers call .name.
  4. Gemini text means different things at different levels of the response graph — joined non-thought parts on a response, the part's own string on a part. A shim property that only did the join returned "" for every part and made streaming yield empty deltas.
  5. Deepgram's credential attribute is api_key, not _api_key, and the scheme differs (Token vs Bearer).
  6. Ollama double-wrapped its own errors — a mapped ProviderError was re-caught by the generic handler and reported as "An unexpected error occurred", burying the actionable "is Ollama running?".

Live validation, both transports

Provider Result
TypeSafe ✅ system_one + models on both (noul 0.88 / 0.86)
Mistral ✅ chat, streaming, 46 models on both
Gemini ✅ identical content, finish reason, keys and streamed text
Deepgram ✅ TTS→STT round-trip, identical transcripts
Anthropic 🟡 both reach the API and return the same credit-balance error — auth and request shaping verified, completions still blocked by the zero balance
Ollama 🟡 no local server here, so error parity only — both now report the same actionable message

Final state

DUAL TRANSPORT (21/23)
SINGLE, declared (2): deepseek, kimi  — no official SDK
REMAINING GAPS: none

2158 tests pass. Lint on every touched provider is byte-identical to its pre-change baseline (checked per file, before vs after).

🤖 Generated with Claude Code

araray and others added 2 commits September 30, 2026 23:49
…Gemini

Four of the five providers that violated llmcore's house rule — call the API
directly, fall back to the vendor SDK where one exists — now follow it.
Deepgram and Ollama are next in this branch.

TypeSafe gains an SDK fallback, correcting a false claim I made in the audit:
the exemption list said "TypeSafe publishes no Python SDK" when typesafe-sdk
0.7.2 is official and was already cloned in the vendor repos. The SDK result is
normalized through a small response shim so neither existing parser changes;
the one honest limitation is that request ids are unavailable on that path,
because the SDK does not surface response headers.

Mistral gains an SDK fallback for chat, streaming, models and embeddings. Its
specialist endpoints — OCR, audio, classification, moderation, FIM — stay on
direct REST even when backend="sdk", because the SDK models those with its own
typed resources rather than OpenAI-shaped dicts and converting them all would
mean maintaining a second translation layer. That choice is logged per call, so
backend="sdk" never silently does nothing for half the surface.

Anthropic and Gemini gain direct REST paths; both keep the SDK as the default,
since those SDKs own prompt-caching headers, beta features, retries and (for
Vertex) an ADC token exchange that llmcore would otherwise have to track.
Gemini's Vertex mode is forced onto the SDK for exactly that reason.

The discipline throughout is one normalization path per provider rather than one
per transport. Anthropic's streaming normalizer now consumes event dicts from
either source. Gemini's reads a wire shim that presents REST JSON the way the
SDK's typed objects look, so the trickiest logic in that provider is not
duplicated.

Three bugs found by live calls rather than by reading:

The Mistral SDK appends its own version prefix, so passing llmcore's base_url
verbatim produced /v1/v1/... and a "no Route match" 404.

The two Mistral streaming paths returned different shapes — the SDK path handed
back an async generator while the existing httpx path returns a coroutine that
resolves to one. Matched to the existing contract rather than "improved", since
callers depend on it.

Gemini's wire shim had two collisions with the SDK's object model. finish_reason
arrives as a plain string but readers call .name, so enum-valued fields are now
wrapped. And `text` means different things at different levels of the graph — the
joined non-thought parts on a response, the part's own string on a part — so a
property that only did the join returned "" for every part and made streaming
yield empty deltas.

Live-validated on both transports: TypeSafe (system_one and models), Mistral
(chat, stream, models), Gemini (identical content, finish reason, keys and
streamed text). Anthropic reaches the API on both and returns the same
credit-balance error, so auth and request shaping are verified but completions
are not — that account still has no balance.

1729 tests pass. The two remaining audit failures are deepgram and ollama, which
is the audit holding this branch to its own promise.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Completes the transport work. Deepgram gains a direct REST path for its batch
endpoints (/v1/listen, /v1/speak, /v1/models); its realtime surfaces — streaming
STT and TTS, the voice agent, Flux v2 — stay on the SDK because they are duplex
WebSocket protocols with their own framing and finalize semantics, and that is
logged per call so backend="httpx" never silently does nothing. Ollama gains
/api/chat with NDJSON streaming and /api/tags.

Coverage is now 21 of 23 providers. The two that remain single-transport are
DeepSeek and Kimi, which publish no official Python SDK — the PyPI names that
turn up (`deepseek` from Deskpai.com, `deepseek-sdk` from Sifat Hasan,
`kimi-sdk` with no stated author or repository) are third-party, and the
exemption list now cites that rather than asserting it.

Two bugs found by live calls. Deepgram's credential attribute is `api_key`, not
`_api_key`, and the auth scheme differs between an API key (Token) and an access
token (Bearer). And Ollama double-wrapped its own errors: a ProviderError raised
by the direct path was re-caught by the generic handler and reported as "An
unexpected error occurred", which buried the actionable "is Ollama running?"
message and made the two transports describe the same condition differently.

The audit grew checks for the four providers whose SDK stays the default
(anthropic, gemini, deepgram, ollama): each must still have a direct client, a
transport selector, and its own error mapper. Keeping the SDK as the default
there is a documented deviation — those SDKs own prompt-caching headers, ADC
token exchange, WebSocket framing and host resolution — so the audit lists them
by name to keep the deviation visible rather than implicit.

Live-validated on both transports: Deepgram TTS and STT round-trip with
identical transcripts. Ollama has no local server here, so only error parity is
verified — both transports now report the same actionable message.

2158 tests pass. Lint on every touched provider is identical to its pre-change
baseline.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Copilot AI balanced review requested due to automatic review settings October 1, 2026 03:02

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Several direct paths mishandle responses or configuration, leak clients, and expose inconsistent streaming errors.

Review effort: Balanced
Findings: 1 High severity · 17 Medium severity · 3 Low severity

Open (21)

And 1 more that still need to be addressed.

What changed in this PR

Adds dual-transport support to six providers, reaching 21 of 23 providers.

Changes:

  • Adds SDK fallbacks for TypeSafe and Mistral.
  • Adds direct REST transports for Anthropic, Gemini, Deepgram, and Ollama.
  • Updates dependencies, transport audits, and documentation.
File Description
tests/​providers/​test_transport_duality.py Expands transport audits.
src/​llmcore/​providers/​typesafe_provider.py Adds TypeSafe SDK transport.
src/​llmcore/​providers/​ollama_provider.py Adds direct Ollama REST transport.
src/​llmcore/​providers/​mistral_provider.py Adds Mistral SDK fallback.
src/​llmcore/​providers/​gemini_provider.py Adds Gemini REST transport and response shim.
src/​llmcore/​providers/​deepgram_provider.py Adds direct batch audio endpoints.
src/​llmcore/​providers/​anthropic_provider.py Adds direct Messages API support.
pyproject.toml Adds SDK dependencies.
docs/​PROVIDER_SUPPORT_MATRIX.md Documents transport coverage.
CHANGELOG.md Records provider changes.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines 505 to +507
self._client = AsyncDeepgramClient(**client_kwargs)
self._backend = self._resolve_backend(config.get("backend"))
self._http: Any = None
Comment on lines +355 to +357
async with client.stream(
"POST", "/v1/messages", json={**body, "stream": True}
) as resp:
Comment on lines +567 to +568
credential = self.api_key or self._access_token or ""
scheme = "Token" if self.api_key else "Bearer"
Comment on lines +569 to +573
self._http = httpx.AsyncClient(
base_url=self._direct_base_url(),
headers={"Authorization": f"{scheme} {credential}"},
timeout=getattr(self, "_timeout", 60.0),
)
Comment on lines +569 to +573
self._http = httpx.AsyncClient(
base_url=self._direct_base_url(),
headers={"Authorization": f"{scheme} {credential}"},
timeout=getattr(self, "_timeout", 60.0),
)
"""

DUAL = ("fal", "elevenlabs", "replicate", "higgsfield")
DUAL = ("fal", "elevenlabs", "replicate", "higgsfield", "typesafe", "mistral")
def test_the_direct_path_maps_its_own_errors(self, provider):
"""A direct path that raised raw httpx errors would make the two
transports report the same condition differently."""
assert "_raise_direct_status" in _module_source(provider), (
Comment thread CHANGELOG.md
Each direct path maps failures to the same exceptions the SDK path raises,
because a dual transport that reports failures differently is not really dual.

### Fixed — five bugs found by live calls rather than by reading
Comment thread pyproject.toml
# SDK today — `mistralai` v3.x is a candidate second backend, see
# docs/PROVIDER_MODERNIZATION_PLAN.md).
mistral = ["httpx>=0.27.0"]
mistral = ["httpx>=0.27.0", "mistralai>=3.0.0"]
#: message framing, keepalives and finalize semantics. Reimplementing those
#: would be rebuilding the part of the SDK that genuinely earns its keep, so
#: they stay on the SDK even when ``backend = "httpx"``, and say so.
_DIRECT_CAPABILITIES: frozenset[str] = frozenset({"transcribe", "speak", "models"})
@araray araray self-assigned this Oct 1, 2026
@araray
araray merged commit bf3d961 into main Oct 1, 2026
2 of 4 checks passed
@araray
araray deleted the av/transport_duality branch October 1, 2026 03:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants