feat(providers): dual transport for every provider that has a vendor SDK (21/23) - #32
Merged
Merged
Conversation
…Gemini Four of the five providers that violated llmcore's house rule — call the API directly, fall back to the vendor SDK where one exists — now follow it. Deepgram and Ollama are next in this branch. TypeSafe gains an SDK fallback, correcting a false claim I made in the audit: the exemption list said "TypeSafe publishes no Python SDK" when typesafe-sdk 0.7.2 is official and was already cloned in the vendor repos. The SDK result is normalized through a small response shim so neither existing parser changes; the one honest limitation is that request ids are unavailable on that path, because the SDK does not surface response headers. Mistral gains an SDK fallback for chat, streaming, models and embeddings. Its specialist endpoints — OCR, audio, classification, moderation, FIM — stay on direct REST even when backend="sdk", because the SDK models those with its own typed resources rather than OpenAI-shaped dicts and converting them all would mean maintaining a second translation layer. That choice is logged per call, so backend="sdk" never silently does nothing for half the surface. Anthropic and Gemini gain direct REST paths; both keep the SDK as the default, since those SDKs own prompt-caching headers, beta features, retries and (for Vertex) an ADC token exchange that llmcore would otherwise have to track. Gemini's Vertex mode is forced onto the SDK for exactly that reason. The discipline throughout is one normalization path per provider rather than one per transport. Anthropic's streaming normalizer now consumes event dicts from either source. Gemini's reads a wire shim that presents REST JSON the way the SDK's typed objects look, so the trickiest logic in that provider is not duplicated. Three bugs found by live calls rather than by reading: The Mistral SDK appends its own version prefix, so passing llmcore's base_url verbatim produced /v1/v1/... and a "no Route match" 404. The two Mistral streaming paths returned different shapes — the SDK path handed back an async generator while the existing httpx path returns a coroutine that resolves to one. Matched to the existing contract rather than "improved", since callers depend on it. Gemini's wire shim had two collisions with the SDK's object model. finish_reason arrives as a plain string but readers call .name, so enum-valued fields are now wrapped. And `text` means different things at different levels of the graph — the joined non-thought parts on a response, the part's own string on a part — so a property that only did the join returned "" for every part and made streaming yield empty deltas. Live-validated on both transports: TypeSafe (system_one and models), Mistral (chat, stream, models), Gemini (identical content, finish reason, keys and streamed text). Anthropic reaches the API on both and returns the same credit-balance error, so auth and request shaping are verified but completions are not — that account still has no balance. 1729 tests pass. The two remaining audit failures are deepgram and ollama, which is the audit holding this branch to its own promise. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Completes the transport work. Deepgram gains a direct REST path for its batch endpoints (/v1/listen, /v1/speak, /v1/models); its realtime surfaces — streaming STT and TTS, the voice agent, Flux v2 — stay on the SDK because they are duplex WebSocket protocols with their own framing and finalize semantics, and that is logged per call so backend="httpx" never silently does nothing. Ollama gains /api/chat with NDJSON streaming and /api/tags. Coverage is now 21 of 23 providers. The two that remain single-transport are DeepSeek and Kimi, which publish no official Python SDK — the PyPI names that turn up (`deepseek` from Deskpai.com, `deepseek-sdk` from Sifat Hasan, `kimi-sdk` with no stated author or repository) are third-party, and the exemption list now cites that rather than asserting it. Two bugs found by live calls. Deepgram's credential attribute is `api_key`, not `_api_key`, and the auth scheme differs between an API key (Token) and an access token (Bearer). And Ollama double-wrapped its own errors: a ProviderError raised by the direct path was re-caught by the generic handler and reported as "An unexpected error occurred", which buried the actionable "is Ollama running?" message and made the two transports describe the same condition differently. The audit grew checks for the four providers whose SDK stays the default (anthropic, gemini, deepgram, ollama): each must still have a direct client, a transport selector, and its own error mapper. Keeping the SDK as the default there is a documented deviation — those SDKs own prompt-caching headers, ADC token exchange, WebSocket framing and host resolution — so the audit lists them by name to keep the deviation visible rather than implicit. Live-validated on both transports: Deepgram TTS and STT round-trip with identical transcripts. Ollama has no local server here, so only error parity is verified — both transports now report the same actionable message. 2158 tests pass. Lint on every touched provider is identical to its pre-change baseline. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Several direct paths mishandle responses or configuration, leak clients, and expose inconsistent streaming errors.
Review effort: Balanced
Findings: 1
Open (21)
Persist the configured direct REST base URL · New Map streaming transport errors inside the generator · New Prefer access tokens over API keys for direct transport · New Send the configured Deepgram session ID · New Close and clear the Deepgram HTTP client · New Preserve response ID and model version in REST normalization · New Close and clear the Gemini HTTP client · New Map Gemini streaming transport errors · New Close the Mistral SDK during teardown · New Normalize Ollama REST model responses before parsing · New Close and clear the Ollama HTTP client · New Encode direct Ollama images as base64 · New Map streaming httpx errors to ProviderError · New Pass configured headers to the TypeSafe SDK · New Preserve SDK error status and retryability · New Forward timeout and extra headers to the SDK · New Add behavioral tests for TypeSafe and Mistral SDKs · New Add behavioral streaming error mapping tests · New Correct the section bug count · New Update Mistral transport package comments · New
And 1 more that still need to be addressed.
What changed in this PR
Adds dual-transport support to six providers, reaching 21 of 23 providers.
Changes:
- Adds SDK fallbacks for TypeSafe and Mistral.
- Adds direct REST transports for Anthropic, Gemini, Deepgram, and Ollama.
- Updates dependencies, transport audits, and documentation.
| File | Description |
|---|---|
tests/providers/test_transport_duality.py |
Expands transport audits. |
src/llmcore/providers/typesafe_provider.py |
Adds TypeSafe SDK transport. |
src/llmcore/providers/ollama_provider.py |
Adds direct Ollama REST transport. |
src/llmcore/providers/mistral_provider.py |
Adds Mistral SDK fallback. |
src/llmcore/providers/gemini_provider.py |
Adds Gemini REST transport and response shim. |
src/llmcore/providers/deepgram_provider.py |
Adds direct batch audio endpoints. |
src/llmcore/providers/anthropic_provider.py |
Adds direct Messages API support. |
pyproject.toml |
Adds SDK dependencies. |
docs/PROVIDER_SUPPORT_MATRIX.md |
Documents transport coverage. |
CHANGELOG.md |
Records provider changes. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Comment on lines
505
to
+507
| self._client = AsyncDeepgramClient(**client_kwargs) | ||
| self._backend = self._resolve_backend(config.get("backend")) | ||
| self._http: Any = None |
Comment on lines
+355
to
+357
| async with client.stream( | ||
| "POST", "/v1/messages", json={**body, "stream": True} | ||
| ) as resp: |
Comment on lines
+567
to
+568
| credential = self.api_key or self._access_token or "" | ||
| scheme = "Token" if self.api_key else "Bearer" |
Comment on lines
+569
to
+573
| self._http = httpx.AsyncClient( | ||
| base_url=self._direct_base_url(), | ||
| headers={"Authorization": f"{scheme} {credential}"}, | ||
| timeout=getattr(self, "_timeout", 60.0), | ||
| ) |
Comment on lines
+569
to
+573
| self._http = httpx.AsyncClient( | ||
| base_url=self._direct_base_url(), | ||
| headers={"Authorization": f"{scheme} {credential}"}, | ||
| timeout=getattr(self, "_timeout", 60.0), | ||
| ) |
| """ | ||
|
|
||
| DUAL = ("fal", "elevenlabs", "replicate", "higgsfield") | ||
| DUAL = ("fal", "elevenlabs", "replicate", "higgsfield", "typesafe", "mistral") |
| def test_the_direct_path_maps_its_own_errors(self, provider): | ||
| """A direct path that raised raw httpx errors would make the two | ||
| transports report the same condition differently.""" | ||
| assert "_raise_direct_status" in _module_source(provider), ( |
| Each direct path maps failures to the same exceptions the SDK path raises, | ||
| because a dual transport that reports failures differently is not really dual. | ||
|
|
||
| ### Fixed — five bugs found by live calls rather than by reading |
| # SDK today — `mistralai` v3.x is a candidate second backend, see | ||
| # docs/PROVIDER_MODERNIZATION_PLAN.md). | ||
| mistral = ["httpx>=0.27.0"] | ||
| mistral = ["httpx>=0.27.0", "mistralai>=3.0.0"] |
| #: message framing, keepalives and finalize semantics. Reimplementing those | ||
| #: would be rebuilding the part of the SDK that genuinely earns its keep, so | ||
| #: they stay on the SDK even when ``backend = "httpx"``, and say so. | ||
| _DIRECT_CAPABILITIES: frozenset[str] = frozenset({"transcribe", "speak", "models"}) |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.



You were right on both counts: TypeSafe does have an official SDK, and the four SDK-only providers should follow the pattern. All six are fixed here.
Coverage: 21 of 23 providers. The two that remain single-transport have no official vendor SDK — which matches your read on Kimi and DeepSeek.
First, the false claim
My audit asserted "TypeSafe publishes no Python SDK."
typesafe-sdk0.7.2 is official, and the evidence was in llmcore's own support matrix, which liststypesafe-sdk-pythonv0.7.2 as a vendor clone. I wrote that table and then contradicted it.The remaining exemptions now cite what was actually checked rather than asserting. For Kimi and DeepSeek the PyPI names that turn up are not vendor packages:
deepseekdeepseek-sdkkimi-sdkSDK fallbacks added (direct stays default)
typesafe— via a response shim so neither existing parser changes. One honest limitation: request ids are unavailable on that path, because the SDK doesn't surface response headers.mistral—mistralaiv3 for chat, streaming, models, embeddings. OCR/audio/classification/moderation/FIM stay direct, because the SDK models those with its own typed resources rather than OpenAI-shaped dicts. Logged per call, sobackend="sdk"never silently does nothing for half the surface.Direct REST paths added (SDK stays default, for stated reasons)
anthropic/v1/messages+ streaminggeminigenerateContent+ SSEdeepgram/v1/listen,/v1/speakollama/api/chat(NDJSON),/api/tagsAll four still offer
backend = "httpx", and the audit now requires each to have a direct client, a selector, and its own error mapper — because a dual transport that reports failures differently isn't really dual.The discipline: one normalization path per provider, not per transport
Six bugs found by live calls, not by reading
base_urlverbatim gave/v1/v1/...and a "no Route match" 404.finish_reasonarrives as a plain string but readers call.name.textmeans different things at different levels of the response graph — joined non-thought parts on a response, the part's own string on a part. A shim property that only did the join returned""for every part and made streaming yield empty deltas.api_key, not_api_key, and the scheme differs (TokenvsBearer).ProviderErrorwas re-caught by the generic handler and reported as "An unexpected error occurred", burying the actionable "is Ollama running?".Live validation, both transports
system_one+modelson both (noul 0.88 / 0.86)Final state
2158 tests pass. Lint on every touched provider is byte-identical to its pre-change baseline (checked per file, before vs after).
🤖 Generated with Claude Code