docs(readme): refresh for v0.53, correct stale and overstated claims - #33
Merged
Merged
Conversation
The README still described v0.50: 13 providers, GPT-4o and o3-mini as flagships, Claude Opus 4.5, Gemini 2.5, and a "What's New" section three releases old. The media and runtimes subsystems were absent entirely. Rewrites the header, provider tables, installation, architecture diagram, model-card and testing sections, and adds sections for the two missing subsystems. The tagline drops "Production-Ready Framework" for a description of what the library actually does. Two claims were not just stale but wrong, and both are corrected rather than restated. The coverage badge said 85%. Measured coverage on the profile CI runs is 66%. The 85% figure is a `fail_under` target in pyproject that the suite does not meet, so the badge was reporting an aspiration as a fact. The badge is gone and the real number is in the statistics table, with a note explaining the gap and that CI does not enforce the gate. The provider table listed 13 providers when there are 23, and named models that no longer exist in the bundled card registry. Every number in the new README is measured from the repository rather than estimated: 23 providers, 2,319 model cards, 19 media capabilities, 35 extras, 6,396 tests collected, ~156k lines across 318 modules. Model ids in the provider tables are exact card ids, not family prefixes — a verification pass caught eight that were prefixes (`claude-haiku-4-5`, `grok-4.1`, `magistral-medium`, `qwen3-vl` and others) under a heading claiming they came from the registry. Also fixes a real test-isolation bug the coverage run exposed: the cardctl coverage guard read the live PROVIDER_MAP, which several suites mutate with monkeypatch.setitem, so a leaked test double could make it report a missing adapter. Both that guard and the transport audit now count only classes defined under llmcore.providers, which makes them immune to leakage while still catching a genuinely unadaptered provider. A verification script checked, before commit, that every model id named resolves to a real card, every docs/ link resolves, every named extra exists, and the provider/card/capability/extra counts match the code. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Several README claims are inaccurate, including automatic runtime limit enforcement, runnable runtime examples, consent coverage, and model-card provider counts.
Review effort: Balanced
Findings: 4
Open (5)
What changed in this PR
Refreshes the README for v0.53 and hardens repository-wide provider audits against injected test doubles.
Changes:
- Updates provider, model-card, media, runtime, installation, and testing documentation.
- Adds media/runtime architecture and usage guidance.
- Filters provider audits to llmcore provider classes.
| File | Description |
|---|---|
README.md |
Comprehensive v0.53 documentation refresh. |
tests/tools/test_cardctl_coverage.py |
Excludes injected doubles from adapter coverage checks. |
tests/providers/test_transport_duality.py |
Excludes injected doubles from transport audits. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
| - **Test hygiene**: The suite now treats warning cleanup as part of release readiness; optional SDK and infrastructure tests are skipped only behind explicit environment gates. | ||
| | Category | What it does | | ||
| |----------|--------------| | ||
| | **🔌 Providers** | 23 vendors, one `chat()` call. Streaming, tool calling, structured output, reasoning extraction, vision, exact tokenizers where the vendor exposes one | |
| |----------|--------------| | ||
| | **🔌 Providers** | 23 vendors, one `chat()` call. Streaming, tool calling, structured output, reasoning extraction, vision, exact tokenizers where the vendor exposes one | | ||
| | **🎨 Generative media** | `llm.media` — image generate/edit/upscale, video generate/interpolate, TTS (+streaming), ASR, music, SFX, voice design. Async jobs with polling, a webhook receiver, and a content-addressed artifact store | | ||
| | **🖥️ Remote GPU runtimes** | `llm.runtimes` — size an open-weights model, provision compute, serve it, and attach the endpoint as a provider instance. Spend ceilings and idle reaping are enforced, not optional | |
Comment on lines
+103
to
+105
| - **`VoiceConsent`** — synthetic speech carries whose voice it is and whether | ||
| the vendor considers it cleared, so callers can refuse unverified clones | ||
| without a second API call. |
| is reachable through the same `llm.chat()` as a hosted API. | ||
|
|
||
| ```python | ||
| plan = await llm.runtimes.estimate("Qwen/Qwen3-30B-A3B-Instruct-2507", context_length=32768) |
| |---|---| | ||
| | **Providers** | **23** behind one interface (plus 11 alias spellings) | | ||
| | **Transport** | **21 of 23** have two transports — 17 direct-first with an SDK fallback, 4 SDK-first with direct available. The other 2 have no vendor SDK | | ||
| | **Model cards** | **2,319** across 22 providers, generated from live APIs | |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.


The README still described v0.50: 13 providers, GPT-4o and o3-mini as flagships, Claude Opus 4.5, Gemini 2.5, and a "What's New" section three releases old. The media and runtimes subsystems were missing entirely.
Two claims were wrong, not just stale
The coverage badge said 85%. Measured coverage is 66%.
That 85% is a
fail_undertarget inpyproject.tomlthat the suite doesn't meet — the badge was reporting an aspiration as a fact. I measured it rather than carrying the number forward:66.28%statement+branch on the profile CI runs. Badge removed, real number in the statistics table, with a note explaining the gap and that CI doesn't enforce the gate.The provider table listed 13 providers. There are 23.
Every number is measured, not estimated
Plus a card-type breakdown (1,716 chat, 156 image-generation, 125 tts, …) and per-provider card counts.
Model names now come from the registry
Rather than naming models from memory, the provider tables list what's actually in the bundled cards:
gpt-6-sol,claude-opus-5-5,gemini-3.8-flash,glm-5.3,deepseek-v4-pro,kimi-k3,mistral-large-3,grok-4.1-20251117.A verification pass caught eight ids that were family prefixes rather than real cards —
claude-haiku-4-5,grok-4.1,magistral-medium,qwen3-vl,gemma3,llama3.3and others — sitting under a heading that claimed they came from the registry. All pinned to exact ids.There's also an explicit caveat that model names move fast, pointing readers at
llm.list_models()over anything written in a README.New sections
VoiceConsent.cardctl— including whydoctorexists (three media providers once shipped with no adapter and nothing complained).Architecture diagram, installation (all 35 extras), and project structure updated too.
A real bug the coverage run exposed
The cardctl coverage guard read the live
PROVIDER_MAP, which several suites mutate withmonkeypatch.setitem. A leaked test double could make it report a missing adapter — it failed in the full run while passing alone. Both that guard and the transport audit now count only classes defined underllmcore.providers, immune to leakage while still catching a genuinely unadaptered provider.Verified before commit
A script checked that every model id named resolves to a real card, every
docs/link resolves, every named extra exists, and the provider/card/capability/extra counts match the code:1,743 tests pass across providers and media; 171 across tools and the transport audit.
Note on tone
Kept technical and descriptive per your ask — no "blazing fast", no "enterprise-grade". The one place I softened rather than removed is the tagline, now describing what the library does instead of asserting it's production-ready. The runtimes section says outright that the backend is unimplemented, and the coverage note says outright that the project misses its own gate; both seemed more useful than leaving them flattering.
🤖 Generated with Claude Code