Skip to content

docs(readme): refresh for v0.53, correct stale and overstated claims - #33

Merged
araray merged 1 commit into
mainfrom
av/readme_refresh
Oct 1, 2026
Merged

araray merged 1 commit into
mainfrom
av/readme_refresh

Conversation

@araray

@araray araray commented Oct 1, 2026

Copy link
Copy Markdown
Owner

The README still described v0.50: 13 providers, GPT-4o and o3-mini as flagships, Claude Opus 4.5, Gemini 2.5, and a "What's New" section three releases old. The media and runtimes subsystems were missing entirely.

Two claims were wrong, not just stale

The coverage badge said 85%. Measured coverage is 66%.

That 85% is a fail_under target in pyproject.toml that the suite doesn't meet — the badge was reporting an aspiration as a fact. I measured it rather than carrying the number forward: 66.28% statement+branch on the profile CI runs. Badge removed, real number in the statistics table, with a note explaining the gap and that CI doesn't enforce the gate.

The provider table listed 13 providers. There are 23.

Every number is measured, not estimated

Providers 23 (+11 alias spellings)
Transport 21 of 23 have two — 17 direct-first w/ SDK fallback, 4 SDK-first w/ direct available
Model cards 2,319 across 22 providers
Media capabilities 19
Install extras 35
Tests 6,396 collected; 5,622 in the default profile
Source ~156k lines / 318 modules

Plus a card-type breakdown (1,716 chat, 156 image-generation, 125 tts, …) and per-provider card counts.

Model names now come from the registry

Rather than naming models from memory, the provider tables list what's actually in the bundled cards: gpt-6-sol, claude-opus-5-5, gemini-3.8-flash, glm-5.3, deepseek-v4-pro, kimi-k3, mistral-large-3, grok-4.1-20251117.

A verification pass caught eight ids that were family prefixes rather than real cards — claude-haiku-4-5, grok-4.1, magistral-medium, qwen3-vl, gemma3, llama3.3 and others — sitting under a heading that claimed they came from the registry. All pinned to exact ids.

There's also an explicit caveat that model names move fast, pointing readers at llm.list_models() over anything written in a README.

New sections

  • Generative media — routing, the per-capability execution class (why the same call returns a result from OpenAI and a job from fal), the 19 capabilities, artifacts and VoiceConsent.
  • Remote GPU runtimes — with the safety model stated plainly, including that nothing there can spend money today because the Colab backend isn't implemented.
  • Invariant guards — the four suites that make whole-repo mistakes fail loudly.
  • cardctl — including why doctor exists (three media providers once shipped with no adapter and nothing complained).

Architecture diagram, installation (all 35 extras), and project structure updated too.

A real bug the coverage run exposed

The cardctl coverage guard read the live PROVIDER_MAP, which several suites mutate with monkeypatch.setitem. A leaked test double could make it report a missing adapter — it failed in the full run while passing alone. Both that guard and the transport audit now count only classes defined under llmcore.providers, immune to leakage while still catching a genuinely unadaptered provider.

Verified before commit

A script checked that every model id named resolves to a real card, every docs/ link resolves, every named extra exists, and the provider/card/capability/extra counts match the code:

ALL CLAIMS VERIFIED
checked: 40 model ids, 23 providers, 2319 cards, 35 extras

1,743 tests pass across providers and media; 171 across tools and the transport audit.

Note on tone

Kept technical and descriptive per your ask — no "blazing fast", no "enterprise-grade". The one place I softened rather than removed is the tagline, now describing what the library does instead of asserting it's production-ready. The runtimes section says outright that the backend is unimplemented, and the coverage note says outright that the project misses its own gate; both seemed more useful than leaving them flattering.

🤖 Generated with Claude Code

The README still described v0.50: 13 providers, GPT-4o and o3-mini as flagships,
Claude Opus 4.5, Gemini 2.5, and a "What's New" section three releases old. The
media and runtimes subsystems were absent entirely.

Rewrites the header, provider tables, installation, architecture diagram,
model-card and testing sections, and adds sections for the two missing
subsystems. The tagline drops "Production-Ready Framework" for a description of
what the library actually does.

Two claims were not just stale but wrong, and both are corrected rather than
restated.

The coverage badge said 85%. Measured coverage on the profile CI runs is 66%.
The 85% figure is a `fail_under` target in pyproject that the suite does not
meet, so the badge was reporting an aspiration as a fact. The badge is gone and
the real number is in the statistics table, with a note explaining the gap and
that CI does not enforce the gate.

The provider table listed 13 providers when there are 23, and named models that
no longer exist in the bundled card registry.

Every number in the new README is measured from the repository rather than
estimated: 23 providers, 2,319 model cards, 19 media capabilities, 35 extras,
6,396 tests collected, ~156k lines across 318 modules. Model ids in the provider
tables are exact card ids, not family prefixes — a verification pass caught
eight that were prefixes (`claude-haiku-4-5`, `grok-4.1`, `magistral-medium`,
`qwen3-vl` and others) under a heading claiming they came from the registry.

Also fixes a real test-isolation bug the coverage run exposed: the cardctl
coverage guard read the live PROVIDER_MAP, which several suites mutate with
monkeypatch.setitem, so a leaked test double could make it report a missing
adapter. Both that guard and the transport audit now count only classes defined
under llmcore.providers, which makes them immune to leakage while still catching
a genuinely unadaptered provider.

A verification script checked, before commit, that every model id named resolves
to a real card, every docs/ link resolves, every named extra exists, and the
provider/card/capability/extra counts match the code.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Copilot AI balanced review requested due to automatic review settings October 1, 2026 03:41
@araray araray self-assigned this Oct 1, 2026
@araray
araray merged commit c2cd0c2 into main Oct 1, 2026
3 of 4 checks passed
@araray
araray deleted the av/readme_refresh branch October 1, 2026 03:43

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Several README claims are inaccurate, including automatic runtime limit enforcement, runnable runtime examples, consent coverage, and model-card provider counts.

Review effort: Balanced
Findings: 4 Medium severity · 1 Low severity

Open (5)
What changed in this PR

Refreshes the README for v0.53 and hardens repository-wide provider audits against injected test doubles.

Changes:

  • Updates provider, model-card, media, runtime, installation, and testing documentation.
  • Adds media/runtime architecture and usage guidance.
  • Filters provider audits to llmcore provider classes.
File Description
README.md Comprehensive v0.53 documentation refresh.
tests/​tools/​test_cardctl_coverage.py Excludes injected doubles from adapter coverage checks.
tests/​providers/​test_transport_duality.py Excludes injected doubles from transport audits.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread README.md
- **Test hygiene**: The suite now treats warning cleanup as part of release readiness; optional SDK and infrastructure tests are skipped only behind explicit environment gates.
| Category | What it does |
|----------|--------------|
| **🔌 Providers** | 23 vendors, one `chat()` call. Streaming, tool calling, structured output, reasoning extraction, vision, exact tokenizers where the vendor exposes one |
Comment thread README.md
|----------|--------------|
| **🔌 Providers** | 23 vendors, one `chat()` call. Streaming, tool calling, structured output, reasoning extraction, vision, exact tokenizers where the vendor exposes one |
| **🎨 Generative media** | `llm.media` — image generate/edit/upscale, video generate/interpolate, TTS (+streaming), ASR, music, SFX, voice design. Async jobs with polling, a webhook receiver, and a content-addressed artifact store |
| **🖥️ Remote GPU runtimes** | `llm.runtimes` — size an open-weights model, provision compute, serve it, and attach the endpoint as a provider instance. Spend ceilings and idle reaping are enforced, not optional |
Comment thread README.md
Comment on lines +103 to +105
- **`VoiceConsent`** — synthetic speech carries whose voice it is and whether
the vendor considers it cleared, so callers can refuse unverified clones
without a second API call.
Comment thread README.md
is reachable through the same `llm.chat()` as a hosted API.

```python
plan = await llm.runtimes.estimate("Qwen/Qwen3-30B-A3B-Instruct-2507", context_length=32768)
Comment thread README.md
|---|---|
| **Providers** | **23** behind one interface (plus 11 alias spellings) |
| **Transport** | **21 of 23** have two transports — 17 direct-first with an SDK fallback, 4 SDK-first with direct available. The other 2 have no vendor SDK |
| **Model cards** | **2,319** across 22 providers, generated from live APIs |
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants