An async Python framework for LLM applications: chat, agents, RAG, generative media, and remote GPU runtimes — behind one interface.
At a glance • Providers • Quickstart • Media • Runtimes • Installation • Architecture
llmcore gives you one async interface over 23 model providers, plus the subsystems that usually get rebuilt per project: conversation persistence, retrieval, a generative-media layer (image/audio/video), tool-using agents with sandboxed execution, and provisioning for remotely served open-weights models.
It is a library, not a service. There is no daemon, no broker, and no required
infrastructure — pip install llmcore[openai] and a config file are enough. Every
subsystem beyond chat is optional and off until configured.
Counts are measured from this repository, not estimated.
| Providers | 23 behind one interface (plus 11 alias spellings) |
| Transport | 21 of 23 have two transports — 17 direct-first with an SDK fallback, 4 SDK-first with direct available. The other 2 have no vendor SDK |
| Model cards | 2,319 across 22 providers, generated from live APIs |
| Media capabilities | 19 — image, video, speech, music, SFX, OCR, voice design |
| Install extras | 35, so you install only the providers you use |
| Tests | 6,396 collected; 5,622 run in the default (no-infrastructure) profile |
| Coverage | 66% statement+branch, measured on that profile |
| Source | ~156k lines across 318 modules |
Coverage is reported as measured.
pyproject.tomlsets afail_underof 85%, which the suite does not currently meet — that gate is a target, not a description of today. CI runs the suite without the coverage gate, so this number is informational rather than enforced.
Cards carry context windows, capability flags, pricing and lifecycle for every model llmcore can reach, so cost and capability checks happen locally instead of by trial and error.
| Model type | Cards | Model type | Cards | |
|---|---|---|---|---|
| chat | 1,716 | audio | 71 | |
| image-generation | 156 | stt | 58 | |
| tts | 125 | multimodal | 13 | |
| embedding | 87 | video-generation | 8 | |
| vision | 72 | ocr / code / other | 13 |
| Category | What it does |
|---|---|
| 🔌 Providers | 23 vendors, one chat() call. Streaming, tool calling, structured output, reasoning extraction, vision, exact tokenizers where the vendor exposes one |
| 🎨 Generative media | llm.media — image generate/edit/upscale, video generate/interpolate, TTS (+streaming), ASR, music, SFX, voice design. Async jobs with polling, a webhook receiver, and a content-addressed artifact store |
| 🖥️ Remote GPU runtimes | llm.runtimes — size an open-weights model, provision compute, serve it, and attach the endpoint as a provider instance. Spend ceilings and idle reaping are enforced, not optional |
| 💬 Sessions | Persistent conversations over SQLite/PostgreSQL/JSON, transient sessions, per-call usage via chat_with_usage() |
| 🔍 RAG | ChromaDB/pgvector, semantic search, context injection, external-RAG bridge |
| 🌐 Web search | Bright Data, Serper.dev, SerpApi, Semantic Scholar (keyless) |
| 🤖 Agents | 8-phase cognitive cycle, goal classification, fast-path execution, circuit breaker, personas |
| 🔒 Sandboxing | Docker and VM/SSH isolation with security policies and output tracking |
| 👤 Human-in-the-loop | Risk assessment, approval workflows, audit logging |
| 📊 Observability | Structured events, metrics, execution replay, context diagnostics |
| 📚 Model cards | 2,319 cards with capability validation and cost estimation, refreshed by cardctl |
llmcore.media— a first-class generative-media subsystem. Capability protocols rather than per-vendor methods, sollm.media.images.generate(...)routes to whichever configured provider can serve it. Execution class is declared per capability, so the same call returns a result from OpenAI and a pollable job from fal, andmedia.wait()absorbs the difference.- Media adapters for OpenAI, Google (Veo), Deepgram, ElevenLabs, fal, Replicate, Hugging Face and Higgsfield.
VoiceConsent— synthetic speech carries whose voice it is and whether the vendor considers it cleared, so callers can refuse unverified clones without a second API call.- Webhook receiver — signed single-use callbacks for long jobs. Polling stays the fallback, so no public ingress is required.
llmcore.runtimes— provision remote GPU compute and attach it as a provider. Off by default;up()refuses without explicit spend confirmation.- Dual transport — 21 of 23 providers now call the vendor API directly with
the SDK as a fallback (or the reverse, where an SDK owns something llmcore
should not reimplement). See
PROVIDER_SUPPORT_MATRIX.md§7.1. cardctl doctor— audits that every registered provider has a card adapter, so a new provider cannot ship without model cards.
import asyncio
from llmcore import LLMCore
async def main():
async with await LLMCore.create() as llm:
# Simple question
response = await llm.chat("What is the capital of France?")
print(response)
# Streaming response
stream = await llm.chat("Tell me a story.", stream=True)
async for chunk in stream:
print(chunk, end="", flush=True)
asyncio.run(main())async def conversation():
async with await LLMCore.create() as llm:
# First message - sets context
await llm.chat(
"My name is Alex and I love astronomy.",
session_id="alex_chat",
system_message="You are a friendly science tutor."
)
# Follow-up - LLM remembers context
response = await llm.chat(
"What should I observe tonight?",
session_id="alex_chat"
)
print(response)async def rag_example():
async with await LLMCore.create() as llm:
# Add documents to vector store
await llm.add_documents_to_vector_store(
documents=[
{"content": "LLMCore supports multiple providers...", "metadata": {"source": "docs"}},
{"content": "Configuration uses the confy library...", "metadata": {"source": "docs"}},
],
collection_name="my_docs"
)
# Query with RAG
response = await llm.chat(
"How does LLMCore handle configuration?",
enable_rag=True,
rag_collection_name="my_docs",
rag_retrieval_k=3
)
print(response)from llmcore.agents import AgentManager, AgentMode
async def agent_example():
async with await LLMCore.create() as llm:
# Create agent manager
agent = AgentManager(
provider_manager=llm._provider_manager,
memory_manager=llm._memory_manager,
storage_manager=llm._storage_manager
)
# Run agent with a goal
result = await agent.run(
goal="Research the top 3 Python web frameworks and compare them",
mode=AgentMode.SINGLE
)
print(result.final_answer)llm.media routes image, audio and video work to whichever configured provider
can serve it. Adapters are the chat providers, so there is one credential per
vendor and no separate media config tree.
llm = await LLMCore.create(config)
# Routed to the first configured provider that can do it
image = await llm.media.images.generate("a calico cat asleep on books")
speech = await llm.media.audio.speak("Consent is not an afterthought.")
text = await llm.media.audio.transcribe(audio=MediaRef.from_path("call.mp3"))
# Ask what the current configuration can actually do
llm.media.capabilities() # every capability available
llm.media.who_can(MediaCapability.VIDEO_GENERATE) # ['gemini', 'fal', 'replicate']Image generation answers in one request on OpenAI and is a queued job on fal. So
a router call returns either a MediaResult or a pollable MediaJob, and
media.wait() absorbs the difference — the same code works against both.
job = await llm.media.video.generate("a drone shot over a fjord", provider="fal")
result = await llm.media.wait(job, timeout=600) # polls with capped backoff
print(result.artifacts[0].uri)Long jobs can also arrive by webhook: media.jobs issues signed, single-use
callback URLs and create_webhook_app() returns a plain ASGI app to receive
them. Polling remains the fallback, so no public ingress is required.
| Group | Capabilities |
|---|---|
| Image | image_generate, image_edit, image_upscale, image_variate |
| Video | video_generate, video_edit, video_interpolate, video_reframe, video_upscale, video_extend |
| Audio | tts, tts_stream, asr, asr_stream, music, sfx, voice_design, voice_agent |
| Document | ocr |
Artifacts carry bytes or a URI, checksums, dimensions/duration, usage, and
provenance. A content-addressed store can materialize remote artifacts before
vendor URLs expire. For synthetic speech, provenance includes a VoiceConsent
record: whose voice it is, whether it is a clone, and whether the vendor
considers it cleared — with None meaning the vendor did not say, which is
deliberately distinct from no.
llm.runtimes provisions compute elsewhere, serves an open-weights model on it,
and registers the endpoint as a provider instance — so a remotely served model
is reachable through the same llm.chat() as a hosted API.
plan = await llm.runtimes.estimate("Qwen/Qwen3-30B-A3B-Instruct-2507", context_length=32768)
print(plan.sku, plan.vram_required_gb, plan.fits) # free; nothing is provisioned
handle = await llm.runtimes.up(plan.spec.repo_id, name="qwen30", confirm_spend=True)
answer = await llm.chat("Explain GQA briefly.", provider_name="qwen30")
await llm.runtimes.down("qwen30") # unregister, then releaseThis subsystem bills per minute from the moment compute is assigned, which makes it unlike every other provider here. The safety rules are enforced in code, not left to the caller:
- off by default —
LLMCore.create()contacts no backend regardless of config; up()refuses without explicit spend confirmation, whileestimate()is free and works even while the subsystem is disabled;- state is written before provisioning returns, as inspectable JSON under
~/.llmcore/runtimes, so a runtime can always be found and killed; - bounded by default — idle and hard-lifetime deadlines, plus an optional compute-unit ceiling, because an idle reaper does not stop a runtime that is busy in a loop;
close()detaches but does not tear down — a process exiting is not a reason to destroy compute you are paying for.down_all()is explicit.
Phase R1 (core, protocols, state, provider attachment) ships with a
FakeRuntime for testing. The Colab backend is specified in
COLAB_RUNTIME_SPEC.md and not yet implemented,
so nothing here can spend money today.
Requires Python 3.11 or later.
pip install llmcoreEach provider is its own extra, so you install only what you use. All 35 extras
are listed in pyproject.toml.
# Chat and reasoning
pip install "llmcore[openai]" # also covers groq, together, xai, deepseek, kimi
pip install "llmcore[anthropic]"
pip install "llmcore[gemini]"
pip install "llmcore[mistral]" # httpx + mistralai v3 fallback
pip install "llmcore[zai]" # GLM family
pip install "llmcore[friendli]"
pip install "llmcore[deepinfra]"
pip install "llmcore[openrouter]"
pip install "llmcore[poe]"
pip install "llmcore[huggingface]"
pip install "llmcore[ollama]" # local, no API key
pip install "llmcore[vllm]" # self-hosted
# Media and voice
pip install "llmcore[elevenlabs]" # TTS, ASR, SFX, music, voice design
pip install "llmcore[deepgram]" # STT, TTS, Voice Agent
pip install "llmcore[fal]" # image, video, audio marketplace
pip install "llmcore[replicate]"
pip install "llmcore[higgsfield]" # image + video
# Typed judgment
pip install "llmcore[typesafe]"
# Web search
pip install "llmcore[brightdata]" "llmcore[serper]" "llmcore[serpapi]" "llmcore[semanticscholar]"# ChromaDB for vector storage
pip install llmcore[chromadb]
# PostgreSQL with pgvector
pip install llmcore[postgres]# Docker sandbox
pip install llmcore[sandbox-docker]
# VM/SSH sandbox
pip install llmcore[sandbox-vm]
# Both sandbox types
pip install llmcore[sandbox]# Everything included
pip install llmcore[all]git clone https://github.com/araray/llmcore.git
cd llmcore
pip install -e ".[dev]"┌───────────────────────────────────────────────────────────────────────────┐
│ LLMCore API Facade │
│ (llmcore.api.LLMCore) │
├───────────────────────────────────────────────────────────────────────────┤
│ │
│ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ ┌──────────────┐ │
│ │ Provider │ │ Session │ │ Memory │ │ Embedding │ │
│ │ Manager │ │ Manager │ │ Manager │ │ Manager │ │
│ └──────┬───────┘ └──────┬───────┘ └──────┬───────┘ └──────┬───────┘ │
│ │ │ │ │ │
│ ┌──────┴───────┐ ┌──────┴───────┐ ┌──────┴───────┐ ┌──────┴───────┐ │
│ │ Providers │ │ Storage │ │ RAG │ │ Embeddings │ │
│ │ 23 vendors │ │ • SQLite │ │ • ChromaDB │ │ • Sentence │ │
│ │ direct REST │ │ • Postgres │ │ • pgvector │ │ Transform │ │
│ │ + SDK │ │ • JSON │ │ • external │ │ • OpenAI │ │
│ │ fallbacks │ │ │ │ bridge │ │ • Google │ │
│ └──────────────┘ └──────────────┘ └──────────────┘ └──────────────┘ │
│ │
│ ┌──────────────┐ ┌──────────────┐ │
│ │ Media │ │ Runtimes │ optional subsystems, off until │
│ │ Manager │ │ Manager │ configured │
│ └──────┬───────┘ └──────┬───────┘ │
│ │ │ │
│ ┌──────┴───────┐ ┌──────┴───────┐ │
│ │ image/audio/ │ │ size → up → │ │
│ │ video router │ │ attach as a │ │
│ │ jobs+webhook │ │ provider │ │
│ │ artifacts │ │ reap/ceiling │ │
│ └──────────────┘ └──────────────┘ │
│ │
├───────────────────────────────────────────────────────────────────────────┤
│ Agent System (Darwin Layer 2) │
│ ┌──────────────────────────────────────────────────────────────────────┐ │
│ │ Cognitive Cycle │ │
│ │ PERCEIVE → PLAN → THINK → VALIDATE → ACT → OBSERVE → REFLECT → UPDATE│ │
│ └──────────────────────────────────────────────────────────────────────┘ │
│ │
│ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ │
│ │ Goal │ │ Fast-Path │ │ Circuit │ │ HITL │ │
│ │ Classifier │ │ Executor │ │ Breaker │ │ Manager │ │
│ └─────────────┘ └─────────────┘ └─────────────┘ └─────────────┘ │
│ │
│ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ │
│ │ Persona │ │ Prompt │ │ Activity │ │ Capability │ │
│ │ Manager │ │ Library │ │ System │ │ Checker │ │
│ └─────────────┘ └─────────────┘ └─────────────┘ └─────────────┘ │
│ │
├───────────────────────────────────────────────────────────────────────────┤
│ Sandbox System │
│ ┌──────────────────────────────────────────────────────────────────────┐ │
│ │ Docker Provider │ VM Provider │ Registry │ Output Tracker │ │
│ └──────────────────────────────────────────────────────────────────────┘ │
│ │
├───────────────────────────────────────────────────────────────────────────┤
│ Supporting Systems │
│ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐ │
│ │ Model Card │ │Observability│ │ Tracing │ │ Logging │ │
│ │ Registry │ │ System │ │ System │ │ Config │ │
│ └─────────────┘ └─────────────┘ └─────────────┘ └─────────────┘ │
└───────────────────────────────────────────────────────────────────────────┘
LLMCore uses confy for layered configuration with the following precedence (highest priority last):
- Package Defaults →
llmcore/config/default_config.toml - User Config →
~/.config/llmcore/config.toml - Custom File →
LLMCore.create(config_file_path="...") - Environment Variables →
LLMCORE_*prefix - Direct Overrides →
LLMCore.create(config_overrides={...})
# ~/.config/llmcore/config.toml
[llmcore]
default_provider = "openai"
default_embedding_model = "text-embedding-3-small"
log_level = "INFO"
[providers.openai]
# API key via: LLMCORE_PROVIDERS__OPENAI__API_KEY or OPENAI_API_KEY
default_model = "gpt-5.4"
timeout = 60
[providers.anthropic]
default_model = "claude-sonnet-5-5"
timeout = 60
[providers.ollama]
# host = "http://localhost:11434"
default_model = "llama3.2:latest"
[providers.vllm]
# Self-hosted vLLM server. base_url is required (no default).
# base_url = "http://localhost:8000/v1"
default_model = "meta-llama/Llama-3.1-8B-Instruct"
timeout = 240
[providers.zai]
# Z.ai Open Platform (GLM family). API key via ZAI_API_KEY.
# backend = "sdk" # "sdk" (native zai-sdk, default) | "openai" | "httpx"
# region = "overseas" # or "china" for the open.bigmodel.cn endpoint
default_model = "glm-5.2"
thinking = "enabled" # "enabled" | "disabled"
reasoning_effort = "high" # none|minimal|low|medium|high|xhigh|max
timeout = 300
[providers.friendli]
# FriendliAI. API key via FRIENDLI_TOKEN (FRIENDLIAI_API_KEY also accepted).
# endpoint_type = "serverless" # "serverless" | "dedicated" | "container"
# backend = "openai" # "openai" (default) | "httpx" | "sdk"
default_model = "zai-org/GLM-5.3"
parse_reasoning = true # split reasoning into reasoning_content
# reasoning_effort = "high" # minimal|low|medium|high|xhigh|max|ultracode
timeout = 300
[providers.typesafe]
# TypeSafe.ai System One (typed judgments, NOT chat). API key via TYPESAFE_API_KEY.
# Use provider.system_one(state, questions) or llm.chat(..., provider_name="typesafe", questions={...}).
default_model = "jev-latest" # alias -> jev-1.13.0; pin the version if you tune thresholds
timeout = 30
max_retries = 2 # 408/429/5xx/529 retried, honours Retry-After
[storage.session]
type = "sqlite"
path = "~/.llmcore/sessions.db"
[storage.vector]
type = "chromadb"
path = "~/.llmcore/chroma_db"
default_collection = "llmcore_default"
[agents]
max_iterations = 10
default_timeout = 600
[agents.sandbox]
mode = "docker"
[agents.sandbox.docker]
enabled = true
image = "python:3.11-slim"
memory_limit = "1g"
cpu_limit = 2.0
network_enabled = false
[agents.hitl]
enabled = true
global_risk_threshold = "medium"
default_timeout_seconds = 300# Provider API Keys
export LLMCORE_PROVIDERS__OPENAI__API_KEY="sk-..."
export LLMCORE_PROVIDERS__ANTHROPIC__API_KEY="sk-ant-..."
export LLMCORE_PROVIDERS__GEMINI__API_KEY="..."
# Or use standard provider env vars
export OPENAI_API_KEY="sk-..."
export ANTHROPIC_API_KEY="sk-ant-..."
# Storage
export LLMCORE_STORAGE__SESSION__TYPE="postgres"
export LLMCORE_STORAGE__SESSION__DB_URL="postgresql://user:pass@localhost/llmcore"
# Logging
export LLMCORE_LOG_LEVEL="DEBUG"
export LLMCORE_LOG_RAW_PAYLOADS="true"
export TYPESAFE_API_KEY="..." # TypeSafe.ai System One23 providers behind one interface. Transport shows which can call the vendor
API directly and which fall back to a vendor SDK — llmcore prefers direct calls
so an SDK lagging the API does not block you, and keeps the SDK where it owns
something non-trivial. Full detail in
PROVIDER_SUPPORT_MATRIX.md.
| Provider | Representative models (from the card registry) | Transport | Notable |
|---|---|---|---|
| OpenAI | gpt-6-sol, gpt-6-astra, gpt-5.6-terra, gpt-5.5-pro, gpt-5.4, o4-mini |
direct + SDK | Streaming, tools, vision, images, TTS/ASR, embeddings, native search |
| Anthropic | claude-opus-5-5, claude-opus-5, claude-sonnet-5-5, claude-opus-4-8, claude-haiku-4-5-20251001 |
SDK + direct | Streaming, tools, vision, extended thinking |
| Google Gemini | gemini-3.8-flash, gemini-3.7-flash, gemini-3.1-pro, gemini-2.5-pro |
SDK + direct | Tools, vision, Imagen, Veo video, native TTS, embeddings |
| xAI | grok-4.1-20251117, grok-4-heavy |
direct + SDK | Streaming, tools, Live Search |
| DeepSeek | deepseek-v4-pro, deepseek-v4-flash, deepseek-v3.2, deepseek-reasoner |
direct only¹ | Reasoning-content extraction, cache usage |
| Kimi (Moonshot) | kimi-k3, kimi-k2.7-code, kimi-k2.6, kimi-k2-thinking |
direct only¹ | Reasoning, exact tokenizer |
| Z.ai (GLM) | glm-5.3, glm-5.2, glm-5.1, glm-4.7, glm-5.3-flash |
SDK → direct | Tools, vision, image/video, TTS/ASR/OCR, embeddings, web search |
| Mistral | mistral-large-3, mistral-large-2512, magistral-medium-latest, magistral-small-latest |
direct + SDK | Tools, vision, FIM, OCR, embeddings, audio |
| Qwen | qwen3-max, qwen3-coder-480b |
direct | Streaming, tools |
| Groq | Llama, Qwen, Whisper, Kimi on LPU hardware | direct + SDK | Low-latency inference |
| Together | Open-weights catalog (Llama, Qwen, DeepSeek, FLUX) | direct + SDK | Chat, images, embeddings |
| FriendliAI | zai-org/GLM-5.3, deepseek-ai/DeepSeek-V3.2 + your dedicated endpoints |
direct → SDK | Reasoning effort/budget, regex-constrained output, exact tokenizer |
| DeepInfra | Qwen3-235B-A22B, DeepSeek, Llama, FLUX, Whisper, Kokoro |
direct + SDK | Chat, vision, TTS/ASR, images, embeddings |
| OpenRouter | 620 cards spanning most vendors | direct + SDK | One key, many vendors |
| Poe | 497 cards across Anthropic/OpenAI/Google and others | direct + SDK | Aggregated access |
| Hugging Face | 305 cards; any Inference Provider model | SDK + direct | Chat, image, TTS, ASR, private Inference Endpoints |
| Ollama | qwen3-vl:4b, llama3.3:70b, gemma3:12b, qwen3-embedding:8b |
SDK + direct | Local, no API key |
| vLLM | Anything you serve | direct + SDK | Self-hosted, guided grammars, structured output |
¹ No official vendor Python SDK exists — their own docs point at the openai
client, which llmcore already speaks. Verified against PyPI rather than assumed.
| Provider | What it serves | Transport |
|---|---|---|
| ElevenLabs | TTS (+streaming), ASR, SFX, music, voice design — with voice-consent metadata | direct + SDK |
| Deepgram | STT (Nova-3, Flux), TTS (Aura-2), Voice Agent, text intelligence | SDK + direct |
| fal | 9 capabilities: FLUX image gen/edit/upscale, video, FILM interpolation, SFX, music, TTS, ASR | direct + SDK |
| Replicate | One generic prediction adapter driven by each model's published schema | direct + SDK |
| Higgsfield | Soul image generation, plus hosted Kling and MiniMax Hailuo video | direct + SDK |
| TypeSafe.ai | jev-1.13.0 (alias jev-latest) — typed judgments: noul (yes/no probability), choice, score |
direct + SDK |
Model names move fast. The table shows what is in the bundled card registry at this release.
cardctl generate <provider>refreshes it from the live API, andllm.list_models()reports what your keys can actually reach — prefer that over anything written here.
# Use default provider
response = await llm.chat("Hello!")
# Override per-request
response = await llm.chat(
"Explain quantum computing",
provider_name="anthropic",
model_name="claude-sonnet-5-5"
)
# With provider-specific parameters
response = await llm.chat(
"Write a poem",
provider_name="openai",
model_name="gpt-5.4",
temperature=0.9,
max_tokens=500
)The Deepgram provider adds real-time voice/audio — speech-to-text (STT),
text-to-speech (TTS), conversational Flux STT, a bidirectional Voice
Agent (STT → LLM → TTS over one socket), and text intelligence. Deepgram
is not a chat-completion provider, so its media methods are called directly on
the provider instance (the LLMCore facade has no audio methods):
from llmcore.providers.deepgram_provider import DeepgramProvider
dg = DeepgramProvider({"api_key": "dg_...", "_instance_name": "deepgram"})
# Pre-recorded STT
result = await dg.transcribe_audio(open("call.wav", "rb").read(),
model="nova-3", smart_format=True, diarize=True)
print(result.text)
# TTS
speech = await dg.generate_speech("Hello from Aura.", voice="aura-2-thalia-en")
open("out.mp3", "wb").write(speech.audio_data)
# Live STT (async byte source -> streamed events)
async for ev in dg.transcribe_stream(mic_frames(), model="nova-3",
encoding="linear16", sample_rate=16000,
interim_results=True):
print(ev.is_final, ev.text)
# Voice Agent (auto-answers client-side tool calls; prompt is never defaulted)
async for event in dg.run_voice_agent(mic_audio(), function_handler=handle,
functions=[weather_fn],
prompt="You are a concise assistant."):
... # CONVERSATION_TEXT / AUDIO / FUNCTION_CALL_REQUEST events
await dg.close()See docs/Deepgram_provider_usage.md for
the full guide (config, streaming, Flux, Voice Agent settings shape, token auth)
and the runnable scripts in examples/ (deepgram_*.py).
The agent system implements an advanced cognitive architecture for autonomous task execution.
The 8-phase cognitive cycle enables sophisticated reasoning:
┌───────────────────────────────────────────────────────────────────┐
│ COGNITIVE CYCLE │
│ │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
│ │ PERCEIVE │ → │ PLAN │ → │ THINK │ → │ VALIDATE │ │
│ │ │ │ │ │ │ │ │ │
│ │ Analyze │ │ Generate │ │ Reason & │ │ Check │ │
│ │ goal & │ │ strategy │ │ decide │ │ validity │ │
│ │ context │ │ │ │ action │ │ │ │
│ └──────────┘ └──────────┘ └──────────┘ └──────────┘ │
│ ↑ │ │
│ │ ↓ │
│ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │
│ │ UPDATE │ ← │ REFLECT │ ← │ OBSERVE │ ← │ ACT │ │
│ │ │ │ │ │ │ │ │ │
│ │ Update │ │ Learn & │ │ Analyze │ │ Execute │ │
│ │ state & │ │ improve │ │ results │ │ action │ │
│ │ memory │ │ │ │ │ │ │ │
│ └──────────┘ └──────────┘ └──────────┘ └──────────┘ │
└───────────────────────────────────────────────────────────────────┘
Phase Descriptions:
| Phase | Purpose |
|---|---|
| PERCEIVE | Analyze goal, extract entities, assess complexity |
| PLAN | Generate execution strategy and select approach |
| THINK | Reason about next action using CoT/ReAct patterns |
| VALIDATE | Check proposed action validity and safety |
| ACT | Execute the chosen action (tool call, code, etc.) |
| OBSERVE | Analyze results and extract observations |
| REFLECT | Learn from outcome, identify improvements |
| UPDATE | Update working memory and iteration state |
Automatic goal complexity assessment for optimal routing:
from llmcore.agents import GoalClassifier, classify_goal
# Classify a goal
classification = classify_goal("What's 2 + 2?")
print(classification.complexity) # GoalComplexity.TRIVIAL
print(classification.execution_strategy) # ExecutionStrategy.FAST_PATH
# Complex goal
classification = classify_goal(
"Research and compare the top 5 cloud providers, "
"analyze their pricing, and create a recommendation report"
)
print(classification.complexity) # GoalComplexity.COMPLEX
print(classification.max_iterations) # 15Complexity Levels:
| Level | Max Iterations | Examples |
|---|---|---|
TRIVIAL |
1 | Greetings, simple math, factual Q&A |
SIMPLE |
5 | Single-step tasks, translations |
MODERATE |
10 | Multi-step tasks, analysis |
COMPLEX |
15 | Research, multi-source synthesis |
Bypass the full cognitive cycle for trivial goals (sub-5s responses):
from llmcore.agents.learning import FastPathExecutor, should_use_fast_path
# Check if fast-path is appropriate
if should_use_fast_path("Hello, how are you?"):
executor = FastPathExecutor(config=fast_path_config)
result = await executor.execute(goal="Hello!")
print(result.response) # Instant responseAutomatic detection and interruption of failing agent loops:
from llmcore.agents import AgentCircuitBreaker, CircuitBreakerConfig
breaker = AgentCircuitBreaker(CircuitBreakerConfig(
max_iterations=15,
max_same_errors=3,
max_execution_time_seconds=300,
max_total_cost=1.0,
progress_stall_threshold=5
))
# Circuit breaker trips on:
# - Maximum iterations exceeded
# - Repeated identical errors
# - Timeout exceeded
# - Cost limit exceeded
# - Progress stall detectedInteractive approval workflows for sensitive operations:
from llmcore.agents.hitl import HITLManager, HITLConfig, ConsoleHITLCallback
# Create HITL manager
hitl = HITLManager(
config=HITLConfig(
enabled=True,
global_risk_threshold="medium",
timeout_policy="reject"
),
callback=ConsoleHITLCallback() # Interactive console prompts
)
# Check if activity needs approval
decision = await hitl.check_approval(
activity_type="execute_shell",
parameters={"command": "rm -rf /tmp/test"}
)
if decision.is_approved:
# Execute the activity
pass
else:
print(f"Rejected: {decision.reason}")Risk Levels:
| Level | Requires Approval | Examples |
|---|---|---|
LOW |
No | Read files, calculations |
MEDIUM |
Configurable | Write files, API calls |
HIGH |
Yes | Delete files, network access |
CRITICAL |
Always | System commands, credentials |
Customize agent behavior and communication style:
from llmcore.agents import PersonaManager, AgentPersona, PersonalityTrait
# Create a custom persona
persona = AgentPersona(
name="DataAnalyst",
description="A meticulous data analyst focused on accuracy",
personality=[
PersonalityTrait.ANALYTICAL,
PersonalityTrait.METHODICAL,
PersonalityTrait.CAUTIOUS
],
communication_style="formal",
risk_tolerance="low",
planning_depth="thorough"
)
# Apply to agent
manager = PersonaManager()
manager.register_persona(persona)
agent_state.persona = manager.get_persona("DataAnalyst")Secure, isolated execution environments for agent-generated code.
┌────────────────────────────────────────────────────────────────────┐
│ Sandbox Registry │
│ (Manages sandbox lifecycle) │
├────────────────────────────────────────────────────────────────────┤
│ │
│ ┌─────────────────────────┐ ┌─────────────────────────┐ │
│ │ Docker Provider │ │ VM Provider │ │
│ │ │ │ │ │
│ │ • Container isolation │ │ • SSH-based access │ │
│ │ • Image management │ │ • Full VM isolation │ │
│ │ • Resource limits │ │ • Network separation │ │
│ │ • Output capture │ │ • Persistent storage │ │
│ └─────────────────────────┘ └─────────────────────────┘ │
│ │
│ ┌─────────────────────────┐ ┌─────────────────────────┐ │
│ │ Output Tracker │ │ Ephemeral Manager │ │
│ │ │ │ │ │
│ │ • File lineage │ │ • Resource cleanup │ │
│ │ • Execution logs │ │ • Timeout handling │ │
│ │ • Artifact collection │ │ • State management │ │
│ └─────────────────────────┘ └─────────────────────────┘ │
└────────────────────────────────────────────────────────────────────┘
from llmcore import (
SandboxRegistry, DockerSandboxProvider, SandboxConfig, SandboxMode
)
# Create sandbox registry
registry = SandboxRegistry()
# Register Docker provider
docker_provider = DockerSandboxProvider(SandboxConfig(
mode=SandboxMode.DOCKER,
image="python:3.11-slim",
memory_limit="1g",
cpu_limit=2.0,
timeout_seconds=300,
network_enabled=False
))
registry.register_provider("docker", docker_provider)
# Create and use sandbox
async with registry.create_sandbox("docker") as sandbox:
# Execute Python code
result = await sandbox.execute_python("""
import math
print(f"Pi = {math.pi}")
""")
print(result.stdout) # "Pi = 3.141592653589793"
# Execute shell command
result = await sandbox.execute_shell("ls -la")
print(result.stdout)
# Save file
await sandbox.save_file("output.txt", "Hello, Sandbox!")
# Read file
content = await sandbox.load_file("output.txt")Pre-built, security-hardened images organized by tier:
| Tier | Image | Description |
|---|---|---|
| Base | llmcore-sandbox-base:1.0.0 |
Minimal Ubuntu 24.04 |
| Specialized | llmcore-sandbox-python:1.0.0 |
Python 3.12 development |
llmcore-sandbox-nodejs:1.0.0 |
Node.js 22 development | |
llmcore-sandbox-shell:1.0.0 |
Shell scripting | |
| Task | llmcore-sandbox-research:1.0.0 |
Research & analysis |
llmcore-sandbox-websearch:1.0.0 |
Web scraping |
| Level | Network | Filesystem | Tools |
|---|---|---|---|
RESTRICTED |
Blocked | Limited | Whitelisted only |
FULL |
Enabled | Extended | All tools |
- Non-root execution: All containers run as
sandboxuser (UID 1000) - No SUID/SGID binaries: Privilege escalation vectors removed
- Resource limits: Memory, CPU, and process limits enforced
- Network isolation: Optional network blocking
- Output tracking: Full lineage and audit trail
- AppArmor/seccomp ready: Compatible with security profiles
2,319 cards across 22 providers, generated from live provider APIs and bundled with the package. Each carries context window, capability flags, pricing and lifecycle — so capability and cost questions are answered locally instead of by trial and error against a paid endpoint.
python -m tools.cardctl doctor # audit coverage — run this first
python -m tools.cardctl generate openai # refresh one provider from its API
python -m tools.cardctl validate # schema-check every card
python -m tools.cardctl diff anthropic # read-only: local cards vs live API
python -m tools.cardctl stats # coverage dashboarddoctor cross-checks llmcore's provider registry against cardctl's adapters and
the cards on disk. It exists because three media providers once shipped with no
adapter at all and nothing complained — generate only reports on the provider
you name, and stats only sees providers that already have cards.
Providers with no catalog endpoint (fal, Higgsfield, Replicate, vLLM) use a
curated adapter and every such card is tagged curated, so a declared entry
is never mistaken for a discovered one.
from llmcore import get_model_card_registry, get_model_card
# Get registry singleton
registry = get_model_card_registry()
# Lookup model card
card = registry.get("openai", "gpt-5.4")
print(f"Context: {card.get_context_length():,} tokens")
print(f"Vision: {card.capabilities.vision}")
print(f"Tools: {card.capabilities.tools}")
# Cost estimation
cost = card.estimate_cost(
input_tokens=50_000,
output_tokens=2_000,
cached_tokens=10_000
)
print(f"Estimated cost: ${cost:.4f}")
# List models by capability
vision_models = registry.list_cards(tags=["vision"])
for model in vision_models:
print(f"{model.provider}/{model.model_id}")
# Alias resolution
card = registry.get("anthropic", "claude-4.5-sonnet") # Resolves alias| Provider | Cards | Provider | Cards | |
|---|---|---|---|---|
| openrouter | 620 | mistral | 66 | |
| poe | 497 | anthropic | 19 | |
| huggingface | 305 | kimi | 16 | |
| deepinfra | 229 | zai | 13 | |
| ollama | 176 | elevenlabs | 11 | |
| deepgram | 147 | fal | 9 | |
| openai | 133 | deepseek / friendli / replicate | 7 each | |
| 44 | higgsfield | 6 | ||
| qwen · xai · typesafe | 3 · 3 · 1 |
- xAI: Grok-4, Grok-4-Heavy
- DeepInfra: DeepSeek-V3/R1, Llama 3.x, Qwen, Mistral, FLUX (image), Whisper (STT), Kokoro (TTS), embeddings
- Deepgram: Nova-3, Nova-2, Whisper, Flux (STT); Aura-2 (TTS); Voice Agent
- TypeSafe.ai: Jev 1.13 (
decisionmodel type; aliasesjev-latest,jev-preview)
Add custom cards in ~/.config/llmcore/model_cards/<provider>/<model>.json:
{
"model_id": "my-custom-model",
"display_name": "My Custom Model",
"provider": "ollama",
"model_type": "chat",
"context": {
"max_input_tokens": 32768,
"max_output_tokens": 4096
},
"capabilities": {
"streaming": true,
"tools": false,
"vision": false
}
}Comprehensive monitoring and debugging for agent executions.
from llmcore.agents.observability import EventLogger, EventCategory
# Events logged to ~/.llmcore/events.jsonl
logger = EventLogger(log_path="~/.llmcore/events.jsonl")
# Event categories
# - LIFECYCLE: Agent start/stop
# - COGNITIVE: Phase execution
# - ACTIVITY: Tool executions
# - HITL: Human approvals
# - ERROR: Exceptions
# - METRIC: Performance data
# - MEMORY: Memory operations
# - SANDBOX: Container lifecycle
# - RAG: Retrieval operationsfrom llmcore.agents.observability import MetricsCollector
collector = MetricsCollector()
# Available metrics
# - Iteration counts
# - LLM call latency (p50, p90, p95, p99)
# - Token usage (input/output)
# - Estimated costs
# - Activity execution times
# - Error counts by typefrom llmcore.agents.observability import ExecutionReplay
replay = ExecutionReplay(events_path="~/.llmcore/events.jsonl")
# List executions
executions = replay.list_executions()
for exec_id, metadata in executions.items():
print(f"{exec_id}: {metadata['goal']}")
# Replay specific execution
events = replay.get_execution_events(exec_id)
for event in events:
print(f"[{event.timestamp}] {event.category}: {event.type}")[agents.observability]
enabled = true
[agents.observability.events]
enabled = true
log_path = "~/.llmcore/events.jsonl"
min_severity = "info"
categories = [] # Empty = all categories
[agents.observability.events.rotation]
strategy = "size"
max_size_mb = 100
max_files = 10
compress = true
[agents.observability.metrics]
enabled = true
track_cost = true
track_tokens = true
latency_percentiles = [50, 90, 95, 99]
[agents.observability.replay]
enabled = true
cache_enabled = true
cache_max_executions = 50Retrieval-Augmented Generation for knowledge-enhanced responses.
# Add documents with metadata
await llm.add_documents_to_vector_store(
documents=[
{
"content": "LLMCore is a Python library...",
"metadata": {
"source": "documentation",
"version": "0.53.0",
"category": "overview"
}
},
{
"content": "Configuration uses the confy library...",
"metadata": {
"source": "documentation",
"category": "configuration"
}
}
],
collection_name="project_docs"
)# Direct similarity search
results = await llm.search_vector_store(
query="How does configuration work?",
k=5,
collection_name="project_docs",
metadata_filter={"category": "configuration"}
)
for doc in results:
print(f"Score: {doc.score:.4f}")
print(f"Content: {doc.content[:100]}...")
print(f"Metadata: {doc.metadata}")response = await llm.chat(
"Explain how to configure providers",
enable_rag=True,
rag_collection_name="project_docs",
rag_retrieval_k=3,
system_message="Answer based ONLY on the provided context."
)LLMCore can serve as an LLM backend for external RAG engines:
# External engine (e.g., semantiscan) handles retrieval
relevant_docs = await external_rag_engine.retrieve(query)
# Construct prompt with retrieved context
context = format_documents(relevant_docs)
full_prompt = f"Context:\n{context}\n\nQuestion: {query}"
# Use LLMCore for generation only
response = await llm.chat(
message=full_prompt,
enable_rag=False, # Disable internal RAG
explicitly_staged_items=[] # Optional additional context
)| Class | Description |
|---|---|
LLMCore |
Main facade for all LLM operations |
AgentManager |
Manages autonomous agent execution |
StorageManager |
Handles session and vector storage |
ProviderManager |
Manages LLM provider connections |
| Model | Description |
|---|---|
ChatSession |
Conversation session with messages |
Message |
Individual chat message (user/assistant/system) |
ContextDocument |
Document for RAG/context |
Tool |
Function/tool definition |
ToolCall |
Tool invocation by LLM |
ToolResult |
Result of tool execution |
ModelCard |
Model metadata and capabilities |
from llmcore import (
LLMCoreError, # Base exception
ConfigError, # Configuration issues
ProviderError, # LLM provider errors
StorageError, # Storage operations
SessionStorageError, # Session storage specific
VectorStorageError, # Vector storage specific
SessionNotFoundError, # Session lookup failure
ContextError, # Context management
ContextLengthError, # Context exceeds limits
EmbeddingError, # Embedding generation
SandboxError, # Sandbox execution
SandboxInitializationError,
SandboxExecutionError,
SandboxTimeoutError,
SandboxAccessDenied,
SandboxResourceError,
SandboxConnectionError,
SandboxCleanupError,
)Reference
| Document | What it covers |
|---|---|
| Configuration reference | Every config key |
| Model cards | Card schema, the cardctl workflow, doctor |
chat_with_usage guide |
Per-call token and cost accounting |
| Agentic system guide | Cognitive cycle, tools, HITL |
| External RAG integration | Bringing your own retrieval |
Design and status
| Document | What it covers |
|---|---|
| Provider support matrix | Per provider: tracked SDK version and commit, transport duality, capability matrix, and a live-validation log recording what was actually called |
| Provider modernization plan | Phased plan for the remaining gaps in that matrix |
| Media subsystem spec | Design, the rollout, and what each vendor taught the abstraction |
| Remote runtime spec | Runtime safety model and the Colab backend design |
Per-provider guides
fal · ElevenLabs · Replicate · Hugging Face media · Deepgram · FriendliAI · TypeSafe.ai · Search providers (rationale)
llmcore/
├── src/llmcore/
│ ├── __init__.py # Public API exports
│ ├── api.py # Main LLMCore class
│ ├── models.py # Core data models
│ ├── exceptions.py # Exception hierarchy
│ ├── config/ # Configuration system
│ │ ├── default_config.toml
│ │ └── models.py
│ ├── providers/ # 23 provider adapters
│ │ ├── openai_provider.py # base for deepinfra/vllm/poe/openrouter
│ │ ├── anthropic_provider.py
│ │ ├── gemini_provider.py
│ │ ├── fal_provider.py # + elevenlabs, replicate, higgsfield,
│ │ └── ... # deepgram, zai, friendli, mistral, …
│ ├── media/ # Generative media subsystem
│ │ ├── manager.py # routing + capability discovery
│ │ ├── protocols.py # 19 capability protocols
│ │ ├── jobs.py # async job polling
│ │ ├── webhooks.py # signed single-use callbacks
│ │ └── artifacts.py # content-addressed store
│ ├── runtimes/ # Remote GPU runtimes
│ │ ├── manager.py # spend ceilings, attach/detach
│ │ ├── protocols.py # ComputeRuntime
│ │ └── state.py # inspectable on-disk records
│ ├── storage/ # Storage backends
│ │ ├── sqlite_session.py
│ │ ├── postgres_session_storage.py
│ │ ├── chromadb_vector.py
│ │ └── pgvector_storage.py
│ ├── embedding/ # Embedding models
│ │ ├── sentence_transformer.py
│ │ ├── openai.py
│ │ └── google.py
│ ├── agents/ # Agent system
│ │ ├── manager.py # AgentManager
│ │ ├── cognitive/ # 8-phase cognitive cycle
│ │ ├── hitl/ # Human-in-the-loop
│ │ ├── sandbox/ # Sandbox execution
│ │ ├── learning/ # Learning mechanisms
│ │ ├── persona/ # Persona system
│ │ ├── prompts/ # Prompt library
│ │ ├── observability/ # Monitoring & logging
│ │ └── routing/ # Model routing
│ ├── model_cards/ # Model metadata
│ │ ├── registry.py
│ │ ├── schema.py
│ │ └── default_cards/
│ └── memory/ # Memory management
├── tools/cardctl/ # Model-card generator (adapters + doctor)
├── container_images/ # Sandbox Docker images
├── examples/ # Usage examples
├── tests/ # Test suite
├── docs/ # Documentation
└── pyproject.toml # Project configuration
# The profile CI runs: no databases, containers, local servers or API keys
pytest tests --ignore=tests/integration --ignore=tests/adhoc_checks \
-m "not slow and not integration and not docker and not vm and not sandbox \
and not requires_postgres and not requires_pgvector and not requires_ollama"
# Everything (needs PostgreSQL + pgvector; Ollama for the local-model tests)
pytest
# Targeted
pytest tests/providers # provider adapters and transports
pytest tests/media # media subsystem
pytest tests/runtimes # runtime safety model
pytest -m sandbox # sandbox tests onlyInfrastructure-dependent tests are marked and excluded by default, so a clean checkout runs green without Postgres, Docker or a GPU. Tests that need those report why they skipped rather than silently passing.
A few suites exist to make whole-repository mistakes fail loudly rather than ship quietly:
| Guard | What it enforces |
|---|---|
tests/providers/test_transport_duality.py |
Every provider either offers both transports or declares why not. A declared SDK backend must actually be constructed and called — written after one shipped that silently did nothing |
tests/tools/test_cardctl_coverage.py |
Every registered provider has a cardctl adapter, so a new provider cannot ship without model cards |
tests/media/test_media_subsystem.py |
Every declared media capability is backed by its protocol |
tests/providers/test_context_length_error_mapping.py |
Static AST check that providers raise ContextLengthError with the right keywords |
Contributions are welcome! Please follow these guidelines:
- Fork the repository
- Create a feature branch (
git checkout -b feature/amazing-feature) - Write tests for your changes
- Ensure all tests pass (
pytest) - Follow the existing code style (
ruff check .) - Commit with conventional commits (
feat: add amazing feature) - Push to your branch
- Open a Pull Request
git clone https://github.com/araray/llmcore.git
cd llmcore
pip install -e ".[dev]"
pre-commit install- Formatter: Ruff
- Type Hints: Required for all public APIs
- Docstrings: Google style
- Line Length: 100 characters
This project is licensed under the MIT License. See the LICENSE file for details.
- llmchat - CLI interface for llmcore
- semantiscan - Advanced RAG engine
- confy - Configuration management library
- Documentation: docs/
- Issues: GitHub Issues
- Discussions: GitHub Discussions
Built with ❤️ by Araray Velho