Spanish (es) is the primary documentation language; this README is in English for GitHub.
Swarmind is a Python multi-agent orchestration harness for coordinating AI coding agents (opencode, Claude Code, Codex) through a single, modular engine. It is 100% TDD, adversarial, evolutionary, and built to the frontier of 2026 practices: a fan-out orchestrator with governed voting, heuristic model routing with fallback, real PBT and mutation validation oracles, central portable memory with vector search, token economics, and CUDA GPU acceleration.
- Quickstart
- Features
- Architecture
- Installation
- Usage
- Memory & Backup
- Token Economics
- GPU Acceleration
- Development & Quality Gates
- Project Structure
- Documentation
- License
Clone, set up, and verify the harness in a few steps:
# 1. Clone
git clone https://github.com/MauricioFCC/SWARMIND.git
cd SWARMIND
# 2. One-command cross-platform setup (verifies Python 3.12+, installs uv, uv sync,
# symlinks config to opencode global, installs hooks, verifies import)
./scripts/install.sh # Linux / macOS
# .\scripts\install.ps1 # Windows (PowerShell, symlink fallback to copy)
# 3. Run the test suite
uv run python -m pytest harness/tests/ -q
# 4. Optional: enable CUDA GPU acceleration (reinstalls the torch CUDA wheel)
python scripts/enable_gpu.py
# 5. Launch the interactive menu
launcher.bat # Windows
# or: python -m harnessAlternative:
python scripts/setup_swarmind.pyperforms the same auto-setup (Windows-oriented, uses copy instead of symlinks). Useinstall.sh/install.ps1for idempotent, symlink-based setup on any OS.
- ParallelExecutors with native fan-out (
harness/orchestrator/parallel_executor.py,ThreadPoolExecutor,max_workers=3) and governed voting (gate score β₯ 70 and confidence < 0.7, N=3). Inspired by the 2026 ORCA analysis: stablyai/orca was evaluated and discarded as a tool; its parallelism/voting was adopted natively instead. - Task planning and orchestration, agent bus, MARS scheduler, MetaClaw, adaptive planning, debate orchestration, worktable, and workflows.
harness/run_commands/package for interactive commands (!rag,!db,!iteration) and a multi-harness layer with adapters + CLI.
The difference isn't the model. It's the harness. An agent without a harness is an isolated department: duplicated effort, no shared memory, no scaling, no measurement. Every new tool/MCP/model is adopted as an orchestrated process (fan-out, governed voting β₯70, SSOT memory, PBT/mutation oracles) or discarded β see the ORCA 2026 case (stablyai/orca dropped as a tool, its parallel/voting process adopted natively).
- Heuristic
ModelRouter(harness/model_router/complexity_router/) with small/frontier signals androute_with_fallback(confidence < 0.7 falls back to the frontier model). MultiAPIProviderwith failover and health-checks (harness/model_router/multi_provider/,harness/model_router/provider_health/).- Token budget SSOT (
token_budgets.yaml), cache-shape (-38%), structured compaction (-41%), governed voting, and budget enforcement viaTokenBudgetManager.
- The harness can delegate tasks to local models through Ollama (
harness/model_router/ollama_client.py+harness/model_router/ollama_tiers.py), with four capability tiers plus a coding tier, all running current 2026 models installed locally: fast (qwen3:4b), quality (deepseek-r1:8b), coding (qwen2.5-coder:7b), embedding/RAG (qwen3-embedding:0.6b) and vision (qwen3-vl:4b) β fully configurable (no hardcode) in.opencode/config/ollama_models.yaml(base_url, timeout, warm_on_start, per-tier keep_alive/auto_pull). - Hot models are kept resident with
keep_alive: "5m"(warm/unload via/api/ps) and are auto-installed withollama pullwhen missing (auto_pull: true), so simple tasks run fully local: 0 cloud tokens (TKN). - If Ollama is unavailable or a tier's model is missing, the router degrades to the existing cloud
ModelRouter/SlmRouterfallback.
- LLM-grep for code (
harness/memory_rag/llm_grep.py): ripgrep-first 3-layer search (lexicalrgβ structuralast-grepβ semanticHybridRetrieverlast resort), compaction-friendly output (path:line+ 2 context lines, dedup, byte budget), auditable routing report (semantic_ratiomisrouting alert) β ADR-0067. - Post-compaction re-anchor (
harness/memory_rag/reanchor.py): condensed<<RE-ANCHOR>>block (N1 + active role + skills + task) re-injected after every compaction; summaries retain ~17% of session constraints, the block restores >90% (65% of enterprise agent failures are context drift, not token exhaustion) β ADR-0070. - Cascade routing (STEER-lite) (
harness/model_router/cascade_router.py): try small first, escalate to frontier when confidence < 0.7,force_tierescape hatch, per-attempt cost accounting for offline threshold calibration β ADR-0068. - Session-affinity routing (SAAR) (
harness/model_router/session_affinity.py): sticky model tier per session with TTL; avoids repeated model switches (frontier: β79% switches, β78.7% cost) and keeps prefix caches warm β ADR-0073. - Batch voting k-in-1 (
harness/orchestrator/batch_vote.py): k votes in one API call via thenparameter (input charged once instead of kΓ), automatic fallback to k sequential calls, quorum-gated majority β ADR-0073 (arXiv 2604.13717). - Structured-output enforcer (
harness/orchestrator/structured_enforcer.py): JSON-schema validation with error-feedback retries (99.9% schema adherence vs <70% unconstrained; 30Γ fewer parse failures) β ADR-0073. - Cache health diagnostics (
TokenUsageTracker.cache_health): flags structural cache-busters (hit ratio < 60% with β₯10K volume β timestamps in system, reordered few-shots, dynamic tool lists) β ADR-0068. - Context optimization (ADR-0074):
artifact_store.py(tool results >4K chars β disk + handle/offset, access preserved),cue_ledger.py(cue-anchored index with injection dedup + staleness, arXiv 2607.20972: β42% tokens),compaction_calibration.py(AgeMem warn 0.75/critical 0.90 zones),model_efficiency_report()(tokens/call per model). - Skills/agents frontier (ADR-0075 + wiring):
skill_composition.py(lazycalls:+ cycle detection + invocation tiers + conflict pruning, Pocock β63%),competence_model.py(Beta posterior per agentΓskill + Thompson anti-collapse + imp@k),fanout_gate.py(anti-over-decomposition: baseline β₯80% β single, avoids Γ17.2 noise). Wired: composition pilot, competence re-rank inAgentSelector, baseline gate invote_on_task+adaptive_planner. - Distributed-systems tooling (ADR-0076):
tool_output_filter.py(rtk wrapper: up to 90% less bash output, opt-in passthrough),idempotency_guard.py(effect dedup by key+payload hash, replay cache),structured_enforcerstrict keys (rejects trojan keys),llm_grep.TgrepBackend(Microsoft tgrep, opt-in); principles mandate Python/bash scripts (PowerShell corrupts UTF-8) + Linux-first tools. - Verify-replan + trace replay (ADR-0079):
verify_replan_gate.py(VMAO stop thresholds),trace_viewer.py(export trace.jsonl + deterministic replay without LLM), per-agent permissions inopencode.json,agent-rigorskill (PEC-35: pre-merge gates, anti-greenwashing). - Competition harness (ADR-0080):
cp_spec_gate.py(4-pillar pre-code gate: edges/invariants/complexity/io_constraints; attacks the 44% design+boundary at the gate) +dual_verify.py(fast vs brute-force with indexed mismatches). - Local execution real (ADR-0078):
LocalExecutorcloses the loop (closed-task allowlist + cloud fallback; trivial = 0 cloud tokens),pressure()meter,prune_then_summarizepipeline, R1 reasoning discipline in principles. - PEC universal in skills: all 35 skills carry an expert persona + canonical frontier references per specialty (OWASP for security, HL7 FHIR for healthtech, Rust API Guidelines, RICOUI Brands for UI...) + anti-hedging rule (
scripts/apply_pec.py, 176 tests) β ADR-0072. - Universal principles v3.1.0: 36 numbered principles with an adherence taxonomy (CHECK vs GUIDE, IFEval/DRFR), post-compaction re-pin rule (RPA), competition-programming fundamentals (CPD) and adversarial TDD/mutants/PBT/pairwise/BVA (TST/PBT) β ADR-0070/0077.
- Ollama CODING tier: local
qwen2.5-coder:7btier with precedence over QUALITY for code tasks + frontier-only keyword filter (is_frontier_only) β ADR-0069. - opencode runs local by default:
"model": "ollama/qwen3:4b"in.opencode/opencode.jsonwith 6 registered local models (fast/quality/coding/instruct/vision/ultra-fast).
- anydoc β document ingestion for RAG (
harness/memory_rag/doc_converter.py): aDocumentConverterprotocol plusAnyDocConverter(lazy, backed byfirecrawl-anydoc>=0.1.9) converts 21 binary/text extensions (pdf, docx, doc, pptx, ppt, xlsx, xls, odt, odp, ods, rtf, epub, csv, tsv, html, htm, md, txt, json, yaml, yml) to Markdown before chunking.DocumentChunkeraccepts an injected converter (DI, defaultAnyDocConverter) and raisesDocumentConversionError(path, reason)when conversion fails β errors are never swallowed. Enable it withharness/scripts/rag_ingest.py --include-docsor the interactive!rag ingest --docs. - deepseek-harness patterns β plugin lifecycle + session replay:
harness/plugins/registry.pyextendsPluginBasewithon_load()/on_unload()/events(no-op defaults),ToolRegistryaccepts an optionalevent_bus(DI) and auto-subscribes plugins toon_{event}handlers;load_all()/unload_all()are idempotent.harness/observability/session_replay.pyprovidesSessionReplayto replay recorded sessions (Markdown/JSON export) withSessionNotFoundError. Demo:harness/plugins/tools/example_tool.py(GreeterTool).
- Central portable memory with LanceDB vector store (
harness/memory_rag/lance_vector_store.py), semantic cache, SQLite-vec adapter (edge/offline backend), federated search, context window management, andshapley_flowoptimization. - Hybrid RAG (RRF) (
harness/memory_rag/hybrid_retriever.py):HybridRetrieverfuses dense vector (LanceDB embeddings) and sparse BM25 (SQLite FTS5) rankings with Reciprocal Rank Fusion (k=60) β documents present in both rankings rank higher; DI overFTSSearch+LanceVectorStore. - Corrective RAG (CRAG) (
harness/memory_rag/corrective_retriever.py):CorrectiveRetrievervalidates retrieval quality before generation (arXiv:2401.15884) β if poor, applies query rewrite or falls back to an alternative source, reporting thecorrective_actiontaken (none/rewrite/fallback). AgentKPITracker, compression strategies, context assembler, token budget managers, and skill loader.
harness/validation/pbt_stage.py: a real Hypothesis oracle running in a subprocess with 5 invariants (no_crash, returns_value, output_list, deterministic, commutative).harness/validation/mutation_stage.py: AST mutation with isolated subprocess execution.
- CUDA 12.6, RTX 4060 8GB, torch 2.13.0+cu126.
harness/gpu_accel.py+harness/gpu_optimize.pywith measured speedups: vector search x10.9 (10k), x9.2 (100k), and embeddings at 41Β΅s/msg.scripts/enable_gpu.pyreinstalls the torch CUDA wheel after everyuv sync.
- One-command setup (
scripts/setup_swarmind.py) and config menu (config_swarmind.py). - CPU-only PyPI torch wheel replaced by the CUDA wheel through a single script, portable across Linux/Mac/Windows.
- Full test suite, ruff-clean code, zero dead code (vulture), and zero architecture debt.
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β AGENTS LAYER β
β opencode Β· Claude Code Β· Codex (multi-harness) β
βββββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββββββ
β
βββββββββββββββββββββββββββββΌββββββββββββββββββββββββββββββββββ
β ORCHESTRATION LAYER β
β harness/orchestrator/ β
β parallel_executor (fan-out + voting) Β· task_orchestrator β
β task_planner Β· agent_bus Β· mars_scheduler Β· metaclaw β
β debate_orchestrator Β· workflows Β· worktable β
β run_commands/ (commands) Β· multi_harness (adapters+cli) β
βββββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββββββ
β
βββββββββββββββββββββββββββββΌββββββββββββββββββββββββββββββββββ
β MODEL ROUTING LAYER β
β harness/model_router/ β
β complexity_router (small/frontier) Β· route_with_fallback β
β multi_provider (failover + health-checks) Β· provider_healthβ
βββββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββββββ
β
βββββββββββββββββββββββββββββΌββββββββββββββββββββββββββββββββββ
β MEMORY & RAG LAYER β
β harness/memory_rag/ β
β lance_vector_store Β· semantic_cache Β· sqlite_vec_adapter β
β federated_search Β· shapley_flow Β· context_window_manager β
β doc_converter Β· doc_ingester (binaries β Markdown β RAG) β
β hybrid_retriever (RRF dense+sparse) Β· corrective_retriever β
β token_budget Β· token_budget_manager β
βββββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββββββ
β
βββββββββββββββββββββββββββββΌββββββββββββββββββββββββββββββββββ
β VALIDATION LAYER β
β harness/validation/ β
β pbt_stage.py (Hypothesis oracle, subprocess, 5 invariants) β
β mutation_stage.py (AST mutation, isolated subprocess) β
βββββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββββββ
β
βββββββββββββββββββββββββββββΌββββββββββββββββββββββββββββββββββ
β ACCELERATION LAYER β
β harness/gpu_accel.py Β· harness/gpu_optimize.py β
β CUDA 12.6 Β· torch 2.13.0+cu126 Β· RTX 4060 8GB β
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Each layer is a dedicated package under harness/: the orchestration layer coordinates agents with parallel execution and governed voting; the routing layer dispatches requests to small or frontier models with automatic fallback and provider failover; the memory layer centralizes knowledge in a portable vector store; the validation layer enforces quality with real property-based and mutation oracles; and the acceleration layer offloads compute to the GPU where available.
- Python 3.12+
- uv as package/dependency manager
- Optional: NVIDIA GPU with CUDA 12.6 (e.g. RTX 4060 8GB)
One command, any OS (idempotent, symlink-based, fallback to copy on Windows without Developer Mode):
git clone https://github.com/MauricioFCC/SWARMIND.git
cd SWARMIND
./scripts/install.sh # Linux / macOS
# .\scripts\install.ps1 # Windows (PowerShell)The installer verifies Python 3.12+, installs uv
if missing, runs uv sync, creates symlinks from ~/.config/opencode/ to the repo
(config, agents, skills), installs the pre-commit hooks (git config core.hooksPath .githooks), and verifies the harness imports. Re-running it is safe.
Alternative (Windows-oriented, copy-based):
python scripts/setup_swarmind.pysetup_swarmind.py verifies Python 3.12+, installs uv if missing, runs uv sync,
syncs the configuration to the opencode global directory, and creates the central
memory store. An interactive config menu is available via config_swarmind.py.
The PyPI torch wheel resolved by the lockfile is CPU-only on Windows/Linux. After setting up, activate CUDA with:
python scripts/enable_gpu.pyThis reinstalls the torch CUDA wheel and verifies it. Important: do not run uv sync after enabling CUDA, otherwise the CPU wheel is restored and the CUDA binary is lost; re-run python scripts/enable_gpu.py after any uv sync.
The harness is portable across Linux, macOS, and Windows. Path resolution uses a resilient _safe_home() that does not depend on HOME, and the central memory store plus its backups live under a configurable root.
Launch the interactive harness and delegate tasks to agents:
python -m harnessAvailable entry points include:
delegateβ assign a task to a specific coding agent (opencode, Claude Code, Codex).runβ execute a task through the orchestration pipeline.schedulerβ schedule recurring agent runs (harness/scheduler/).- Iteration pipeline β run a full iteration loop: plan, parallel fan-out, governed voting, validation, and memory persistence.
Interactive commands (harness/run_commands/):
!ragβ query the vector store.!dbβ database inspection / maintenance.!iterationβ trigger the iteration pipeline on demand.
Central memory is the single source of truth (SSOT):
- Default root:
~/Documents/Memory_Proyects(override withMEMORY_ROOT). - Contains
data/lancedbplus backups.
Resolution priority in memory_config.py:
LANCEDB_PATHenvironment variable.swarmind_config.json(MEMORY_ROOT)- Legacy
harness/db/lancedb
Backups are managed with scripts/backup_memory.py:
python scripts/backup_memory.py --list # list existing backups
python scripts/backup_memory.py --schedule # register a scheduled backupDuplicated databases were eliminated (7.5 GB reclaimed). The memory layout is portable across Linux, macOS, and Windows via a resilient _safe_home() that works without HOME.
- Model routing:
harness/model_router/complexity_router/emits small/frontier signals;route_with_fallbacksends low-confidence requests (confidence < 0.7) to the frontier model. - Cascade routing:
cascade_routertries small first and escalates on low confidence (STEER-lite), with per-attempt cost accounting;session_affinitykeeps the tier sticky per session to avoid prefill re-payments. - Governed voting:
ParallelExecutorfans out to N=3 agents when the gate score is β₯ 70 and confidence < 0.7, with a budget ofMAX_TOKENS_BY_AGENT Γ 3;batch_votecharges input once for k votes (thenparameter) with fallback. - Structured outputs:
structured_enforcervalidates every machine-readable verdict against a JSON schema with error-feedback retries (30Γ fewer parse failures). - Token budgets: budgets are the single source of truth (
token_budgets.yaml), enforced byTokenBudgetManager;cache_healthflags structural cache-busters (hit < 60% with volume). - Measured savings: cache-shape -38%, structured compaction -41%.
The harness auto-detects the GPU via harness/gpu_accel.py and optimizes tensor operations with harness/gpu_optimize.py.
- Environment: CUDA 12.6, RTX 4060 8GB, torch 2.13.0+cu126.
- Measured speedups:
- Vector search x10.9 (10k vectors), x9.2 (100k vectors)
- Embeddings 41Β΅s/msg
Enable it with:
python scripts/enable_gpu.pyWarning: never run uv sync after enabling CUDA β the lockfile resolves the CPU torch wheel from PyPI and the CUDA binary is lost. Re-run python scripts/enable_gpu.py after every uv sync.
Quality is enforced continuously, not at the end:
- Test suite: 5383 tests collected (TDD suite), mutation testing mutmut gate β₯70%.
- Lint: ruff β all checks passed.
- Dead code: vulture β 0 dead code.
- Architecture debt (AGR): 0 files over 500 lines in non-test code; 32 flat modules refactored into packages with re-exporting
__init__.py; mixins limited to β€ 2 bases; SOLID corrected in 9 classes. - Validation oracles: real Hypothesis property-based tests (
pbt_stage.py, 5 invariants) and AST mutation testing (mutation_stage.py) run in isolated subprocesses. - TDD: strictly always-on (RED β GREEN β REFACTOR), backed by WAL before expensive runs.
SWARMIND/
βββ harness/ # Core engine (Python 3.12+, packages per domain)
β βββ orchestrator/ # agent_bus, task_planner, task_orchestrator,
β β # mars_scheduler, metaclaw, adaptive_planner,
β β # natural_language_tools, tool_guardian, hitl,
β β # multi_user_governance, organizational_layer,
β β # health, federated_memory, agent_discovery,
β β # debate_orchestrator, worktable, workflows,
β β # multi_harness (adapters+cli), parallel_executor.py
β βββ model_router/ # complexity_router, multi_provider,
β β # provider_health, ollama_client, ollama_tiers
β βββ memory_rag/ # lance_vector_store, semantic_cache,
β β # sqlite_vec_adapter, federated_search,
β β # agent_kpi_tracker, vector_store_adapter,
β β # context_window_manager, compression_strategies,
β β # shapley_flow, optimization_pipeline,
β β # context_assembler, token_budget,
β β # token_budget_manager, skill_loader,
β β # doc_converter, doc_ingester (anydoc),
β β # hybrid_retriever (RRF), corrective_retriever (CRAG)
β βββ evolve_loop/ # agent_builder, skill_generator, prompt_evolver,
β β # gepa_mutator, nudge_system, evaluator,
β β # self_improver, procedural_memory, cognition_sync
β βββ validation/ # pbt_stage.py, mutation_stage.py (oracles reales)
β βββ security/ # zero_trust.py
β βββ hooks/ # hook_manager, hook_registry, builtin_hooks
β βββ observability/ # OpenTelemetry, logging, session_replay
β βββ qa/ # detector, generator, orchestrator, predictor
β βββ gateway/ # Slack/Telegram/CLI gateways
β βββ parallel/ # adaptive_pool, io_fusion, pipeline_macu
β βββ benchmarks/ # bench_memory, bench_routing, bench_cache, ...
β βββ plugins/ # plugin lifecycle (on_load/on_unload/events) + tools
β βββ run_commands/ # interactive commands (!rag, !db, !iteration)
β βββ scheduler/ # scheduled runs (Simple/Lance schedulers)
β βββ db/ # migrate_engine, iteration_reports
β βββ tools_sandbox/ # MCP client, mcp_executor, mcp_manager
β βββ aifactory/ # factory, agent_factory
β βββ guardrails/ # guardrail_engine
β βββ evals/ # eval_factory
β βββ context/ # token_budget_router, skill_contract
β βββ scripts/ # init, rag_ingest, end_of_iteration, ...
β βββ tests/ # 224 test files (5383 tests)
βββ scripts/ # Repo-level tooling
β βββ setup_swarmind.py # one-command setup (Python 3.12+, uv, uv sync, sync global, central memory)
β βββ enable_gpu.py # reinstall torch CUDA wheel after uv sync
β βββ backup_memory.py # central memory backups (--list / --schedule)
β βββ config_swarmind.py # interactive config menu
β βββ sync_opencode_global.py # sync to opencode global
β βββ deploy_all.py # propagate .opencode to 10 projects
β βββ gen_docs_api.py # API reference Markdown via Griffe (AST)
β βββ audit_docstrings.py # DOC gate: 0 functions without docstring
β βββ quality_audit.py # AGR audit (<900LC, except:pass, docstrings)
β βββ tdad_select.py # test dependency graph (AST)
β βββ validate_skills.py # skills spec validation (--strict)
βββ docs/
β βββ src/es/ # Documentation (Spanish, primary language)
β β βββ api/ # API reference generated by gen_docs_api.py
β β βββ roadmap/estado.md
β β βββ guide/, technical/, reference/, skills/
β βββ src/en/SUMMARY.md
β βββ .MEJORAS_SWARMIND.md
βββ .opencode/ # agents (22), skills (35, PEC universal), config β SSOT
βββ CHANGELOG.md
βββ pyproject.toml
βββ README.md
- Documentation (ES) β primary language β full docs in Spanish, the main documentation language.
- English summary
- Roadmap
- Skills registry + residency tiers β 35 skills, PEC universal, REFERENCE/SAVED/INSTALLED
- CHANGELOG
- Improvements log