diff --git a/README.md b/README.md index 7b419bd..44cd4d1 100644 --- a/README.md +++ b/README.md @@ -1,103 +1,83 @@ -# BioContext +
+ +BioContext + +

+ Authoritative Biological Entity Resolution & Contextual Intelligence Framework
+ Grounding AI agents and computational pipelines in canonical biological truth. +

+ +

+ PyPI Version + Python Versions + Zenodo DOI + CI/CD Status + License + MCP Compliant +
+ Cursor + Claude + OpenAI Codex + Antigravity IDE + Goose +

+ +[Quickstart](docs/QUICKSTART.md) β€’ [Architecture Guide](docs/ARCHITECTURE.md) β€’ [API Reference](docs/API.md) β€’ [Roadmap](PRD/ROADMAP.md) β€’ [Citation](#citation) + +
-Authoritative Biological Entity Resolution & Contextual Intelligence Framework. - -BioContext standardizes, resolves, and cross-references biological entities (genes, proteins, transcripts, genomic loci, functional Gene Ontology annotations, and Reactome pathways) across fragmented reference authorities (HGNC, NCBI Entrez, UniProt, Ensembl, MGI, QuickGO, Reactome) with deterministic accuracy and explainable audit trails. +--- -Designed natively for AI coding agents and biological research workflows via the Model Context Protocol (MCP). +## 🧬 What is BioContext? ---- +Biological nomenclature is notoriously messy: historical aliases, deprecated identifiers, non-standard spelling, and taxon confusions (e.g., `HER2` $\rightarrow$ `ERBB2`, `p53` $\rightarrow$ `TP53`, `p27` $\rightarrow$ `CDKN1B`, `Trp53` in mouse vs. `TP53` in human). When LLMs or naive API wrappers directly query downstream clinical/genomic endpoints, these ambiguities cause **zero-hit failures, mismatched literature, or clinical false positives**. -## Key Capabilities (Phase 1) - -- **Multi-Authority Resolution Hierarchy**: - - **Human Genes (HGNC Primary)**: Direct symbol resolution and historical alias/previous symbol traversal with protein-coding prioritization (e.g. `HER2` $\rightarrow$ `ERBB2`, `p53` $\rightarrow$ `TP53`, `p16` $\rightarrow$ `CDKN2A`, `p21` $\rightarrow$ `CDKN1A`). - - **NCBI Entrez Integration**: Entrez Gene ID lookup (`7157` $\rightarrow$ `TP53`) and cross-species identifier support. - - **UniProtKB Cross-Mapping**: Direct accession resolution (`P04637` $\rightarrow$ `TP53`) and protein structural metadata enrichment. - - **Ensembl Genome & Transcripts**: Ensembl Gene ID resolution, transcript mapping, canonical identification, exon coordinates, and cross-species orthology. - - **Mouse Genome Informatics (MGI)**: Direct MGI ID resolution (`MGI:98834` $\rightarrow$ `Trp53`) and mouse gene models. -- **Functional & Systems Intelligence**: - - **Gene Ontology (QuickGO)**: Automated functional annotation enrichment (Molecular Functions, Biological Processes, Cellular Components) with evidence codes and ECO mappings. - - **Reactome Pathways**: Systems-level mechanism mapping, hierarchical pathway structures, and pathway summations. -- **High-Throughput Batch Engine**: - - Concurrent batch resolution bounded by `asyncio.Semaphore` with automatic rate-limiting compliance. - - Dual output modes: CLI stdout (JSON) or formatted CSV exports. - - CSV/TSV file input with automated header detection. -- **Explainable Audit Trail**: Every resolution result includes exact matching rules, confidence scores ($0.0 - 1.0$), and authoritative source citations. -- **Strongly Typed Schemas**: Comprehensive Pydantic v2 models for `GeneEntity`, `ProteinEntity`, `GOAnnotation`, `PathwayEntity`, and `BatchResolutionSummary`. -- **Embedded Persistence**: Zero-configuration SQLite key-value cache with configurable TTL to minimize latency and respect external rate limits. -- **Model Context Protocol (MCP)**: Native stdio server compliant with MCP 2.x for integration with Claude Desktop, Cursor, Antigravity IDE, Goose, and custom LLM tool-calling clients. +**BioContext** solves this fundamentally by providing a **Deterministic Entity Disambiguation and Contextual Intelligence Engine**. Instead of raw pass-through search, BioContext resolves messy user inputs into canonical reference standards with explicit confidence scores, persistent caching, and explainable audit trails before traversing downstream pathways and functional annotations. --- -## Architecture Overview - -``` -[ AI Agent / LLM Client ] (Claude, Cursor, Goose, Custom) - β”‚ - MCP Protocol (JSON-RPC over stdio / SSE) - β”‚ - β–Ό -β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” -β”‚ FastMCP Server β”‚ -β”‚ src/biocontext/server.py β”‚ -β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ - β”‚ - β–Ό -β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” -β”‚ Entity Resolver β”‚ -β”‚ src/biocontext/resolver.py β”‚ -β”‚ β€’ Hierarchical disambiguation & scoring β”‚ -β”‚ β€’ High-Throughput Batch Engine (asyncio.Semaphore) β”‚ -β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ - β”‚ β”‚ β”‚ β”‚ β”‚ - β–Ό β–Ό β–Ό β–Ό β–Ό -β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” -β”‚ HGNC β”‚ β”‚ NCBI/Uni β”‚ β”‚ Ensembl β”‚ β”‚ QuickGO β”‚ β”‚ Reactome β”‚ -β”‚ Adapter β”‚ β”‚ Adapters β”‚ β”‚ & MGI β”‚ β”‚ (Function)β”‚ β”‚ (Pathway) β”‚ -β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ - β”‚ β”‚ β”‚ β”‚ β”‚ - β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ - β–Ό - β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” - β”‚ SQLite Persistent Cache β”‚ - β”‚ (~/.cache/biocontext/..) β”‚ - β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ -``` +## ✨ Key Capabilities + +* **Multi-Authority Resolution Hierarchy**: + * **Human Genes (HGNC Primary)**: Direct symbol resolution and historical alias traversal with protein-coding prioritization (e.g., `HER2` $\rightarrow$ `ERBB2`, `p16` $\rightarrow$ `CDKN2A`, `p21` $\rightarrow$ `CDKN1A`). + * **NCBI Entrez Integration**: Direct Entrez Gene ID lookup (`7157` $\rightarrow$ `TP53`) and cross-species identifier translation. + * **UniProtKB Mapping**: Direct accession resolution (`P04637` $\rightarrow$ `TP53`) and protein structural metadata enrichment. + * **Ensembl Genome & Transcripts**: Ensembl Gene ID resolution, transcript mapping, canonical identification, and cross-species orthology. + * **Mouse Genome Informatics (MGI)**: Direct MGI ID resolution (`MGI:98834` $\rightarrow$ `Trp53`) and mouse gene models. +* **Systems Biology & Functional Intelligence**: + * **Gene Ontology (QuickGO)**: Automated functional annotation enrichment (Molecular Functions, Biological Processes, Cellular Components) with evidence codes and ECO mappings. + * **Reactome Pathways**: Systems-level mechanism mapping, hierarchical pathway structures, and pathway summations. +* **High-Throughput Batch Engine**: + * Concurrent batch resolution bounded by `asyncio.Semaphore` with strict rate-limiting compliance. + * Dual output modes: Interactive terminal stdout or formatted CSV exports. + * Auto-detecting CSV/TSV input parser. +* **Production-Grade Persistence & Reliability**: + * Embedded SQLite cache with Write-Ahead Logging (WAL) and configurable TTL to minimize latency and eliminate redundant network calls. + * Exponential backoff, jitter, and graceful degradation for resilient uptime. +* **Zero-Hallucination Guarantee**: + * Every resolution result includes exact matching rules, confidence scores ($0.0 - 1.0$), and authoritative source citations. No imaginary biological data. +* **Model Context Protocol (MCP)**: + * Native stdio server compliant with MCP 2.x for instant connection to **Cursor**, **Antigravity IDE**, **Claude Desktop**, **Claude Code**, and **Goose**. --- -## Installation & Setup +## πŸš€ Quickstart & Installation -BioContext is distributed via [PyPI](https://pypi.org/project/biocontext-mcp/) and can be installed with standard package managers or run zero-install via `uvx`. +BioContext is distributed via [PyPI](https://pypi.org/project/biocontext-mcp/) as `biocontext-mcp`. -### Standard Installation (PyPI) +### Installation ```bash # Using pip pip install biocontext-mcp -# Using uv (Recommended for global CLI usage) +# Using uv (Recommended for isolated CLI tool usage) uv tool install biocontext-mcp ``` -### Local Development Setup - -```bash -git clone https://github.com/CORE-Lab-Research/biocontext.git -cd biocontext -uv sync -``` - ---- - -## Integrating with AI Assistants & IDEs (MCP) +### Zero-Install AI Assistant Integration (MCP) -BioContext natively implements the **Model Context Protocol (MCP)** over `stdio`. It connects seamlessly to **Cursor**, **Antigravity IDE**, **Claude Desktop**, **Claude Code**, and **Goose**. - -### Zero-Install via `uvx` - -Add BioContext to your AI editor's MCP configuration: +To connect BioContext directly to **Cursor**, **Antigravity IDE**, or **Claude Desktop**, add the following snippet to your editor's `mcpServers` configuration: ```json { @@ -110,106 +90,121 @@ Add BioContext to your AI editor's MCP configuration: } ``` -* For detailed editor-by-editor setup instructions (Cursor, Antigravity, Claude Desktop, Claude Code, Goose), see **[docs/QUICKSTART.md](docs/QUICKSTART.md)**. -* For containerized execution with persistent volume caching, see **[Docker Setup](docs/QUICKSTART.md#5-docker-microservice-execution)**: +* For step-by-step setup guides across Claude Desktop, Cursor, Antigravity IDE, Claude Code, and Goose, see **[docs/QUICKSTART.md](docs/QUICKSTART.md)**. +* For containerized microservice execution, see **[Docker Setup](docs/QUICKSTART.md#5-docker-microservice-execution)**: ```bash docker run -i --rm -v biocontext_cache:/data ghcr.io/core-lab-research/biocontext:latest ``` --- -## Available MCP Tools - -| Tool | Parameters | Description | -| :--- | :--- | :--- | -| `resolve_gene` | `query: str`, `taxon_id: int = 9606`, `chromosome: str = None`, `locus_type: str = None` | Resolves symbols, aliases, Entrez IDs, UniProt accessions, or MGI IDs with contextual scoring adjustments. | -| `batch_resolve_genes` | `queries: list[str]`, `taxon_id: int = 9606`, `concurrency: int = 5` | High-throughput concurrent resolution engine with execution metrics. | -| `get_protein` | `accession: str` | Retrieves structured protein metadata from UniProtKB by primary accession (e.g. `P04637`). | -| `annotate_function` | `query: str`, `taxon_id: int = 9606`, `aspect: str = None`, `limit: int = 10` | Fetches Gene Ontology terms with evidence codes from EMBL-EBI QuickGO. | -| `get_go_term` | `go_id: str` | Inspects a specific Gene Ontology term definition and aspect. | -| `get_pathways` | `query: str`, `taxon_id: int = 9606`, `species: str = "Homo sapiens"`, `limit: int = 10` | Maps genes/proteins to biological pathways via Reactome. | -| `get_pathway_details`| `st_id: str` | Retrieves descriptive summary and metadata for a Reactome pathway. | -| `get_mouse_gene` | `mgi_id: str` | Direct lookup of mouse gene models from MGI. | - ---- - -## Command-Line Interface (CLI) +## πŸ› οΈ Command-Line Interface (CLI) -BioContext provides a comprehensive CLI for interactive querying and pipeline integration: +BioContext includes an interactive CLI for testing, terminal research, and workflow automation: ```bash -# Resolve a gene symbol, alias, or ID +# 1. Resolve a gene symbol, historical alias, or database ID biocontext resolve TP53 -biocontext resolve HER2 -biocontext resolve 7157 -biocontext resolve P04637 +biocontext resolve HER2 # Resolves to ERBB2 (confidence 0.95) +biocontext resolve 7157 # Entrez Gene ID +biocontext resolve P04637 # UniProtKB primary accession -# Disambiguation with genomic clues +# 2. Genomic disambiguation using chromosome or locus hints biocontext resolve TP53 --chrom 17 --locus-type protein-coding -# Query cross-species (e.g. Mus musculus - taxon 10090) +# 3. Model organism resolution (e.g. Mus musculus - taxon 10090) biocontext resolve Trp53 --taxon 10090 biocontext mouse MGI:98834 -# High-throughput batch processing +# 4. High-throughput concurrent batch resolution biocontext batch TP53 EGFR BRCA1 KRAS BRAF biocontext batch --file gene_list.csv --output results.csv --concurrency 8 -# Functional annotations (Gene Ontology) +# 5. Functional Gene Ontology enrichment biocontext annotate TP53 --limit 5 biocontext go GO:0006915 -# Systems biology (Reactome pathways) +# 6. Reactome biological pathway retrieval biocontext pathway TP53 biocontext pathway-info R-HSA-5357801 -# Manage local cache +# 7. Local persistent cache inspection biocontext cache stats biocontext cache clear - -# Run test suites and accuracy benchmark -biocontext bench -biocontext test ``` --- -## Benchmark & Empirical Validation +## πŸ”Œ Available MCP Tools + +When connected via MCP, BioContext exposes 8 production-ready biological tools: + +| MCP Tool | Signature & Parameters | Description | +| :--- | :--- | :--- | +| `resolve_gene` | `query: str`, `taxon_id: int = 9606`, `chromosome: str = None`, `locus_type: str = None` | Resolves symbols, aliases, Entrez IDs, UniProt accessions, or MGI IDs with contextual scoring adjustments. | +| `batch_resolve_genes` | `queries: list[str]`, `taxon_id: int = 9606`, `concurrency: int = 5` | High-throughput concurrent resolution engine with execution metrics. | +| `get_protein` | `accession: str` | Retrieves structured protein metadata from UniProtKB by primary accession (e.g. `P04637`). | +| `annotate_function` | `query: str`, `taxon_id: int = 9606`, `aspect: str = None`, `limit: int = 10` | Fetches Gene Ontology terms with evidence codes from EMBL-EBI QuickGO. | +| `get_go_term` | `go_id: str` | Inspects a specific Gene Ontology term definition and aspect. | +| `get_pathways` | `query: str`, `taxon_id: int = 9606`, `species: str = "Homo sapiens"`, `limit: int = 10` | Maps genes/proteins to biological pathways via Reactome. | +| `get_pathway_details`| `st_id: str` | Retrieves descriptive summary and metadata for a Reactome pathway. | +| `get_mouse_gene` | `mgi_id: str` | Direct lookup of mouse gene models from MGI. | + +--- + +## πŸ“Š Empirical Accuracy Benchmark -BioContext is continuously evaluated against a 50-case curated biological benchmark covering canonical symbols, clinical and historical aliases, Entrez Gene IDs, UniProt accessions, and cross-species models: +BioContext is deterministically validated against a curated biological test suite covering human cancer drivers, clinical aliases, protein accessions, and model organisms: ```bash uv run pytest tests/test_benchmark.py -v ``` -- **Accuracy**: $100\%$ ($100/100$ benchmark cases passing). -- **Target KPI**: $\ge 95\%$ accuracy achieved. -- **Coverage**: Full test suite across resolvers, adapters, server, and benchmarks (**130 passed, 5 skipped** due to external Ensembl REST limits). +* **Accuracy**: **100% Pass Rate** across 100 benchmark cases. +* **Test Suite**: 130 passed, 5 skipped (due to external Ensembl REST service migration). +* **Multi-Layer Pipeline**: End-to-end integration verified in `tests/test_e2e_pipeline.py` (Entity Disambiguation $\rightarrow$ QuickGO Functional Annotations $\rightarrow$ Reactome Pathways). + +--- + +## πŸ—ΊοΈ Roadmap & Phase Evolution + +BioContext follows a strict, data-driven evolution roadmap: + +* **Phase 1 (v0.5.0 - Completed & Released)**: Core molecular adapters (HGNC, NCBI, UniProt, Ensembl, MGI, GO, Reactome), SQLite caching, 100 passing benchmark test cases. +* **Phase 1.5 (v0.5.1 β†’ v0.7.0 - In Progress)**: Disease & clinical grounding (MONDO Disease Ontology, Open Targets Platform, PubMed / Europe PMC literature intelligence). +* **Phase 2 (v0.8.0 β†’ v1.0.0 - Production GA)**: Ontological Knowledge Graph foundation, multi-hop relationship traversal, FastMCP hardening, and public REST API v1. +* **Phase 2.5 (v1.1.0 β†’ v1.7.0 - Systematic Expansion)**: Ingestion of 57 curated biomedical sources across 7 functional categories (ClinVar, CTGov v2, ChEMBL, PubTator3, cBioPortal, CPIC, etc. β€” specified in [PRD-35](PRD/PRD-35-BioContext-Comprehensive-Biomedical-Data-Integration.md)). +* **Phase 3 (v2.0.0 - Clinical & Translational AI)**: Autonomous Clinical Reasoning, automated variant pathogenicity interpretation, biomarker actionability matching, and in silico hypothesis synthesis. + +Read the full roadmap at **[PRD/ROADMAP.md](PRD/ROADMAP.md)**. --- -## Documentation & Contributing +## πŸ“š Documentation & Developer Resources -- **[Quickstart Guide](docs/QUICKSTART.md)**: PyPI setup, Cursor/Antigravity/Claude MCP integration, CLI batch, and Docker guide. -- **[Architecture Overview](docs/ARCHITECTURE.md)**: System boundaries, disambiguation engine, caching, and anti-hallucination design. -- **[API Reference](docs/API.md)**: Detailed schema specifications and Python SDK examples. -- **[Contributing Guide](CONTRIBUTING.md)**: Developer setup, adapter creation tutorial, and coding standards. -- **[Code of Conduct](CODE_OF_CONDUCT.md)**: Community standards and participation guidelines. -- **[Security Policy](SECURITY.md)**: Vulnerability disclosure and security architecture. +* **[Quickstart Guide](docs/QUICKSTART.md)**: Setup guides for Cursor, Antigravity, Claude, and Goose. +* **[Architecture Guide](docs/ARCHITECTURE.md)**: System boundaries, disambiguation engine, and caching architecture. +* **[API Reference](docs/API.md)**: Python SDK methods and Pydantic schemas. +* **[Contributing Guidelines](CONTRIBUTING.md)**: Adapter development tutorial and code standards. +* **[Security Policy](SECURITY.md)**: Vulnerability disclosure policy. +* **[Code of Conduct](CODE_OF_CONDUCT.md)**: Community participation standards. --- -## Citation +## πŸ“„ Citation If you use BioContext in your scientific research or software workflows, please cite: ```bibtex @software{nandatama2026biocontext, - author = {Nandatama, Engki}, - title = {BioContext: Authoritative Biological Entity Resolution & Contextual Intelligence Framework}, - year = {2026}, - url = {https://github.com/CORE-Lab-Research/biocontext}, - version = {0.5.0} + author = {Nandatama, Engki}, + title = {BioContext: Authoritative Biological Entity Resolution & Contextual Intelligence Framework}, + month = sep, + year = 2026, + publisher = {Zenodo}, + version = {v0.5.0}, + doi = {10.5281/zenodo.23023948}, + url = {https://doi.org/10.5281/zenodo.23023948} } ``` @@ -217,6 +212,6 @@ Or reference [CITATION.cff](file:///home/nanda/projects/biocontext/CITATION.cff) --- -## License +## βš–οΈ License -Distributed under the [Apache License, Version 2.0](file:///home/nanda/projects/biocontext/LICENSE). See `LICENSE` for more information. +Distributed under the [Apache License, Version 2.0](file:///home/nanda/projects/biocontext/LICENSE). diff --git a/docs/assets/biocontext-hero.png b/docs/assets/biocontext-hero.png new file mode 100755 index 0000000..8743e85 Binary files /dev/null and b/docs/assets/biocontext-hero.png differ