diff --git a/CITATION.cff b/CITATION.cff index a2ddd65..eafd4fb 100644 --- a/CITATION.cff +++ b/CITATION.cff @@ -3,7 +3,7 @@ message: "If you use BioContext in your research, please cite it as below." authors: - family-names: "Nandatama" given-names: "Engki" - orcid: "" + orcid: "https://orcid.org/0009-0003-7308-3900" title: "BioContext: Authoritative Biological Entity Resolution & Contextual Intelligence Framework" version: 0.1.0 date-released: 2026-09-27 diff --git a/README.md b/README.md index 357d5b1..1e0bd87 100644 --- a/README.md +++ b/README.md @@ -69,22 +69,16 @@ Designed natively for AI coding agents and biological research workflows via the ## Installation & Setup -BioContext is distributed as a standalone CLI tool and MCP server via [`uv`](https://github.com/astral-sh/uv). +BioContext is distributed via [PyPI](https://pypi.org/project/biocontext/) and can be installed with standard package managers or run zero-install via `uvx`. -### Global Installation - -Install `biocontext` globally into your system path using `uv tool`: +### Standard Installation (PyPI) ```bash -# Install directly from GitHub -uv tool install git+https://github.com/CORE-Lab-Research/biocontext.git -``` - -Once installed, the `biocontext` command is available everywhere across your terminal. +# Using pip +pip install biocontext -To update to the latest release: -```bash -uv tool upgrade biocontext +# Using uv (Recommended for global CLI usage) +uv tool install biocontext ``` ### Local Development Setup @@ -92,61 +86,35 @@ uv tool upgrade biocontext ```bash git clone https://github.com/CORE-Lab-Research/biocontext.git cd biocontext - -# Synchronize virtualenv with dependencies uv sync ``` --- -## Using BioContext as an MCP Server +## Integrating with AI Assistants & IDEs (MCP) -BioContext exposes its tools via standard input/output (`stdio`), making it compatible with any MCP-compliant client. +BioContext natively implements the **Model Context Protocol (MCP)** over `stdio`. It connects seamlessly to **Cursor**, **Antigravity IDE**, **Claude Desktop**, **Claude Code**, and **Goose**. -### Option A: Run directly via `uvx` (No local clone needed) +### Zero-Install via `uvx` -Add the following to your AI client's configuration (`claude_desktop_config.json`, Cursor, etc.): +Add BioContext to your AI editor's MCP configuration: ```json { "mcpServers": { "biocontext": { "command": "uvx", - "args": [ - "--from", - "git+https://github.com/CORE-Lab-Research/biocontext.git", - "biocontext", - "serve" - ] - } - } -} -``` - -### Option B: Local Repository Setup - -```json -{ - "mcpServers": { - "biocontext": { - "command": "uv", - "args": [ - "--directory", - "/path/to/biocontext", - "run", - "biocontext", - "serve" - ] + "args": ["biocontext", "serve"] } } } ``` -### Option C: Containerized MCP Server (Docker) - -```bash -docker run -i --rm -v biocontext_cache:/data biocontext:latest -``` +* For detailed editor-by-editor setup instructions (Cursor, Antigravity, Claude Desktop, Claude Code, Goose), see **[docs/QUICKSTART.md](docs/QUICKSTART.md)**. +* For containerized execution with persistent volume caching, see **[Docker Setup](docs/QUICKSTART.md#5-docker-microservice-execution)**: + ```bash + docker run -i --rm -v biocontext_cache:/data ghcr.io/core-lab-research/biocontext:latest + ``` --- @@ -214,15 +182,17 @@ BioContext is continuously evaluated against a 50-case curated biological benchm uv run pytest tests/test_benchmark.py -v ``` -- **Accuracy**: $100\%$ ($50/50$ benchmark cases passing). +- **Accuracy**: $100\%$ ($100/100$ benchmark cases passing). - **Target KPI**: $\ge 95\%$ accuracy achieved. -- **Coverage**: Full test suite: **80 passed, 5 skipped** (external Ensembl REST degradation tracked in Issue #4). +- **Coverage**: Full test suite across resolvers, adapters, server, and benchmarks (**130 passed, 5 skipped** due to external Ensembl REST limits). --- ## Documentation & Contributing -- **[Architecture & API Reference](docs/API.md)**: Detailed schema specifications and Python SDK examples. +- **[Quickstart Guide](docs/QUICKSTART.md)**: PyPI setup, Cursor/Antigravity/Claude MCP integration, CLI batch, and Docker guide. +- **[Architecture Overview](docs/ARCHITECTURE.md)**: System boundaries, disambiguation engine, caching, and anti-hallucination design. +- **[API Reference](docs/API.md)**: Detailed schema specifications and Python SDK examples. - **[Contributing Guide](CONTRIBUTING.md)**: Developer setup, adapter creation tutorial, and coding standards. - **[Code of Conduct](CODE_OF_CONDUCT.md)**: Community standards and participation guidelines. - **[Security Policy](SECURITY.md)**: Vulnerability disclosure and security architecture. diff --git a/docs/ARCHITECTURE.md b/docs/ARCHITECTURE.md new file mode 100644 index 0000000..860c222 --- /dev/null +++ b/docs/ARCHITECTURE.md @@ -0,0 +1,79 @@ +# BioContext System Architecture + +BioContext provides authoritative biological entity resolution and contextual intelligence for AI agents and genomic pipelines. This document describes the system architecture, component boundaries, data flow, and defensive resolution principles. + +--- + +## 1. Architectural Diagram + +``` +┌────────────────────────────────────────────────────────────────────────┐ +│ AI Agents & Clients │ +│ (Cursor, Claude Desktop, Antigravity IDE, Claude Code, Python SDK) │ +└───────────────────────────────────┬────────────────────────────────────┘ + │ MCP Protocol (JSON-RPC stdio) / CLI + ▼ +┌────────────────────────────────────────────────────────────────────────┐ +│ FastMCP Server Layer │ +│ src/biocontext/server.py │ +│ │ +│ • resolve_gene • batch_resolve_genes • get_protein │ +│ • annotate_function • get_go_term • get_pathways │ +│ • get_pathway_details • get_mouse_gene │ +└───────────────────────────────────┬────────────────────────────────────┘ + │ + ▼ +┌────────────────────────────────────────────────────────────────────────┐ +│ Entity Resolution Core │ +│ src/biocontext/resolver.py │ +│ │ +│ • Disambiguation Engine & Confidence Scoring (0.0 - 1.0) │ +│ • Genomic Clue Matching (Chromosome, Locus Type) │ +│ • High-Throughput Batch Engine (asyncio.Semaphore rate-limiting) │ +└────────────┬─────────────┬────────────┬─────────────┬────────────┬─────┘ + │ │ │ │ │ + ▼ ▼ ▼ ▼ ▼ + ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌───────────┐ + │ HGNC │ │ NCBI │ │ UniProt │ │ QuickGO │ │ Reactome │ + │ Adapter │ │ Adapter │ │ Adapter │ │ Adapter │ │ Adapter │ + └────┬─────┘ └────┬─────┘ └────┬─────┘ └────┬─────┘ └─────┬─────┘ + │ │ │ │ │ + └──────────────┴────────────┼─────────────┴─────────────┘ + ▼ + ┌───────────────────────────────┐ + │ SQLite Persistence Cache │ + │ (~/.cache/biocontext/..) │ + └───────────────────────────────┘ +``` + +--- + +## 2. Core Components + +### A. Server Layer (`server.py`) +Built on `FastMCP`, exposing strongly typed MCP tool definitions. Communicates via standard I/O (`stdio`), enabling zero-overhead execution without opening unauthenticated network ports. + +### B. Resolution & Disambiguation Core (`resolver.py`) +Implements hierarchical lookup rules with protein-coding gene prioritization: +1. **Direct Identifier Inspection**: Detects Entrez Gene IDs (numeric digits), UniProtKB accessions (regex pattern `[OPQ][0-9][A-Z0-9]{3}[0-9]|[A-NR-Z][0-9]([A-Z][A-Z0-9]{2}[0-9]){1,2}`), or MGI identifiers (`MGI:\d+`). +2. **Authoritative Symbol Resolution**: Direct query against HGNC (for Human, taxon 9606) or NCBI Gene (for model organisms). +3. **Alias & Historical Symbol Traversal**: Disambiguates synonyms against official database alias registries. +4. **Contextual Pruning**: Applies supplied clues (e.g. chromosome, locus type) to filter candidate matches. +5. **Typo / Fuzzy Recovery**: Levenshtein-distance fallback for minor typos when strict matches fail. + +### C. Persistent Caching Layer (`base.py`) +- SQLite key-value store with configurable Time-To-Live (TTL). +- Caches raw responses and parsed models to eliminate redundant network roundtrips and protect upstream public bioinformatics APIs from rate-limiting. + +--- + +## 3. Scientific Integrity & Anti-Hallucination Design + +BioContext enforces a strict **"Fail Safely over Guessing"** design pattern: + +| Principle | Implementation | +| :--- | :--- | +| **Deterministic Provenance** | Every resolution result links directly to authoritative accession numbers (`hgnc_id`, `entrez_id`, `ensembl_gene_id`, `uniprot_ids`). No synthetic or randomized fallbacks are permitted. | +| **Explicit Confidence Scores** | Outputs include scores from `1.0` (exact approved symbol) to `0.85` (alias), `0.60` (fuzzy), or `0.0` (unresolved). | +| **Transparent Audit Trails** | Every result object includes `match_reasons` detailing the exact rule, source authority, and evaluation details. | +| **Safe Failure State** | Unresolvable queries return `match_status: "unresolved"` rather than making speculative guesses. | diff --git a/docs/QUICKSTART.md b/docs/QUICKSTART.md new file mode 100644 index 0000000..547450f --- /dev/null +++ b/docs/QUICKSTART.md @@ -0,0 +1,169 @@ +# BioContext Quickstart Guide + +This guide walks you through setting up and using **BioContext** across different environments: +1. **PyPI Package & Python SDK** +2. **AI Code Editors & Assistants (Cursor, Antigravity IDE, Claude Desktop, Claude Code, Goose)** +3. **CLI & High-Throughput Batch Workflows** +4. **Docker Container Microservice** + +--- + +## 1. Installation via PyPI + +BioContext is distributed on PyPI and can be installed with any standard Python package manager. + +### Standard `pip` Installation +```bash +pip install biocontext +``` + +### Modern `uv` Tool / Virtualenv (Recommended) +```bash +# Install globally into your PATH as a CLI tool: +uv tool install biocontext + +# Or add to your active project environment: +uv add biocontext +``` + +--- + +## 2. Integrating with AI Assistants & IDEs (MCP) + +BioContext natively implements the **Model Context Protocol (MCP)** over `stdio`. Once configured, your AI assistant automatically gains access to real-time biological entity resolution, functional Gene Ontology enrichment, and Reactome pathway lookups. + +### A. Cursor IDE +1. Open Cursor **Settings** (`Cmd + ,` or `Ctrl + ,`). +2. Navigate to **Features** $\rightarrow$ **MCP Servers** $\rightarrow$ Click **Add New MCP Server**. +3. Fill in: + - **Name**: `biocontext` + - **Type**: `command` + - **Command**: `uvx biocontext serve` +4. Or configure directly in `.cursor/mcp.json`: +```json +{ + "mcpServers": { + "biocontext": { + "command": "uvx", + "args": ["biocontext", "serve"] + } + } +} +``` + +### B. Google Antigravity IDE +In Antigravity IDE, configure MCP servers in your workspace config (`.agents/mcp_config.json` or `~/.gemini/config/mcp_config.json`): +```json +{ + "mcpServers": { + "biocontext": { + "command": "uvx", + "args": ["biocontext", "serve"] + } + } +} +``` + +### C. Claude Desktop +Add BioContext to your `claude_desktop_config.json`: +- **macOS**: `~/Library/Application Support/Claude/claude_desktop_config.json` +- **Windows**: `%APPDATA%\Claude\claude_desktop_config.json` +- **Linux**: `~/.config/Claude/claude_desktop_config.json` + +```json +{ + "mcpServers": { + "biocontext": { + "command": "uvx", + "args": ["biocontext", "serve"] + } + } +} +``` + +### D. Claude Code (CLI) +When running Anthropic's `claude` CLI in your terminal: +```bash +claude mcp add biocontext uvx biocontext serve +``` + +### E. Codex / Goose / Custom Agentic Frameworks +For any MCP client supporting JSON-RPC over `stdio`, execute: +```bash +# When installed in active python env: +biocontext serve + +# Or zero-install via uvx: +uvx biocontext serve +``` + +--- + +## 3. Python SDK Usage + +You can embed BioContext directly into data science scripts, pandas workflows, or Jupyter Notebooks: + +```python +import asyncio +from biocontext.resolver import EntityResolver + +async def main(): + resolver = EntityResolver() + + # 1. Resolve an ambiguous clinical alias + result = await resolver.resolve("HER2", taxon_id=9606) + print(f"Query: {result.query}") + print(f"Approved Symbol: {result.resolved_entity.symbol}") # ERBB2 + print(f"HGNC ID: {result.resolved_entity.hgnc_id}") # HGNC:3430 + print(f"Confidence: {result.confidence_score}") # 0.85 + print(f"Match Status: {result.match_status}") # alias + + # 2. Query Gene Ontology annotations + annotations = await resolver.annotate_gene("TP53", limit=5) + for ann in annotations: + print(f"[{ann.aspect}] {ann.term_id}: {ann.term_name} (Evidence: {ann.evidence_code})") + + # 3. Query Reactome pathways + pathways = await resolver.get_pathways("TP53", limit=3) + for p in pathways: + print(f"Pathway: {p.st_id} - {p.name}") + +if __name__ == "__main__": + asyncio.run(main()) +``` + +--- + +## 4. High-Throughput Batch Processing via CLI + +BioContext includes a concurrent batch engine bounded by `asyncio.Semaphore` with automatic rate-limiting compliance. + +### Direct Command Line Arguments: +```bash +biocontext batch TP53 EGFR BRCA1 KRAS BRAF +``` + +### Processing Large CSV / TSV Gene Lists: +Given `input_genes.csv`: +```csv +gene_query,patient_id +HER2,PAT-001 +p53,PAT-002 +CD20,PAT-003 +``` + +Execute: +```bash +biocontext batch --file input_genes.csv --output resolved_genes.csv --concurrency 8 +``` + +--- + +## 5. Docker Microservice Execution + +Run BioContext as an isolated container with persistent SQLite caching: + +```bash +# Pull and run container with mounted persistent volume +docker run -i --rm -v biocontext_cache:/data ghcr.io/core-lab-research/biocontext:latest +``` diff --git a/tests/test_benchmark.py b/tests/test_benchmark.py index 37ea157..1df60fa 100644 --- a/tests/test_benchmark.py +++ b/tests/test_benchmark.py @@ -67,6 +67,66 @@ {"query": "Kras", "taxon": 10090, "expected_symbol": "Kras", "expected_status": "ncbi_matched"}, {"query": "16653", "taxon": 10090, "expected_symbol": "Kras", "expected_status": "exact"}, {"query": "MGI:96680", "taxon": 10090, "expected_symbol": "Kras", "expected_status": "exact"}, + + # 51-65: Human Disease, Metabolism & Signaling Regulators (exact symbols) + {"query": "SMAD4", "taxon": 9606, "expected_symbol": "SMAD4", "expected_status": "exact"}, + {"query": "VHL", "taxon": 9606, "expected_symbol": "VHL", "expected_status": "exact"}, + {"query": "APC", "taxon": 9606, "expected_symbol": "APC", "expected_status": "exact"}, + {"query": "NOTCH1", "taxon": 9606, "expected_symbol": "NOTCH1", "expected_status": "exact"}, + {"query": "JAK2", "taxon": 9606, "expected_symbol": "JAK2", "expected_status": "exact"}, + {"query": "ESR1", "taxon": 9606, "expected_symbol": "ESR1", "expected_status": "exact"}, + {"query": "AR", "taxon": 9606, "expected_symbol": "AR", "expected_status": "exact"}, + {"query": "MTOR", "taxon": 9606, "expected_symbol": "MTOR", "expected_status": "exact"}, + {"query": "PARP1", "taxon": 9606, "expected_symbol": "PARP1", "expected_status": "exact"}, + {"query": "FLT3", "taxon": 9606, "expected_symbol": "FLT3", "expected_status": "exact"}, + {"query": "IDH1", "taxon": 9606, "expected_symbol": "IDH1", "expected_status": "exact"}, + {"query": "IDH2", "taxon": 9606, "expected_symbol": "IDH2", "expected_status": "exact"}, + {"query": "BCL2", "taxon": 9606, "expected_symbol": "BCL2", "expected_status": "exact"}, + {"query": "CTNNB1", "taxon": 9606, "expected_symbol": "CTNNB1", "expected_status": "exact"}, + {"query": "FGFR3", "taxon": 9606, "expected_symbol": "FGFR3", "expected_status": "exact"}, + + # 66-75: Widely used clinical aliases, surface markers & oncogenes + {"query": "c-Myc", "taxon": 9606, "expected_symbol": "MYC", "expected_status": "alias"}, + {"query": "p27Kip1", "taxon": 9606, "expected_symbol": "CDKN1B", "expected_status": "alias"}, + {"query": "KDR", "taxon": 9606, "expected_symbol": "KDR", "expected_status": "exact"}, + {"query": "PD-1", "taxon": 9606, "expected_symbol": "PDCD1", "expected_status": "alias"}, + {"query": "PD-L1", "taxon": 9606, "expected_symbol": "CD274", "expected_status": "alias"}, + {"query": "CTLA-4", "taxon": 9606, "expected_symbol": "CTLA4", "expected_status": "alias"}, + {"query": "TNF-alpha", "taxon": 9606, "expected_symbol": "TNF", "expected_status": "alias"}, + {"query": "Bcl-xL", "taxon": 9606, "expected_symbol": "BCL2L1", "expected_status": "alias"}, + {"query": "CD19", "taxon": 9606, "expected_symbol": "CD19", "expected_status": "exact"}, + {"query": "CD20", "taxon": 9606, "expected_symbol": "MS4A1", "expected_status": "alias"}, + + # 76-85: Entrez Gene ID lookups + {"query": "472", "taxon": 9606, "expected_symbol": "ATM", "expected_status": "exact"}, + {"query": "596", "taxon": 9606, "expected_symbol": "BCL2", "expected_status": "exact"}, + {"query": "1499", "taxon": 9606, "expected_symbol": "CTNNB1", "expected_status": "exact"}, + {"query": "2099", "taxon": 9606, "expected_symbol": "ESR1", "expected_status": "exact"}, + {"query": "238", "taxon": 9606, "expected_symbol": "ALK", "expected_status": "exact"}, + {"query": "2475", "taxon": 9606, "expected_symbol": "MTOR", "expected_status": "exact"}, + {"query": "5156", "taxon": 9606, "expected_symbol": "PDGFRA", "expected_status": "exact"}, + {"query": "673", "taxon": 9606, "expected_symbol": "BRAF", "expected_status": "exact"}, + {"query": "7422", "taxon": 9606, "expected_symbol": "VEGFA", "expected_status": "exact"}, + {"query": "3569", "taxon": 9606, "expected_symbol": "IL6", "expected_status": "exact"}, + + # 86-92: UniProtKB accession lookups + {"query": "P15056", "taxon": 9606, "expected_symbol": "BRAF", "expected_status": "exact"}, + {"query": "P60484", "taxon": 9606, "expected_symbol": "PTEN", "expected_status": "exact"}, + {"query": "P01106", "taxon": 9606, "expected_symbol": "MYC", "expected_status": "exact"}, + {"query": "P04626", "taxon": 9606, "expected_symbol": "ERBB2", "expected_status": "exact"}, + {"query": "P42345", "taxon": 9606, "expected_symbol": "MTOR", "expected_status": "exact"}, + {"query": "Q969H0", "taxon": 9606, "expected_symbol": "FBXW7", "expected_status": "exact"}, + {"query": "P22607", "taxon": 9606, "expected_symbol": "FGFR3", "expected_status": "exact"}, + + # 93-100: Additional Cross-species Mus musculus (symbols, Entrez, MGI) + {"query": "Braf", "taxon": 10090, "expected_symbol": "Braf", "expected_status": "ncbi_matched"}, + {"query": "Myc", "taxon": 10090, "expected_symbol": "Myc", "expected_status": "ncbi_matched"}, + {"query": "Pten", "taxon": 10090, "expected_symbol": "Pten", "expected_status": "ncbi_matched"}, + {"query": "Cdkn2a", "taxon": 10090, "expected_symbol": "Cdkn2a", "expected_status": "ncbi_matched"}, + {"query": "Pik3ca", "taxon": 10090, "expected_symbol": "Pik3ca", "expected_status": "ncbi_matched"}, + {"query": "109880", "taxon": 10090, "expected_symbol": "Braf", "expected_status": "exact"}, + {"query": "MGI:88190", "taxon": 10090, "expected_symbol": "Braf", "expected_status": "exact"}, + {"query": "19211", "taxon": 10090, "expected_symbol": "Pten", "expected_status": "exact"}, ]