Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion CITATION.cff
Original file line number Diff line number Diff line change
Expand Up @@ -3,7 +3,7 @@ message: "If you use BioContext in your research, please cite it as below."
authors:
- family-names: "Nandatama"
given-names: "Engki"
orcid: ""
orcid: "https://orcid.org/0009-0003-7308-3900"
title: "BioContext: Authoritative Biological Entity Resolution & Contextual Intelligence Framework"
version: 0.1.0
date-released: 2026-09-27
Expand Down
72 changes: 21 additions & 51 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -69,84 +69,52 @@ Designed natively for AI coding agents and biological research workflows via the

## Installation & Setup

BioContext is distributed as a standalone CLI tool and MCP server via [`uv`](https://github.com/astral-sh/uv).
BioContext is distributed via [PyPI](https://pypi.org/project/biocontext/) and can be installed with standard package managers or run zero-install via `uvx`.

### Global Installation

Install `biocontext` globally into your system path using `uv tool`:
### Standard Installation (PyPI)

```bash
# Install directly from GitHub
uv tool install git+https://github.com/CORE-Lab-Research/biocontext.git
```

Once installed, the `biocontext` command is available everywhere across your terminal.
# Using pip
pip install biocontext

To update to the latest release:
```bash
uv tool upgrade biocontext
# Using uv (Recommended for global CLI usage)
uv tool install biocontext
```

### Local Development Setup

```bash
git clone https://github.com/CORE-Lab-Research/biocontext.git
cd biocontext

# Synchronize virtualenv with dependencies
uv sync
```

---

## Using BioContext as an MCP Server
## Integrating with AI Assistants & IDEs (MCP)

BioContext exposes its tools via standard input/output (`stdio`), making it compatible with any MCP-compliant client.
BioContext natively implements the **Model Context Protocol (MCP)** over `stdio`. It connects seamlessly to **Cursor**, **Antigravity IDE**, **Claude Desktop**, **Claude Code**, and **Goose**.

### Option A: Run directly via `uvx` (No local clone needed)
### Zero-Install via `uvx`

Add the following to your AI client's configuration (`claude_desktop_config.json`, Cursor, etc.):
Add BioContext to your AI editor's MCP configuration:

```json
{
"mcpServers": {
"biocontext": {
"command": "uvx",
"args": [
"--from",
"git+https://github.com/CORE-Lab-Research/biocontext.git",
"biocontext",
"serve"
]
}
}
}
```

### Option B: Local Repository Setup

```json
{
"mcpServers": {
"biocontext": {
"command": "uv",
"args": [
"--directory",
"/path/to/biocontext",
"run",
"biocontext",
"serve"
]
"args": ["biocontext", "serve"]
}
}
}
```

### Option C: Containerized MCP Server (Docker)

```bash
docker run -i --rm -v biocontext_cache:/data biocontext:latest
```
* For detailed editor-by-editor setup instructions (Cursor, Antigravity, Claude Desktop, Claude Code, Goose), see **[docs/QUICKSTART.md](docs/QUICKSTART.md)**.
* For containerized execution with persistent volume caching, see **[Docker Setup](docs/QUICKSTART.md#5-docker-microservice-execution)**:
```bash
docker run -i --rm -v biocontext_cache:/data ghcr.io/core-lab-research/biocontext:latest
```

---

Expand Down Expand Up @@ -214,15 +182,17 @@ BioContext is continuously evaluated against a 50-case curated biological benchm
uv run pytest tests/test_benchmark.py -v
```

- **Accuracy**: $100\%$ ($50/50$ benchmark cases passing).
- **Accuracy**: $100\%$ ($100/100$ benchmark cases passing).
- **Target KPI**: $\ge 95\%$ accuracy achieved.
- **Coverage**: Full test suite: **80 passed, 5 skipped** (external Ensembl REST degradation tracked in Issue #4).
- **Coverage**: Full test suite across resolvers, adapters, server, and benchmarks (**130 passed, 5 skipped** due to external Ensembl REST limits).

---

## Documentation & Contributing

- **[Architecture & API Reference](docs/API.md)**: Detailed schema specifications and Python SDK examples.
- **[Quickstart Guide](docs/QUICKSTART.md)**: PyPI setup, Cursor/Antigravity/Claude MCP integration, CLI batch, and Docker guide.
- **[Architecture Overview](docs/ARCHITECTURE.md)**: System boundaries, disambiguation engine, caching, and anti-hallucination design.
- **[API Reference](docs/API.md)**: Detailed schema specifications and Python SDK examples.
- **[Contributing Guide](CONTRIBUTING.md)**: Developer setup, adapter creation tutorial, and coding standards.
- **[Code of Conduct](CODE_OF_CONDUCT.md)**: Community standards and participation guidelines.
- **[Security Policy](SECURITY.md)**: Vulnerability disclosure and security architecture.
Expand Down
79 changes: 79 additions & 0 deletions docs/ARCHITECTURE.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,79 @@
# BioContext System Architecture

BioContext provides authoritative biological entity resolution and contextual intelligence for AI agents and genomic pipelines. This document describes the system architecture, component boundaries, data flow, and defensive resolution principles.

---

## 1. Architectural Diagram

```
┌────────────────────────────────────────────────────────────────────────┐
│ AI Agents & Clients │
│ (Cursor, Claude Desktop, Antigravity IDE, Claude Code, Python SDK) │
└───────────────────────────────────┬────────────────────────────────────┘
│ MCP Protocol (JSON-RPC stdio) / CLI
▼
┌────────────────────────────────────────────────────────────────────────┐
│ FastMCP Server Layer │
│ src/biocontext/server.py │
│ │
│ • resolve_gene • batch_resolve_genes • get_protein │
│ • annotate_function • get_go_term • get_pathways │
│ • get_pathway_details • get_mouse_gene │
└───────────────────────────────────┬────────────────────────────────────┘
│
▼
┌────────────────────────────────────────────────────────────────────────┐
│ Entity Resolution Core │
│ src/biocontext/resolver.py │
│ │
│ • Disambiguation Engine & Confidence Scoring (0.0 - 1.0) │
│ • Genomic Clue Matching (Chromosome, Locus Type) │
│ • High-Throughput Batch Engine (asyncio.Semaphore rate-limiting) │
└────────────┬─────────────┬────────────┬─────────────┬────────────┬─────┘
│ │ │ │ │
▼ ▼ ▼ ▼ ▼
┌──────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ ┌───────────┐
│ HGNC │ │ NCBI │ │ UniProt │ │ QuickGO │ │ Reactome │
│ Adapter │ │ Adapter │ │ Adapter │ │ Adapter │ │ Adapter │
└────┬─────┘ └────┬─────┘ └────┬─────┘ └────┬─────┘ └─────┬─────┘
│ │ │ │ │
└──────────────┴────────────┼─────────────┴─────────────┘
▼
┌───────────────────────────────┐
│ SQLite Persistence Cache │
│ (~/.cache/biocontext/..) │
└───────────────────────────────┘
```

---

## 2. Core Components

### A. Server Layer (`server.py`)
Built on `FastMCP`, exposing strongly typed MCP tool definitions. Communicates via standard I/O (`stdio`), enabling zero-overhead execution without opening unauthenticated network ports.

### B. Resolution & Disambiguation Core (`resolver.py`)
Implements hierarchical lookup rules with protein-coding gene prioritization:
1. **Direct Identifier Inspection**: Detects Entrez Gene IDs (numeric digits), UniProtKB accessions (regex pattern `[OPQ][0-9][A-Z0-9]{3}[0-9]|[A-NR-Z][0-9]([A-Z][A-Z0-9]{2}[0-9]){1,2}`), or MGI identifiers (`MGI:\d+`).
2. **Authoritative Symbol Resolution**: Direct query against HGNC (for Human, taxon 9606) or NCBI Gene (for model organisms).
3. **Alias & Historical Symbol Traversal**: Disambiguates synonyms against official database alias registries.
4. **Contextual Pruning**: Applies supplied clues (e.g. chromosome, locus type) to filter candidate matches.
5. **Typo / Fuzzy Recovery**: Levenshtein-distance fallback for minor typos when strict matches fail.

### C. Persistent Caching Layer (`base.py`)
- SQLite key-value store with configurable Time-To-Live (TTL).
- Caches raw responses and parsed models to eliminate redundant network roundtrips and protect upstream public bioinformatics APIs from rate-limiting.

---

## 3. Scientific Integrity & Anti-Hallucination Design

BioContext enforces a strict **"Fail Safely over Guessing"** design pattern:

| Principle | Implementation |
| :--- | :--- |
| **Deterministic Provenance** | Every resolution result links directly to authoritative accession numbers (`hgnc_id`, `entrez_id`, `ensembl_gene_id`, `uniprot_ids`). No synthetic or randomized fallbacks are permitted. |
| **Explicit Confidence Scores** | Outputs include scores from `1.0` (exact approved symbol) to `0.85` (alias), `0.60` (fuzzy), or `0.0` (unresolved). |
| **Transparent Audit Trails** | Every result object includes `match_reasons` detailing the exact rule, source authority, and evaluation details. |
| **Safe Failure State** | Unresolvable queries return `match_status: "unresolved"` rather than making speculative guesses. |
169 changes: 169 additions & 0 deletions docs/QUICKSTART.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,169 @@
# BioContext Quickstart Guide

This guide walks you through setting up and using **BioContext** across different environments:
1. **PyPI Package & Python SDK**
2. **AI Code Editors & Assistants (Cursor, Antigravity IDE, Claude Desktop, Claude Code, Goose)**
3. **CLI & High-Throughput Batch Workflows**
4. **Docker Container Microservice**

---

## 1. Installation via PyPI

BioContext is distributed on PyPI and can be installed with any standard Python package manager.

### Standard `pip` Installation
```bash
pip install biocontext
```

### Modern `uv` Tool / Virtualenv (Recommended)
```bash
# Install globally into your PATH as a CLI tool:
uv tool install biocontext

# Or add to your active project environment:
uv add biocontext
```

---

## 2. Integrating with AI Assistants & IDEs (MCP)

BioContext natively implements the **Model Context Protocol (MCP)** over `stdio`. Once configured, your AI assistant automatically gains access to real-time biological entity resolution, functional Gene Ontology enrichment, and Reactome pathway lookups.

### A. Cursor IDE
1. Open Cursor **Settings** (`Cmd + ,` or `Ctrl + ,`).
2. Navigate to **Features** $\rightarrow$ **MCP Servers** $\rightarrow$ Click **Add New MCP Server**.
3. Fill in:
- **Name**: `biocontext`
- **Type**: `command`
- **Command**: `uvx biocontext serve`
4. Or configure directly in `.cursor/mcp.json`:
```json
{
"mcpServers": {
"biocontext": {
"command": "uvx",
"args": ["biocontext", "serve"]
}
}
}
```

### B. Google Antigravity IDE
In Antigravity IDE, configure MCP servers in your workspace config (`.agents/mcp_config.json` or `~/.gemini/config/mcp_config.json`):
```json
{
"mcpServers": {
"biocontext": {
"command": "uvx",
"args": ["biocontext", "serve"]
}
}
}
```

### C. Claude Desktop
Add BioContext to your `claude_desktop_config.json`:
- **macOS**: `~/Library/Application Support/Claude/claude_desktop_config.json`
- **Windows**: `%APPDATA%\Claude\claude_desktop_config.json`
- **Linux**: `~/.config/Claude/claude_desktop_config.json`

```json
{
"mcpServers": {
"biocontext": {
"command": "uvx",
"args": ["biocontext", "serve"]
}
}
}
```

### D. Claude Code (CLI)
When running Anthropic's `claude` CLI in your terminal:
```bash
claude mcp add biocontext uvx biocontext serve
```

### E. Codex / Goose / Custom Agentic Frameworks
For any MCP client supporting JSON-RPC over `stdio`, execute:
```bash
# When installed in active python env:
biocontext serve

# Or zero-install via uvx:
uvx biocontext serve
```

---

## 3. Python SDK Usage

You can embed BioContext directly into data science scripts, pandas workflows, or Jupyter Notebooks:

```python
import asyncio
from biocontext.resolver import EntityResolver

async def main():
resolver = EntityResolver()

# 1. Resolve an ambiguous clinical alias
result = await resolver.resolve("HER2", taxon_id=9606)
print(f"Query: {result.query}")
print(f"Approved Symbol: {result.resolved_entity.symbol}") # ERBB2
print(f"HGNC ID: {result.resolved_entity.hgnc_id}") # HGNC:3430
print(f"Confidence: {result.confidence_score}") # 0.85
print(f"Match Status: {result.match_status}") # alias

# 2. Query Gene Ontology annotations
annotations = await resolver.annotate_gene("TP53", limit=5)
for ann in annotations:
print(f"[{ann.aspect}] {ann.term_id}: {ann.term_name} (Evidence: {ann.evidence_code})")

# 3. Query Reactome pathways
pathways = await resolver.get_pathways("TP53", limit=3)
for p in pathways:
print(f"Pathway: {p.st_id} - {p.name}")

if __name__ == "__main__":
asyncio.run(main())
```

---

## 4. High-Throughput Batch Processing via CLI

BioContext includes a concurrent batch engine bounded by `asyncio.Semaphore` with automatic rate-limiting compliance.

### Direct Command Line Arguments:
```bash
biocontext batch TP53 EGFR BRCA1 KRAS BRAF
```

### Processing Large CSV / TSV Gene Lists:
Given `input_genes.csv`:
```csv
gene_query,patient_id
HER2,PAT-001
p53,PAT-002
CD20,PAT-003
```

Execute:
```bash
biocontext batch --file input_genes.csv --output resolved_genes.csv --concurrency 8
```

---

## 5. Docker Microservice Execution

Run BioContext as an isolated container with persistent SQLite caching:

```bash
# Pull and run container with mounted persistent volume
docker run -i --rm -v biocontext_cache:/data ghcr.io/core-lab-research/biocontext:latest
```
Loading
Loading