CLI middleware for reducing Claude API token costs. Four features, composable by flag.
| Feature | Flag | What it does |
|---|---|---|
| Prune | --no-prune to skip |
Removes stale conversation turns below a cosine-similarity threshold |
| Compress | --compress to enable |
Compresses old history into a factual summary via Claude Haiku |
| Cache | --no-cache to skip |
Returns cached responses for near-identical prompts (no API call) |
| Inject | inject command |
Ranks file chunks by relevance, injects only the top-K |
pip install -r requirements.txt
cp .env.example .env # add your ANTHROPIC_API_KEY# Prune stale turns and show savings (no API call)
python main.py slim convo.json --query "fix the auth bug" --dry-run
# Prune + write slimmed output
python main.py slim convo.json -q "AC-2 compliance" -o slim.json
# Prune + compress history (makes a Haiku API call)
python main.py slim convo.json --compress -o slim.json
# Rank file chunks by relevance
python main.py inject -q "AC-2 evidence" policy.md controls.md README.md
# Output a formatted context block ready to paste into a prompt
python main.py inject -q "AC-2 evidence" *.md --context-block
# Semantic cache stats
python main.py cache stats
# Run the self-contained demo (prune + inject + cache; no API key needed)
python main.py demoPrune: System messages are never touched. The last --keep-recent N non-system turns are always kept. Older turns are embedded (hash-based bag-of-words, L2-normalized) and compared to the current query via cosine similarity. Any turn below --stale-threshold is removed.
Compress: The oldest turns (beyond keep_recent) are sent to Claude Haiku with a summarization prompt. The result replaces those turns as a single [CONTEXT SUMMARY: ...] message. Skipped in --dry-run mode (requires an API call).
Cache: Each stored prompt is embedded and persisted to ~/.contextslim/cache.json. On lookup, the incoming prompt is embedded and compared against all stored embeddings. A cosine similarity ≥ 0.92 returns the cached response and skips the API call entirely.
Inject: Files are split on paragraph boundaries. Each chunk is embedded and scored against the query. Only the top-K chunks are returned, keeping injected context focused.
slim [FILE|-]
--query / -q TEXT Current task description (similarity reference)
--no-prune Skip stale-turn pruning
--compress Compress old history via Claude Haiku
--no-cache Disable semantic cache lookup
--dry-run Show savings without writing anything
--output / -o FILE Write result to FILE instead of stdout
--stale-threshold FLOAT Cosine similarity below this → stale [default: 0.30]
--keep-recent INT Always keep last N non-system turns [default: 3]
inject
--query / -q TEXT Current task or question (required)
--top-k INT Number of chunks to return [default: 5]
--context-block Output formatted block ready for prompt injection
pytest tests/ -v57 tests across 8 modules (embedder, token_counter, pruner, compressor, cache, injector, slim, CLI).
- No heavy deps: numpy for cosine similarity; no FAISS, no sentence-transformers, no ONNX. Embedder uses a hash-based bag-of-words approach — fast, deterministic, no model downloads.
- Dry-run safety:
--dry-runnever writes files, never mutates the cache, never calls the API. Only the pruner runs (pure local computation). - Compression is opt-in: Default
compress=Falsemeans no surprise API calls. The--compressflag is explicit. - Cache is human-readable: Stored as JSON in
~/.contextslim/cache.json. Embeddings are float lists; you can inspect, edit, or delete entries manually.