Skip to content

About

CLI middleware for Claude API token reduction: prune stale turns, compress history via Haiku, semantic cache, relevance-ranked file injection

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

ContextSlim

CLI middleware for reducing Claude API token costs. Four features, composable by flag.

Features

Feature Flag What it does
Prune --no-prune to skip Removes stale conversation turns below a cosine-similarity threshold
Compress --compress to enable Compresses old history into a factual summary via Claude Haiku
Cache --no-cache to skip Returns cached responses for near-identical prompts (no API call)
Inject inject command Ranks file chunks by relevance, injects only the top-K

Install

pip install -r requirements.txt
cp .env.example .env          # add your ANTHROPIC_API_KEY

Usage

# Prune stale turns and show savings (no API call)
python main.py slim convo.json --query "fix the auth bug" --dry-run

# Prune + write slimmed output
python main.py slim convo.json -q "AC-2 compliance" -o slim.json

# Prune + compress history (makes a Haiku API call)
python main.py slim convo.json --compress -o slim.json

# Rank file chunks by relevance
python main.py inject -q "AC-2 evidence" policy.md controls.md README.md

# Output a formatted context block ready to paste into a prompt
python main.py inject -q "AC-2 evidence" *.md --context-block

# Semantic cache stats
python main.py cache stats

# Run the self-contained demo (prune + inject + cache; no API key needed)
python main.py demo

How each feature works

Prune: System messages are never touched. The last --keep-recent N non-system turns are always kept. Older turns are embedded (hash-based bag-of-words, L2-normalized) and compared to the current query via cosine similarity. Any turn below --stale-threshold is removed.

Compress: The oldest turns (beyond keep_recent) are sent to Claude Haiku with a summarization prompt. The result replaces those turns as a single [CONTEXT SUMMARY: ...] message. Skipped in --dry-run mode (requires an API call).

Cache: Each stored prompt is embedded and persisted to ~/.contextslim/cache.json. On lookup, the incoming prompt is embedded and compared against all stored embeddings. A cosine similarity ≥ 0.92 returns the cached response and skips the API call entirely.

Inject: Files are split on paragraph boundaries. Each chunk is embedded and scored against the query. Only the top-K chunks are returned, keeping injected context focused.

Flags

slim [FILE|-]
  --query / -q TEXT         Current task description (similarity reference)
  --no-prune                Skip stale-turn pruning
  --compress                Compress old history via Claude Haiku
  --no-cache                Disable semantic cache lookup
  --dry-run                 Show savings without writing anything
  --output / -o FILE        Write result to FILE instead of stdout
  --stale-threshold FLOAT   Cosine similarity below this → stale  [default: 0.30]
  --keep-recent INT         Always keep last N non-system turns    [default: 3]

inject
  --query / -q TEXT         Current task or question (required)
  --top-k INT               Number of chunks to return             [default: 5]
  --context-block           Output formatted block ready for prompt injection

Tests

pytest tests/ -v

57 tests across 8 modules (embedder, token_counter, pruner, compressor, cache, injector, slim, CLI).

Design notes

  • No heavy deps: numpy for cosine similarity; no FAISS, no sentence-transformers, no ONNX. Embedder uses a hash-based bag-of-words approach — fast, deterministic, no model downloads.
  • Dry-run safety: --dry-run never writes files, never mutates the cache, never calls the API. Only the pruner runs (pure local computation).
  • Compression is opt-in: Default compress=False means no surprise API calls. The --compress flag is explicit.
  • Cache is human-readable: Stored as JSON in ~/.contextslim/cache.json. Embeddings are float lists; you can inspect, edit, or delete entries manually.

About

CLI middleware for Claude API token reduction: prune stale turns, compress history via Haiku, semantic cache, relevance-ranked file injection

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages