Skip to content

Repository files navigation

ToolRoute

ToolRoute is a small, deterministic tool-context router for OpenAI-compatible function calling. It retrieves a compact subset from a large tool catalog before the request reaches the model.

Why ToolRoute?

On a fixed 100-example BFCL stress set, exposing all 128 semantically similar tools reduced Qwen3-8B-AWQ accuracy to 83%. BM25 recovered it to 91%, and the recommended coverage-preserving BM25+MMR route reached 94%, close to the 95% gold-only upper bound.

Architecture

Tool catalog → BM25 → Top-32 pool → optional MMR → fixed K=16 / adaptive K → LLM

Recommended: BM25 → Top-32 → MMR (λ=0.25) → K=16.

ToolRoute Framework

ToolRoute first retrieves a coverage-preserving candidate pool with BM25, then applies MMR reranking to reduce redundant or competing tool schemas before exposing the final tool set to the LLM.

Adaptive routing remains EXPERIMENTAL. Its calibration-frozen policy scored 81% on Hard-128 and 85% on Random-128, so it did not qualify for promotion. The release audit found that Adaptive K=16 uses a different routing/construction path from BM25 K=16: all 100 Hard-128 selected tool sets differ. The token reduction is therefore an implementation-level construction difference, not a claim that the same BM25 schemas have a cheaper representation.

Installation

python -m pip install -e .

Quick start

from toolroute import ToolRouter

router = ToolRouter.from_config("configs/router/bm25_mmr.yaml")
result = router.route("Find the weather in Paris", tools)
print(result.selected_tool_names)

The bundled demo catalog can be used directly:

python examples/basic_route.py
python examples/mmr_route.py
python examples/adaptive_route.py

Python API

ToolRouter() uses the recommended bm25_mmr defaults. For explicit, versioned behavior, load one of the YAML presets with ToolRouter.from_config(...). route(...) returns the selected Tool objects, their names, schema-token accounting, final K, and routing metadata.

CLI

python -m toolroute.route --query "Find the weather in Paris" \
  --tools examples/tools.json --config configs/router/bm25_mmr.yaml

The JSON response includes selected tools, names, schema-token estimate, final K, routing metadata, and ranking information.

Routing modes

  • bm25: deterministic lexical retrieval.
  • bm25_mmr: RECOMMENDED; Top-32 candidate pool, MMR λ=0.25, K=16.
  • adaptive: EXPERIMENTAL efficiency mode; not recommended by frozen tests.

Benchmark

The frozen evaluation uses paired 100-example Hard-128 and Random-128 BFCL conditions. Official BFCL scoring is applied independently at the condition + category level. Details and pinned revisions are in the Phase 1B report and reproducibility guide.

Final results

Method Hard-128 Random-128 Role
Full 83 86 stress baseline
BM25 K16 91 92 retrieval baseline
BM25+MMR K16 94 not evaluated recommended
Adaptive 81 85 experimental
Oracle Gold 95 95 upper bound

Random-128 RRF K16 reached 94%, but RRF and MMR are different methods; the project does not relabel that historical result as MMR.

Accuracy–context trade-off

Hard-128 Full used 11,737 mean schema tokens. BM25 K16 used 2,121 (81.9% reduction). Adaptive used 1,095 but lost accuracy. The recommended MMR K16 prioritizes the validated 94% accuracy result over heuristic adaptive reduction.

Hard-128 accuracy versus schema tokens

Reproducibility

See docs/REPRODUCIBILITY.md and the frozen artifacts under artifacts/phase0, artifacts/phase1a, artifacts/phase1b0, and artifacts/phase1b.

Project structure

Core routing is in src/toolroute; benchmark/evaluation scripts are in scripts; examples and configs are self-contained. The runtime router has no BFCL, vLLM, Qwen, AutoDL, or server dependency.

Limitations

The benchmark is a 100-example controlled pilot, not a leaderboard submission. The released dependency-light MMR fallback uses lexical similarity; the 94% benchmark MMR used the pinned BGE-base-en-v1.5 representation. Adaptive results were frozen without test-time retuning.

Upstream attribution

BFCL/Gorilla supplies the benchmark and official AST evaluator; ToolRet supplies the external tool corpus; BGE-base-en-v1.5 supplies construction/evaluation embeddings. These are upstream works, not ToolRoute contributions. Exact pinned revisions are recorded in docs/UPSTREAM.md; see THIRD_PARTY_NOTICES.md.

License

ToolRoute is licensed under the Apache License 2.0.

About

Scalable tool routing for LLM function calling with BM25 and MMR.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Contributors

Languages