Specification for a reproducible, provenance-bound multi-lane LLM benchmark suite.
-
Updated
Sep 26, 2026 - Python
Specification for a reproducible, provenance-bound multi-lane LLM benchmark suite.
What 200 steps of fully simulated multi-turn tool-use RL do to a 4B policy: every 10th checkpoint scored on BFCL v4, with the pipeline that produced the measurement.
Public benchmark harnesses and reproducible evaluations from PixelSpaceAI
Pre-generation tool-call gating via linear probes on LLM hidden states. F1 ≈ 0.91–0.94 on BFCL v4, 14–22× faster than full generation. Cross-architecture transfer across Llama / Qwen / Phi / Mistral (3B–7B) with ≥96% retention.
A bf16 LoRA fine-tune of Qwen 3.5 4B for function calling on xLAM. v1.0 ships below the BFCL gate with full per-category failure analysis.
Canonical IR, schema validation, and deterministic delivery for function calls.
Independent audit of a fine-tuned LLM tool-calling PoC — BFCL regression decomposition, inference stack risk assessment, and production recommendation for a FinTech client. Qwen-2.5, LoRA, SGLang, H100.
Task-conditioned tail reliability for tool-using agents under equivalent interfaces
The first Turkish-native tool-calling benchmark: BFCL-style AST scoring, Turkish difficulty layer, HF leaderboard
Open LLM leaderboard featuring Xiaomi MiMo v2.5 & MiMo 100T head-to-head with GPT-5, Claude, Gemini, DeepSeek, Llama 4. ARC-AGI · SWE-Bench · MMLU-Pro · GPQA · HumanEval · BFCL.
Paired metamorphic evaluation of tool-calling robustness under realistic user phrasing, built on BFCL.
Single-GPU trajectory SFT for Qwen3-1.7B: +30 pp in-domain BFCL task success, with open data, LoRA weights and reproducible evaluation.
QLoRA fine-tune of Qwen2.5-1.5B for tool calling on one 8 GiB GPU: +8.5% on BFCL call categories, -55 points on the ability to decline. Refusal data recovers ~49 of them, reproduced across 3 seeds.
面向工具调用的大模型后训练与泛化评测系统。 End-to-end tool-calling LLM post-training pipeline with data auditing, QLoRA SFT, grounded on-policy DPO, strict evaluation and BFCL.
OpenEuroLLM snapshot of the Berkeley Function Calling Leaderboard evaluation harness and OLMo evaluation orchestration.
A compact, model-first notation for LLM tool definitions. Re-encodes JSON Schema with a median ~30% input-token reduction and behavior-preserving fallback. Includes the specification, a reference converter, a deployment protocol, and the full evaluation.
Tool-call data and evaluation lab: context isolation, strict JSON scoring, grouped splits, CPU contract evidence, and optional QLoRA training.
To associate your repository with the bfcl topic, visit your repo's landing page and select "manage topics."