Specification for a reproducible, provenance-bound multi-lane LLM benchmark suite.
-
Updated
Sep 27, 2026 - Python
Specification for a reproducible, provenance-bound multi-lane LLM benchmark suite.
Prompt-only boundary prediction for IFEval-style instruction-checker pass/fail behavior.
Reproducing a published IFEval score on one shared GPU - three arms, a pre-registration chain, and a paired test that killed my own conclusion
Code, per-item results and figures for a study dissociating prompt quality from response compliance in automated prompt-engineering assessment. Four-agent MATLAB evaluator on a locally hosted Qwen 2.5-7B judge, over 498 prompts from IFEval, LMSYS-Chat-1M and WildChat.
Task-dependent benchmark gap between /v1/completions and /v1/chat/completions on instruction-tuned LLMs -- two case studies on Qwen3.x GGUFs, reproduction recipe, and probes.
Benchmark the LM Studio models you already have - speed, answer quality and GPU thermals at several concurrency levels. Windows app, results stay local.
LoRA-based parameter-efficient fine-tuning for TinyLlama instruction following and ALBERT IMDb sentiment classification.
Parameter-efficient instruction tuning of TinyLlama-1.1B with LoRA, evaluated on IFEval.
To associate your repository with the ifeval topic, visit your repo's landing page and select "manage topics."