ChoiceBench is a lightweight framework for MCQ evaluation-method research on LLMs, with built-in support for answer-order bias analysis and mitigation methods.
-
Updated
Sep 27, 2026 - Python
ChoiceBench is a lightweight framework for MCQ evaluation-method research on LLMs, with built-in support for answer-order bias analysis and mitigation methods.
Proper scoring rules, reduces LLM overconfidence in multiple-choice QA.
Open LLM leaderboard featuring Xiaomi MiMo v2.5 & MiMo 100T head-to-head with GPT-5, Claude, Gemini, DeepSeek, Llama 4. ARC-AGI · SWE-Bench · MMLU-Pro · GPQA · HumanEval · BFCL.
A Descriptive Analysis of Capacity, Training Stage, and Performance in Open Weight LLMs
To associate your repository with the mmlu-pro topic, visit your repo's landing page and select "manage topics."