Test-Time Calibration Learning (TTCL), a label-free framework that jointly adapts reasoning accuracy and verbalized confidence directly on unlabeled target-task data. Specifically, TTCL derives self-supervision signals for both correctness and calibration from multiple model-generated responses, enabling calibration learning at test time without ground-truth labels. Theoretical analysis further establishes TTCL as a bounded surrogate for the ideal calibration objective. Extensive experiments on mathematical reasoning and factual question answering demonstrate that TTCL consistently improves both accuracy and calibration across diverse models and tasks. On base models, TTCL achieves an average relative accuracy improvement of
Overview of our work: TTCL enables label-free calibration learning directly on unlabeled target-task data, consistently improving reasoning accuracy while reducing calibration error across models and tasks, with confidence progressively aligning with accuracy during test-time learning.
Illustration of TTCL method: Given an unlabeled question, TTCL samples multiple rollouts to construct label-free supervision: the most supported answer serves as the pseudo label, while the support of each answer provides its confidence target. A target cache stabilizes these signals and is updated via EMA across epochs. Correctness and calibration rewards are optimized separately, with the former updating reasoning and answer tokens and the latter updating only confidence tokens.
TTCL effectively aligns confidence with empirical accuracy. Left figure shows the evolution of accuracy and mean confidence during TTCL training. The initial models are generally over-confident, with mean confidence exceeding empirical accuracy. As training proceeds, TTCL narrows this gap while improving accuracy, bringing the two increasingly into alignment. This training behavior empirically supports our theoretical interpretation of TTCL as a surrogate for the ideal calibration objective.
TTCL effectively aligns confidence with empirical accuracy. Beyond the scalar ECE metric, Fig.~\ref{Fig:calibration_bar} visualizes confidence against empirical accuracy. Accuracy across confidence bins becomes closer to the ideal calibration diagonal after TTCL, indicating better alignment between predicted confidence and correctness likelihood. This pattern holds across mathematical reasoning and factual QA, as well as base and pre-calibrated models, showing that reliable confidence estimates of TTCL.
Create a Python environment with a PyTorch/CUDA version appropriate for your hardware, activate it, and install the project dependencies with:
USE_MEGATRON=0 USE_SGLANG=0 bash scripts/install_vllm_sglang_mcore.shAs we use vllm as inference engine and fsdp as training backbone, we do not need to install sglang and megatron.
The training and evaluation scripts require a local Hugging Face model checkpoint.
Datasets are located in data/. Set every model path in the launch scripts to
a local checkpoint directory before running an experiment.
The TTCL training launcher is:
bash scripts/majority_ttcl/run.shThis launcher runs the listed experiments serially. Configure LLM_PATH,
MODEL_NAME, and the experiment list in scripts/majority_ttcl/run.sh before
launching. Training hyperparameters are defined in
scripts/majority_ttcl/run_majority_ttcl.sh; per-dataset settings are in the
corresponding run_majority_ttcl_*.sh files. Configure experiment logging
locally if it is needed.
Set the model paths in the following scripts, then run:
# Mathematical benchmarks
bash eval/run_evaluate_math.sh
# Factual QA benchmarks
bash eval/run_evaluate_factqa.shFor the Lighteval benchmark suite, edit the model list in
eval/lighteval/run_suite.sh and run:
bash eval/lighteval/run_suite.shWe release all checkpoints trained by us on huggingface:
| Model Name | Model Size | Method | Hugging Face Link |
|---|---|---|---|
| TTCL-MATH500-Qwen3-4B | 4B | TTCL | View Model |
| TTCL-AMC23-Qwen3-4B | 4B | TTCL | View Model |
| TTCL-AIME24-Qwen3-4B | 4B | TTCL | View Model |
| TTCL-AIME25-Qwen3-4B | 4B | TTCL | View Model |
| TTCL-SimpleQA-Qwen3-4B | 4B | TTCL | View Model |
| Model Name | Model Size | Method | Hugging Face Link |
|---|---|---|---|
| TTCL-MATH500-Qwen3-8B | 8B | TTCL | View Model |
| TTCL-AMC23-Qwen3-8B | 8B | TTCL | View Model |
| TTCL-AIME24-Qwen3-8B | 8B | TTCL | View Model |
| TTCL-AIME25-Qwen3-8B | 8B | TTCL | View Model |
| TTCL-SimpleQA-Qwen3-8B | 8B | TTCL | View Model |
| Model Name | Model Size | Method | Hugging Face Link |
|---|---|---|---|
| TTCL-MATH500-Llama-3.1-8B-Instruct | 8B | TTCL | View Model |
| TTCL-AMC23-Llama-3.1-8B-Instruct | 8B | TTCL | View Model |
| TTCL-AIME24-Llama-3.1-8B-Instruct | 8B | TTCL | View Model |
| TTCL-AIME25-Llama-3.1-8B-Instruct | 8B | TTCL | View Model |
| TTCL-SimpleQA-Llama-3.1-8B-Instruct | 8B | TTCL | View Model |
| Model Name | Model Size | Method | Hugging Face Link |
|---|---|---|---|
| RLCR-calibrated Qwen3-4B on DAPO-Math-14K | 4B | RLCR | View Model |
| RLCR-calibrated Qwen3-8B on DAPO-Math-14K | 8B | RLCR | View Model |
| RLCR-calibrated Qwen3-4B on FactQA-10K | 4B | TTCL | View Model |
| Model Name | Model Size | Method | Hugging Face Link |
|---|---|---|---|
| TTCL-MATH500-RLCR-Qwen3-4B-DAPO-Math14K | 4B | TTCL | View Model |
| TTCL-AMC23-RLCR-Qwen3-4B-DAPO-Math14K | 4B | TTCL | View Model |
| TTCL-AIME24-RLCR-Qwen3-4B-DAPO-Math14K | 4B | TTCL | View Model |
| TTCL-AIME25-RLCR-Qwen3-4B-DAPO-Math14K | 4B | TTCL | View Model |
| TTCL-AIME26-RLCR-Qwen3-4B-DAPO-Math14K | 4B | TTCL | View Model |
| TTCL-SimpleQA-RLCR-Qwen3-4B-DAPO-Math14K | 4B | TTCL | View Model |
| TTCL-ChineseSimpleQA-RLCR-Qwen3-4B-DAPO-Math14K | 4B | TTCL | View Model |
| TTCL-NQ-RLCR-Qwen3-4B-DAPO-Math14K | 4B | TTCL | View Model |
| TTCL-HotpotQA-RLCR-Qwen3-4B-DAPO-Math14K | 4B | TTCL | View Model |
| TTCL-TriviaQA-RLCR-Qwen3-4B-DAPO-Math14K | 4B | TTCL | View Model |
| Model Name | Model Size | Method | Hugging Face Link |
|---|---|---|---|
| TTCL-MATH500-RLCR-Qwen3-8B-DAPO-Math14K | 8B | TTCL | View Model |
| TTCL-AMC23-RLCR-Qwen3-8B-DAPO-Math14K | 8B | TTCL | View Model |
| TTCL-AIME24-RLCR-Qwen3-8B-DAPO-Math14K | 8B | TTCL | View Model |
| TTCL-AIME25-RLCR-Qwen3-8B-DAPO-Math14K | 8B | TTCL | View Model |
| TTCL-AIME26-RLCR-Qwen3-8B-DAPO-Math14K | 8B | TTCL | View Model |
| TTCL-SimpleQA-RLCR-Qwen3-8B-DAPO-Math14K | 8B | TTCL | View Model |
| TTCL-ChineseSimpleQA-RLCR-Qwen3-8B-DAPO-Math14K | 8B | TTCL | View Model |
| TTCL-NQ-RLCR-Qwen3-8B-DAPO-Math14K | 8B | TTCL | View Model |
| TTCL-HotpotQA-RLCR-Qwen3-8B-DAPO-Math14K | 8B | TTCL | View Model |
| TTCL-TriviaQA-RLCR-Qwen3-8B-DAPO-Math14K | 8B | TTCL | View Model |
| Model Name | Model Size | Method | Hugging Face Link |
|---|---|---|---|
| TTCL-MATH500-RLCR-Qwen3-4B-FactQA10K | 4B | TTCL | View Model |
| TTCL-AMC23-RLCR-Qwen3-4B-FactQA10K | 4B | TTCL | View Model |
| TTCL-AIME24-RLCR-Qwen3-4B-FactQA10K | 4B | TTCL | View Model |
| TTCL-AIME25-RLCR-Qwen3-4B-FactQA10K | 4B | TTCL | View Model |
| TTCL-AIME26-RLCR-Qwen3-4B-FactQA10K | 4B | TTCL | View Model |
If you use our datasets or models, please cite our paper!




