This repository contains the documentation and prompt release for CORTEX (Clinically Organized Reasoning and sTructured EXplanation), a structured reasoning benchmark for trustworthy multimodal large language models working with 3D chest CT.
This release accompanies the paper “CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs” by Hashmat Shadab Malik, Anees Ur Rehman Hashmi, Numan Saeed, Muzammal Naseer, Salman Khan, and Christoph Lippert.
- Paper: arXiv:2606.27264 (PDF)
- Venue: SAFER 2026, MICCAI 2026.
Medical multimodal models can produce fluent answers to imaging questions while leaving unclear whether their conclusions are grounded in the scan. CORTEX addresses this problem for 3D chest CT by organizing each example into four radiologist-inspired stages: task understanding, visual observation, diagnostic reasoning, and answer synthesis. The benchmark pairs these structured traces with a clinician-designed, stage-level evaluation protocol covering reasoning quality rather than final-answer correctness alone. It spans open-ended VQA, closed-ended VQA, and report generation and is built on CT-RATE.
Every generated trace is organized into the following four stages:
| Stage | Tag | Description |
|---|---|---|
| 1 | <task> |
Restate the clinical question, study type, and available clinical context |
| 2 | <observation> |
Systematic visual assessment organized by relevant anatomical region |
| 3 | <reason> |
Diagnostic reasoning over answer options or candidate hypotheses |
| 4 | <answer> |
Concise rationale and final answer or report synthesis |
The dataset files and CT volumes are not included in this repository. This repository contains only dataset documentation and the prompts used to generate and score the released responses.
dep/
├── README.md
├── dataset/
│ └── overview.md
└── prompts/
├── generation/
│ ├── closedended.md
│ ├── openended.md
│ └── report.md
├── scoring/
│ ├── closedended.md
│ ├── openended.md
│ └── report.md
The structure follows the paper’s public codebase conventions while keeping this repository intentionally lightweight: it contains the finalized prompts and documentation, but no model checkpoints, inference scripts, CT volumes, or full dataset files.
See dataset/overview.md for the task breakdown, record schema, release counts, provenance, and access notes. The separately distributed release contains the JSON task files and references to CT volumes; neither the JSON data nor the volumes are committed here.
Use a prompt from prompts/generation/ to produce a structured response for the corresponding task. Use the matching prompt from prompts/scoring/ to evaluate that response. In the paper’s terminology, the scoring prompts are the rubric-based validation prompts.
The prompts expect the model input to provide the CT image context, the relevant question or report task, and the required answer/reference fields for the selected task.
The generation prompts define a four-stage response format using <task>, <observation>, <reason>, and <answer> tags. The scoring prompts evaluate the resulting trace with task-specific rubrics on a 1–10 scale. The prompts are provided for transparency and reproducibility; this repository does not include an inference or scoring executable.
To reproduce a prompt-driven run, select the task family matching the input record, load its generation prompt, and provide the image context and task fields expected by that prompt. For quality assessment, pass the generated response and the corresponding reference fields to the matching scoring prompt. Exact dataset access, model configurations, filtering rules, and generation details should be recorded alongside any derived release.
The dataset is hosted on Hugging Face. The GitHub repository contains documentation and prompts only; dataset JSON files and CT volumes are not committed here.
Download the complete Hugging Face release with:
from huggingface_hub import snapshot_download
local_path = snapshot_download(
repo_id="aneesurhashmi/cortex",
repo_type="dataset",
local_dir="./cortex",
)
print(f"Downloaded to: {local_path}")See the Hugging Face dataset card for the data layout, loading examples, provenance, and access terms.
CORTEX is intended for research on multimodal clinical question answering, radiology report generation, and explicit diagnostic reasoning. It is not a clinical decision-support system. Model outputs are research artifacts and must not be used as a substitute for qualified clinical interpretation or patient care. Users should independently review the source dataset’s data-use terms, privacy documentation, and any applicable institutional requirements before accessing or redistributing the data.
The release artifacts summarized here were assembled from CT-RATE-derived chest CT question-answer tasks. The exact source identifiers and image filenames remain in the separately distributed dataset files.
If you use CORTEX, please cite the accompanying paper:
@article{malik2026cortex,
title = {CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs},
author = {Malik, Hashmat Shadab and Hashmi, Anees Ur Rehman and Saeed, Numan and Naseer, Muzammal and Khan, Salman and Lippert, Christoph},
journal = {arXiv preprint arXiv:2606.27264},
year = {2026}
}