Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 

Repository files navigation

CORTEX

A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs

Paper Venue

This repository contains the documentation and prompt release for CORTEX (Clinically Organized Reasoning and sTructured EXplanation), a structured reasoning benchmark for trustworthy multimodal large language models working with 3D chest CT.

This release accompanies the paper “CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs” by Hashmat Shadab Malik, Anees Ur Rehman Hashmi, Numan Saeed, Muzammal Naseer, Salman Khan, and Christoph Lippert.

Abstract

Medical multimodal models can produce fluent answers to imaging questions while leaving unclear whether their conclusions are grounded in the scan. CORTEX addresses this problem for 3D chest CT by organizing each example into four radiologist-inspired stages: task understanding, visual observation, diagnostic reasoning, and answer synthesis. The benchmark pairs these structured traces with a clinician-designed, stage-level evaluation protocol covering reasoning quality rather than final-answer correctness alone. It spans open-ended VQA, closed-ended VQA, and report generation and is built on CT-RATE.

Reasoning trace format

Every generated trace is organized into the following four stages:

Stage Tag Description
1 <task> Restate the clinical question, study type, and available clinical context
2 <observation> Systematic visual assessment organized by relevant anatomical region
3 <reason> Diagnostic reasoning over answer options or candidate hypotheses
4 <answer> Concise rationale and final answer or report synthesis

The dataset files and CT volumes are not included in this repository. This repository contains only dataset documentation and the prompts used to generate and score the released responses.

Contents

dep/
├── README.md
├── dataset/
│   └── overview.md
└── prompts/
    ├── generation/
    │   ├── closedended.md
    │   ├── openended.md
    │   └── report.md
    ├── scoring/
    │   ├── closedended.md
    │   ├── openended.md
    │   └── report.md

The structure follows the paper’s public codebase conventions while keeping this repository intentionally lightweight: it contains the finalized prompts and documentation, but no model checkpoints, inference scripts, CT volumes, or full dataset files.

Dataset

See dataset/overview.md for the task breakdown, record schema, release counts, provenance, and access notes. The separately distributed release contains the JSON task files and references to CT volumes; neither the JSON data nor the volumes are committed here.

Prompts and evaluation

Use a prompt from prompts/generation/ to produce a structured response for the corresponding task. Use the matching prompt from prompts/scoring/ to evaluate that response. In the paper’s terminology, the scoring prompts are the rubric-based validation prompts.

The prompts expect the model input to provide the CT image context, the relevant question or report task, and the required answer/reference fields for the selected task.

The generation prompts define a four-stage response format using <task>, <observation>, <reason>, and <answer> tags. The scoring prompts evaluate the resulting trace with task-specific rubrics on a 1–10 scale. The prompts are provided for transparency and reproducibility; this repository does not include an inference or scoring executable.

Reproducibility

To reproduce a prompt-driven run, select the task family matching the input record, load its generation prompt, and provide the image context and task fields expected by that prompt. For quality assessment, pass the generated response and the corresponding reference fields to the matching scoring prompt. Exact dataset access, model configurations, filtering rules, and generation details should be recorded alongside any derived release.

Dataset access

The dataset is hosted on Hugging Face. The GitHub repository contains documentation and prompts only; dataset JSON files and CT volumes are not committed here.

Download the complete Hugging Face release with:

from huggingface_hub import snapshot_download

local_path = snapshot_download(
    repo_id="aneesurhashmi/cortex",
    repo_type="dataset",
    local_dir="./cortex",
)

print(f"Downloaded to: {local_path}")

See the Hugging Face dataset card for the data layout, loading examples, provenance, and access terms.

Intended use and limitations

CORTEX is intended for research on multimodal clinical question answering, radiology report generation, and explicit diagnostic reasoning. It is not a clinical decision-support system. Model outputs are research artifacts and must not be used as a substitute for qualified clinical interpretation or patient care. Users should independently review the source dataset’s data-use terms, privacy documentation, and any applicable institutional requirements before accessing or redistributing the data.

Provenance

The release artifacts summarized here were assembled from CT-RATE-derived chest CT question-answer tasks. The exact source identifiers and image filenames remain in the separately distributed dataset files.

Citation

If you use CORTEX, please cite the accompanying paper:

@article{malik2026cortex,
  title   = {CORTEX: A Structured Reasoning Benchmark for Trustworthy 3D Chest CT MLLMs},
  author  = {Malik, Hashmat Shadab and Hashmi, Anees Ur Rehman and Saeed, Numan and Naseer, Muzammal and Khan, Salman and Lippert, Christoph},
  journal = {arXiv preprint arXiv:2606.27264},
  year    = {2026}
}

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors