BrainVLM

A vision-language foundation model for comprehensive preoperative brain tumor diagnosis

Yinong Wang*, Jianwen Chen*, Zhou Chen*, Shuwen Kuang*, Haoning Jiang, Yanzhao Shi, Huichun Yuan, Yan-ran (Joyce) Wang, Bing Wang, Lei Wu, Bin Tang, Li Meng, Baihua Luo, Bin Zhou, Wei Ding, Weiming Zhong, Wei Hou, Yuanbing Chen, Zhiping Wan, Wei Wang, Zhenkun Xiao, Wenwu Wan, Allen He, Yuyin Zhou, Longbo Zhang, Feifei Wang, Zhixiong Liu, Michael Iv, Xuan Gong, Liangqiong Qu

The University of Hong Kong · Xiangya School of Medicine, Central South University · etc.

Accepted for publication in The Lancet Digital Health

Abstract

Background. Preoperative diagnosis of brain tumors from MRI remains challenging because imaging features overlap across tumor types and specialist interpretation requires substantial experience. BrainVLM was developed to support comprehensive and reliable brain tumor diagnosis from preoperative multimodal data.

Methods. BrainVLM classifies all 12 major brain tumor categories defined by WHO CNS5 and provides each diagnosis with a confidence score and a generated radiology report. The model was trained on multimodal data from 40,043 individuals and evaluated on 5,211 pathologically confirmed cases, including 3,877 patients from Xiangya Hospital and 1,334 patients from 11 independent hospitals. Clinical utility was further assessed in a prospective study of 1,009 patients and a blinded multi-reader study involving 12 neuroradiologists and 248 cases.

Findings. BrainVLM achieved a macro-AUC of 0.85 and F1 score of 0.82 in the primary test cohort, and a macro-AUC of 0.80 and F1 score of 0.75 in external validation. Its prospective performance was comparable to or better than neuroradiologists across primary and external cohorts. With BrainVLM assistance, neuroradiologists improved their mean F1 score by 27.6% and reduced diagnostic time by 34.7%.

Interpretation. BrainVLM combines comprehensive tumor classification, diagnostic confidence, and report generation within one framework. Retrospective, prospective, and multi-reader evaluations indicate its potential to support more accurate and efficient preoperative brain tumor assessment.

Project overview. We introduce BrainVLM from four perspectives: Diagnosis, covering comprehensive tumor classification and retrospective validation; Reports, presenting generated radiology reports and diagnosis-associated confidence; Clinical, summarizing prospective and multi-reader studies; and Data & Training, describing dataset construction, preprocessing, and model development.

12WHO CNS5 tumor classes
47,947individuals in BrainTumor48K
Multi-centerretrospective & prospective evaluation

01 · Input

Input

BrainVLM reads T1, T1c, T2, and FLAIR together as one case, while remaining usable when part of the MRI protocol is unavailable. It accepts raw or pre-processed MRI and can incorporate available metadata, allowing the same framework to accommodate variations in routine imaging protocols and data preparation.

View input details →
BrainVLM multi-sequence MRI input and case-level interpretation

02 · Diagnosis

Diagnosis

To our knowledge, BrainVLM is the first MRI-based AI model to provide comprehensive preoperative classification across all 12 major WHO CNS5 brain tumor categories, together with a broad range of their subtypes. Its coverage extends beyond the few common tumor classes typically studied by existing AI systems to include less prevalent and imaging-ambiguous entities.

BrainVLM coverage of major WHO CNS5 brain tumor categories

Diagnostic performance was evaluated in 3,877 held-out patients from Xiangya Hospital and 1,334 patients from 11 independent hospitals. This design examines both performance at the primary center and generalization across institutions with different scanners, populations, and clinical imaging practices.

BrainVLM retrospective and external diagnostic performance

03 · Reports

Reports

BrainVLM provides more than a diagnostic label. For each case, it also generates a radiology report describing relevant MRI findings and presents the confidence associated with the diagnosis. Together, these outputs give clinicians more context for reviewing the result and help identify uncertain cases that may require closer assessment.

View reports & confidence →
BrainVLM MRI findings, imaging report, and diagnostic confidence

04 · Clinical

Clinical

BrainVLM was evaluated beyond retrospective test sets through a prospective real-world study and a blinded multi-reader study. These experiments examine not only diagnostic performance, but also how the model may contribute when used alongside neuroradiologists in clinical review.

BrainVLM prospective study results

Prospective real-world results

The prospective study evaluated consecutive patients before definitive diagnosis, including both surgical cases confirmed by pathology and non-operative cases established through follow-up or multidisciplinary assessment. Primary and external cohorts were analyzed separately to reflect different clinical settings.

Primary F1 0.84
BrainVLM multi-reader clinical evaluation

AI-augmented neuroradiology

In a blinded crossover study, 12 neuroradiologists of different experience levels reviewed 248 cases with and without BrainVLM assistance. Access to the model output improved diagnostic performance across reader groups while reducing the time required for each assessment.

+27.6% F1 · −34.7% time

05 · Data & Training

Data & Training

Online sources. BrainTumor48K integrates 39 online sources, including BraTS23 and TCIA collections, ReMIND, UPENN-GBM, Radiopaedia, OpenNeuro, IXI, and Kaggle datasets such as Br35H, Figshare Brain Tumor Dataset, Brain Tumor Classification, and Brain Tumor MRI Images 44 Classes. PubMed Central and ImageCLEF contribute additional MRI–text pairs, while several OpenNeuro datasets provide healthy controls.

Hospital data. The institutional cohort was collected from Xiangya Hospital and 11 independent centers: Changde First People’s Hospital, Tongji Hospital, the Second Affiliated Hospital of the University of South China, Shenzhen Second People’s Hospital, the Third Xiangya Hospital, Jiangxi Provincial People’s Hospital, Chongqing Traditional Chinese Medicine Hospital, the First and Second Affiliated Hospitals of Nanchang University, Hunan Children’s Hospital, and the First Hospital of Lanzhou University. These data include de-identified T1, T1c, T2, and FLAIR MRI, demographics, radiology reports, and pathological diagnoses when available.

Processing. Structured repositories were converted into a common patient-level format linking MRI, diagnosis labels, metadata, and available reports. Web-derived figures were separated, filtered, and aligned with their captions. Institutional reports were restricted to the four MRI sequences, standardized through a clinician-validated bilingual terminology workflow, and reviewed by neuroradiologists before use. The resulting 2D and 3D data were then organized for progressive model training.

Pipelines for processing structured repositories, web resources, and private medical institutions
01 / PROCESS

Multi-source data processing

Separate workflows handle structured repositories, web-derived MRI–text pairs, and de-identified hospital data.

BrainTumor48K data curation and dataset organization
02 / CURATE

BrainTumor48K construction

The curated dataset contains 47,947 individuals, combining 2D MRI–text data with case-level 3D multimodal MRI.

BrainVLM progressive training and reliability training framework
03 / TRAIN

Progressive training

Training connects visual representation learning, comprehensive diagnosis, report generation, and diagnostic confidence.

BrainVLM

A Vision-Language Foundation Model for Precise and Comprehensive Brain Tumor Diagnosis from Preoperative Multimodal Data

The University of Hong Kong
Xiangya Hospital, Central South University
Stanford University
University of California, Santa Cruz
Tongji Hospital, Tongji University
Jiangxi Medical College, Nanchang University
...

*Indicates Equal Contribution

Corresponding Authors
Paper Code Hugging Face Logo Demo
HKU Logo
Xiangya Hospital Logo
Stanford University Logo
Jiangxi University Logo
Tongji University Logo
Nanchang University Logo
Lanzhou University Logo
Changde Hospital Logo
Hunan Ertong Hospital Logo
UCSC Logo
Shenzhen Hospital Logo
Nanhua2 Hospital Logo
Nanchang2 Hospital Logo
Xiangya3 Hospital Logo
Chongqing Hospital Logo

Overview of the BrainVLM Dataset and Model Architecture

BrainVLM Dataset and Architecture Overview

a, BrainTumor48K aggregates 47,947 individuals from 39 public repositories and 12 collaborating medical centers, including 33,149 patients with pathologically confirmed brain tumors and 14,798 healthy controls. Cases include 2D MRI slices or multi-sequence 3D MRI scans (T1-weighted [T1], T1 contrast-enhanced [T1c], T2-weighted [T2], and T2-FLAIR [T2f]), paired with demographics, radiology reports, and pathological diagnoses. The dataset spans all 12 major brain tumor types defined by WHO CNS5.

b, BrainVLM is a multimodal vision-language foundation model for preoperative brain tumor analysis. It combines multi-parametric MRI, demographic information, and radiology reports to produce a tumor diagnosis, calibrated confidence estimate, and radiology-style report. The framework was evaluated on primary and multi-center external cohorts, a prospective real-world study, and a blinded multi-reader study with neuroradiologists.

Abstract

Background: Preoperative diagnosis of brain tumors from multimodal MRI is challenging because different tumor types can share imaging features and specialist interpretation is time-intensive. Methods: We developed BrainVLM, a vision-language foundation model trained with MRI, demographics, and radiology reports to classify all 12 major WHO CNS5 tumor types, quantify diagnostic uncertainty, and generate radiology-style reports. BrainTumor48K includes 47,947 individuals collected from 39 public repositories and 12 collaborating medical centers. The model was evaluated on 5,211 pathologically confirmed cases, a prospective real-world cohort, and a blinded multi-reader study with 12 neuroradiologists. Findings: BrainVLM achieved a primary-test macro-AUC of 0.85 and F1 score of 0.82, and an external-test macro-AUC of 0.80 and F1 score of 0.75. AI assistance improved neuroradiologists’ mean F1 by 23.3% and reduced diagnostic time by 36.1%. The model also achieved AUCs of 0.95 and 0.88 for adult-type diffuse glioma molecular subtyping in primary and external cohorts. Interpretation: By combining prediction, calibrated confidence, and report generation, BrainVLM is designed to support transparent and efficient clinician review rather than replace clinical judgment.

BrainVLM at a glance

A specialized vision-language foundation model for preoperative brain tumor diagnosis, designed to pair accurate predictions with calibrated confidence and clinically readable reports.

12 WHO CNS5 major brain tumor types covered
47,947 Individuals in BrainTumor48K from 39 repositories and 12 medical centers
0.85 Primary-test macro-AUC across tumor categories
0.80 Multi-center external-test macro-AUC

From MRI to a clinical decision aid

BrainVLM combines complementary clinical signals and makes its output easier to review, communicate, and triage in real-world neuro-oncology workflows.

01

Multimodal input

Processes standard T1, T1c, T2, and T2-FLAIR MRI together with patient demographics and available radiology text.

02

Comprehensive classification

Targets the full spectrum of 12 major WHO CNS5 brain tumor categories, including low-prevalence and imaging-ambiguous tumors.

03

Confidence-aware output

Provides calibrated uncertainty estimates and can surface a supplementary top-two diagnosis when the leading confidence is below 75%.

04

Radiology-style report

Generates a clinical narrative describing imaging findings, supporting interpretation beyond an isolated tumor label.

Evidence across clinical settings

The paper evaluates BrainVLM on retrospective, multi-center external, prospective, and AI-augmented reader cohorts.

Diagnostic performance

Primary test · 3,877 patientsF1 0.82 · AUC 0.85
External test · 1,334 patients / 11 hospitalsF1 0.75 · AUC 0.80
Prospective real-world evaluation4% higher F1
Adult-type diffuse glioma subtypingAUC 0.95 / 0.88

The molecular-subgroup AUCs correspond to primary and external validation cohorts, respectively.

AI-augmented practice

Blinded multi-reader study248 cases · 12 readers
Mean F1 improvement with AI+23.3%
Diagnostic time reduction−36.1%
Confidence calibrationECE 0.036

BrainVLM is intended to support clinician review; lower-confidence cases should receive human assessment.

Research system for decision support, not a standalone medical device. Model performance and generalizability should be assessed in the intended clinical population before deployment.

BrainTumor48K construction and preprocessing pipeline

Research Results Overview

Result

Demo video

If you are unable to play the video, please try using a modern browser or contact the website administrator.