To our knowledge, BrainVLM is the first MRI-based AI model to provide comprehensive preoperative classification across all 12 major WHO CNS5 brain tumor categories, together with a broad range of their subtypes. Its coverage extends beyond the few common tumor classes typically studied by existing AI systems to include less prevalent and imaging-ambiguous entities.
Abstract
Background. Preoperative diagnosis of brain tumors from MRI remains challenging because imaging features overlap across tumor types and specialist interpretation requires substantial experience. BrainVLM was developed to support comprehensive and reliable brain tumor diagnosis from preoperative multimodal data.
Methods. BrainVLM classifies all 12 major brain tumor categories defined by WHO CNS5 and provides each diagnosis with a confidence score and a generated radiology report. The model was trained on multimodal data from 40,043 individuals and evaluated on 5,211 pathologically confirmed cases, including 3,877 patients from Xiangya Hospital and 1,334 patients from 11 independent hospitals. Clinical utility was further assessed in a prospective study of 1,009 patients and a blinded multi-reader study involving 12 neuroradiologists and 248 cases.
Findings. BrainVLM achieved a macro-AUC of 0.85 and F1 score of 0.82 in the primary test cohort, and a macro-AUC of 0.80 and F1 score of 0.75 in external validation. Its prospective performance was comparable to or better than neuroradiologists across primary and external cohorts. With BrainVLM assistance, neuroradiologists improved their mean F1 score by 27.6% and reduced diagnostic time by 34.7%.
Interpretation. BrainVLM combines comprehensive tumor classification, diagnostic confidence, and report generation within one framework. Retrospective, prospective, and multi-reader evaluations indicate its potential to support more accurate and efficient preoperative brain tumor assessment.
Project overview. We introduce BrainVLM from four perspectives: Diagnosis, covering comprehensive tumor classification and retrospective validation; Reports, presenting generated radiology reports and diagnosis-associated confidence; Clinical, summarizing prospective and multi-reader studies; and Data & Training, describing dataset construction, preprocessing, and model development.
01 · Input
Input
BrainVLM reads T1, T1c, T2, and FLAIR together as one case, while remaining usable when part of the MRI protocol is unavailable. It accepts raw or pre-processed MRI and can incorporate available metadata, allowing the same framework to accommodate variations in routine imaging protocols and data preparation.
View input details →02 · Diagnosis
Diagnosis
Diagnostic performance was evaluated in 3,877 held-out patients from Xiangya Hospital and 1,334 patients from 11 independent hospitals. This design examines both performance at the primary center and generalization across institutions with different scanners, populations, and clinical imaging practices.
03 · Reports
Reports
BrainVLM provides more than a diagnostic label. For each case, it also generates a radiology report describing relevant MRI findings and presents the confidence associated with the diagnosis. Together, these outputs give clinicians more context for reviewing the result and help identify uncertain cases that may require closer assessment.
View reports & confidence →04 · Clinical
Clinical
BrainVLM was evaluated beyond retrospective test sets through a prospective real-world study and a blinded multi-reader study. These experiments examine not only diagnostic performance, but also how the model may contribute when used alongside neuroradiologists in clinical review.
Prospective real-world results
The prospective study evaluated consecutive patients before definitive diagnosis, including both surgical cases confirmed by pathology and non-operative cases established through follow-up or multidisciplinary assessment. Primary and external cohorts were analyzed separately to reflect different clinical settings.
AI-augmented neuroradiology
In a blinded crossover study, 12 neuroradiologists of different experience levels reviewed 248 cases with and without BrainVLM assistance. Access to the model output improved diagnostic performance across reader groups while reducing the time required for each assessment.
05 · Data & Training
Data & Training
Online sources. BrainTumor48K integrates 39 online sources, including BraTS23 and TCIA collections, ReMIND, UPENN-GBM, Radiopaedia, OpenNeuro, IXI, and Kaggle datasets such as Br35H, Figshare Brain Tumor Dataset, Brain Tumor Classification, and Brain Tumor MRI Images 44 Classes. PubMed Central and ImageCLEF contribute additional MRI–text pairs, while several OpenNeuro datasets provide healthy controls.
Hospital data. The institutional cohort was collected from Xiangya Hospital and 11 independent centers: Changde First People’s Hospital, Tongji Hospital, the Second Affiliated Hospital of the University of South China, Shenzhen Second People’s Hospital, the Third Xiangya Hospital, Jiangxi Provincial People’s Hospital, Chongqing Traditional Chinese Medicine Hospital, the First and Second Affiliated Hospitals of Nanchang University, Hunan Children’s Hospital, and the First Hospital of Lanzhou University. These data include de-identified T1, T1c, T2, and FLAIR MRI, demographics, radiology reports, and pathological diagnoses when available.
Processing. Structured repositories were converted into a common patient-level format linking MRI, diagnosis labels, metadata, and available reports. Web-derived figures were separated, filtered, and aligned with their captions. Institutional reports were restricted to the four MRI sequences, standardized through a clinician-validated bilingual terminology workflow, and reviewed by neuroradiologists before use. The resulting 2D and 3D data were then organized for progressive model training.
Multi-source data processing
Separate workflows handle structured repositories, web-derived MRI–text pairs, and de-identified hospital data.
BrainTumor48K construction
The curated dataset contains 47,947 individuals, combining 2D MRI–text data with case-level 3D multimodal MRI.
Progressive training
Training connects visual representation learning, comprehensive diagnosis, report generation, and diagnostic confidence.
Abstract
Background: Preoperative diagnosis of brain tumors from multimodal MRI is challenging because different tumor types can share imaging features and specialist interpretation is time-intensive. Methods: We developed BrainVLM, a vision-language foundation model trained with MRI, demographics, and radiology reports to classify all 12 major WHO CNS5 tumor types, quantify diagnostic uncertainty, and generate radiology-style reports. BrainTumor48K includes 47,947 individuals collected from 39 public repositories and 12 collaborating medical centers. The model was evaluated on 5,211 pathologically confirmed cases, a prospective real-world cohort, and a blinded multi-reader study with 12 neuroradiologists. Findings: BrainVLM achieved a primary-test macro-AUC of 0.85 and F1 score of 0.82, and an external-test macro-AUC of 0.80 and F1 score of 0.75. AI assistance improved neuroradiologists’ mean F1 by 23.3% and reduced diagnostic time by 36.1%. The model also achieved AUCs of 0.95 and 0.88 for adult-type diffuse glioma molecular subtyping in primary and external cohorts. Interpretation: By combining prediction, calibrated confidence, and report generation, BrainVLM is designed to support transparent and efficient clinician review rather than replace clinical judgment.
BrainVLM at a glance
A specialized vision-language foundation model for preoperative brain tumor diagnosis, designed to pair accurate predictions with calibrated confidence and clinically readable reports.
From MRI to a clinical decision aid
BrainVLM combines complementary clinical signals and makes its output easier to review, communicate, and triage in real-world neuro-oncology workflows.
Multimodal input
Processes standard T1, T1c, T2, and T2-FLAIR MRI together with patient demographics and available radiology text.
Comprehensive classification
Targets the full spectrum of 12 major WHO CNS5 brain tumor categories, including low-prevalence and imaging-ambiguous tumors.
Confidence-aware output
Provides calibrated uncertainty estimates and can surface a supplementary top-two diagnosis when the leading confidence is below 75%.
Radiology-style report
Generates a clinical narrative describing imaging findings, supporting interpretation beyond an isolated tumor label.
Evidence across clinical settings
The paper evaluates BrainVLM on retrospective, multi-center external, prospective, and AI-augmented reader cohorts.
Diagnostic performance
The molecular-subgroup AUCs correspond to primary and external validation cohorts, respectively.
AI-augmented practice
BrainVLM is intended to support clinician review; lower-confidence cases should receive human assessment.
BrainTumor48K construction and preprocessing pipeline
Dataset preprocessing pipeline.
Overview of the construction pipeline of BrainTumor48K dataset.
Training pipeline for BrainVLM.
Research Results Overview
Overview of BrainVLM's dataset, model architecture and overall performance.
Multi-dimensional evaluation of BrainVLM's diagnostic performance.
Results of prospective study design, uncertainty quantification, molecular prediction, as well as comparative evaluation of radiology report.
Results of multi-reader study.
Subgroup analysis of BrainVLM's performance.
Result
Examples of reports and diagnoses generated by our BrainVLM, compared with those from doctors (expert-curated radiology report) and other state-of-the-art AI models.
Examples of reports and diagnoses generated by our BrainVLM, compared with those from doctors (expert-curated radiology report) and other state-of-the-art AI models.
Demo video
If you are unable to play the video, please try using a modern browser or contact the website administrator.