BrainVLM
02 · Diagnosis and retrospective evidence

Diagnosis & Retrospective Validation

BrainVLM is evaluated across 12 WHO CNS5 tumor categories in primary and independent external cohorts.

12 WHO CNS5 categories3,877 primary-test patients1,334 external-test patients11 external hospitals

One comprehensive diagnostic task

Instead of restricting evaluation to a few common tumors, BrainVLM is trained and tested across the major diagnostic groups defined by WHO CNS5.

Coverage

Major classes and most subtypes

The model targets all 12 major categories and most subtypes, including less prevalent and imaging-ambiguous entities that are rarely covered by existing AI studies.

Case-level input

Standard multimodal MRI

T1, T1c, T2, and T2-FLAIR are interpreted together with age and sex at the study level, without requiring advanced MRI sequences.

Reference standard

Pathology-confirmed evaluation

The retrospective benchmark contains 5,211 pathologically confirmed tumors, supporting direct evaluation of preoperative classification against the definitive diagnosis.

Why preoperative classification matters. The initial diagnosis can change the treatment pathway: diffuse gliomas commonly require maximal safe resection, whereas lymphomas and some germ-cell tumors may be treated primarily with chemotherapy and/or radiotherapy without surgical resection.

Diagnostic scope across WHO CNS5

BrainVLM covers the 12 major WHO CNS5 brain tumor categories and extends toward less prevalent and imaging-ambiguous entities that are underrepresented in existing AI studies.

BrainVLM coverage of major WHO CNS5 brain tumor categories
Coverage view: major WHO CNS5 categories and less prevalent or imaging-ambiguous entities represented in the diagnostic benchmark.

Retrospective and external validation

Performance was first measured in 3,877 held-out patients from the primary hospital and then tested in 1,334 patients from 11 independent hospitals. Postoperative pathology served as the reference standard.

Independent evaluation design. The external test hospitals did not contribute cases to model training or validation. BrainVLM was compared with board-certified neuroradiologists and four AI baselines using sensitivity, precision, F1, Cohen’s κ, and AUC, with 95% confidence intervals estimated by 1,000 bootstrap replicates.
BrainVLM retrospective category-level performance
Category-level sensitivity, precision, F1, and kappa for BrainVLM and board-certified neuroradiologists in the primary and external cohorts.
BrainVLM retrospective confusion matrices
Normalized confusion matrices comparing BrainVLM with neuroradiologists across the 12 tumor categories.
False negative rates for 12 brain tumor types comparing BrainVLM with board-certified neuroradiologists
False-negative rates across 12 tumor types for BrainVLM and board-certified neuroradiologists in the primary and external retrospective cohorts.
0.85Primary macro-AUC · 95% CI 0.84–0.86
0.82Primary F1 · radiologists 0.80
0.80External macro-AUC · 95% CI 0.79–0.82
0.75External F1 · radiologists 0.71
Primary cohort

Higher precision and agreement

BrainVLM reached precision 0.84, sensitivity 0.82, and Cohen’s κ 0.75, compared with 0.71, 0.80, and 0.69 for neuroradiologists.

External cohort

Performance transferred across centers

External F1 remained 0.75, exceeding neuroradiologist consensus at 0.71 and comparator AI models, whose F1 scores ranged from 0.30 to 0.52.

Less prevalent tumors

Gains beyond common entities

In the primary cohort, F1 was higher than neuroradiologists for hematolymphoid (0.64 vs 0.45), embryonal (0.68 vs 0.58), and choroid plexus tumors (0.55 vs 0.47).

How to interpret the results

The category-level analysis provides more context than a single aggregate score and helps identify both clinically useful strengths and remaining boundaries.

Common tumors

Comparable to expert assessment

For meningioma, the glioma/glioneuronal/neuronal group, and cranial or paraspinal nerve tumors, BrainVLM matched or exceeded expert-level F1 in the primary cohort.

Error structure

Model and experts share difficult boundaries

More than 90% of the two most frequent misclassification pairs overlapped between BrainVLM and human experts in both retrospective cohorts.

Input efficiency

Standard sequences were sufficient

BrainVLM used standard T1, T1c, T2, and T2-FLAIR plus demographics, while neuroradiologists consulted additional advanced imaging in 56.1% of cases.

Important boundary. Performance still varied by tumor type, and rare-tumor samples remain limited. The real-world institutional cohorts were also confined to East Asian populations, so broader prospective and demographic validation remains necessary.

Extension to molecularly defined diffuse glioma subtypes

Beyond the 12-category task, BrainVLM was fine-tuned to distinguish IDH-mutant astrocytoma, IDH-mutant oligodendroglioma, and IDH-wildtype glioblastoma—three clinically important adult-type diffuse glioma subgroups.

n=494Primary molecular-subtyping cohort
AUC 0.95Primary three-fold cross-validation
n=138Independent external cohort
AUC 0.88External molecular-subtyping performance
For radiology-style reports and calibrated confidence, continue to Reports & confidence. Prospective and multi-reader experiments are presented on Prospective & readers.