BrainVLM
04 · Prospective and multi-reader evidence

Prospective & Multi-reader Study

Real-world prospective cases and a blinded reader study evaluate clinical performance and AI-assisted interpretation.

Prospective cohort 248 reader-study cases 12 neuroradiologists −36.1% diagnostic time

Prospective study

Real-world evaluation before definitive diagnosis

Cases entered BrainVLM along routine surgical and non-operative pathways, under real acquisition and workflow conditions—before the final reference diagnosis was available. Among 1,162 admissions, 776 cases entered the prospective workflow after predefined exclusions and then followed surgical or non-operative pathways, with pathology or radiology diagnosis as the reference outcome.

Collection flow

Cohort at a glance

Eligible prospective casesn=776
Clinical pathwaysSurgery & non-operative care
Reference outcomePathology or radiology diagnosis
BrainVLM prospective cohort collection and diagnostic flow
Prospective cohort collection flow from admission through surgery or non-operative management and reference diagnosis.
Prospective results

Performance remains strong before definitive diagnosis

In the primary prospective cohort, BrainVLM reached 0.85 sensitivity and 0.87 precision, with a macro-F1 of 0.84. External prospective results remained competitive, with sensitivity 0.78, precision 0.73, and macro-F1 0.75.

BrainVLM primary and external prospective study results
Comparison of primary and external prospective study results across sensitivity, precision, F1, and kappa.
n=776Patients entering BrainVLM prospective workflow
0.87Primary prospective precision
0.84Primary prospective macro-F1
F1 +4%Prospective improvement versus radiologists

Multi-reader study

A blinded reader study tested whether BrainVLM can improve diagnostic performance and efficiency across different levels of neuroradiology experience.

Study design and reader performance

AI assistance improved performance across seniority

In 248 cases, 12 neuroradiologists reviewed the same studies without AI and then with BrainVLM output. Readers were stratified by experience level to test whether assistance generalized across clinical backgrounds.

Junior neuroradiologistsF1 0.52 → 0.67
Senior neuroradiologistsF1 0.55 → 0.74
Expert neuroradiologistsF1 0.79 → 0.85
Diagnostic time−36.1%
Safety and difficult-case analysis. Among 57 cases with an incorrect BrainVLM suggestion, overall reader accuracy remained stable after AI review (50.0% without AI vs 50.9% with AI). In a separate bottom-10% low-confidence subset (N=385), BrainVLM reached 65.0% accuracy, outperforming the two junior readers (51% and 44%) and approaching the expert readers (72% and 66%). These analyses suggest that low confidence identifies genuinely difficult cases and that incorrect suggestions did not produce systematic over-reliance in this study.
BrainVLM multi-reader clinical evaluation
Multi-reader evaluation across reader experience levels, with and without BrainVLM assistance.

Representative clinical impact

Two cases from the multi-reader study illustrate how BrainVLM assistance can shorten review time while correcting or strengthening the expert interpretation.

Two representative cases showing reduced diagnostic time and increased confidence with BrainVLM assistance
Supplementary Figure 16: BrainVLM assistance reduced expert reading time by 70% and 71% in two representative cases; one case corrected an initial misdiagnosis and the other increased diagnostic confidence.
Case 1 · Diagnostic correction

From glioma to the pathology-confirmed brain metastasis

A 62-year-old woman was initially diagnosed by the expert reader as having a glioma with 60% confidence after 132 seconds of review. BrainVLM instead suggested brain metastasis—the pathology-confirmed category—with 75% confidence. With this additional evidence, the expert corrected the diagnosis in 40 seconds, reducing reading time by 70%.

Case 2 · Diagnostic reinforcement

Faster confirmation of a difficult pediatric glioma

An 8-year-old girl was correctly diagnosed with glioma by the expert, but the unaided interpretation required 80 seconds and carried only 60% confidence. BrainVLM independently predicted glioma with 90% confidence; with AI support, the expert finalized the diagnosis in 23 seconds, increased confidence to 70%, and reduced reading time by 71%.

Two complementary forms of assistance. The first case illustrates diagnostic correction, while the second illustrates reinforcement of an already correct but uncertain interpretation. Together, they show how the diagnosis, confidence estimate, and supporting report can act as review cues rather than replacing the expert decision.

Clinical interpretation

These experiments support a workflow in which BrainVLM provides structured evidence while clinicians remain responsible for final interpretation.

23.3%Mean reader F1 improvement with AI assistance
36.1%Diagnostic time reduction
12Neuroradiologists in the reader study
248Cases in the blinded reader evaluation
BrainVLM is a research and clinical decision-support system, not a standalone medical device. Further validation is needed before deployment in a specific clinical population.