BrainVLM
03 · Reports and confidence

Reports & Confidence

Each diagnosis is accompanied by a radiology-style report and a calibrated confidence estimate.

Three linked outputs

Classification, explanation, and reliability are designed as one case-level output.

01 · Diagnosis

Comprehensive tumor prediction

The model predicts across the major WHO CNS5 categories and most clinically relevant subtypes.

02 · Report

Radiology-style rationale

Generated findings describe enhancement, signal intensity, lesion location, and other imaging attributes supporting the impression.

03 · Confidence

Uncertainty-aware review

Confidence helps distinguish straightforward predictions from cases that may benefit from a second candidate or human review.

Report quality and clinical content

Each report metric is paired with a short clinical reading on the left and the corresponding figure on the right.

Automated evaluation

Faithful wording and clinical entities

BrainVLM leads both cohorts on the complementary report metrics: primary / external RaTEScore 0.75 / 0.69, RadGraph-XL F1 0.57 / 0.52, and BLEU-4 0.43 / 0.35. The combination reflects both semantic content and lexical agreement rather than fluency alone.

BrainVLM report quality metric comparison
Report-level comparison using RaTEScore, F1, RadGraph-XL, and BLEU-4.
Clinical content

Imaging evidence survives generation

An LLM-as-a-judge assessment found 80% correctness for contrast enhancement, 71–74% for T1, T2, and T2-FLAIR signal patterns, and 60% for lesion localization. Localization is the hardest attribute, but still exceeds the strongest baseline at 51%.

BrainVLM clinical content accuracy across report attributes
Clinical content accuracy for MRI signal characteristics and lesion localization.

Confidence as a workflow signal

Confidence is evaluated for reliability and used to support confidence-triggered Top-2 review of ambiguous cases.

BrainVLM confidence calibration analysis
Confidence calibration analysis across the model’s prediction outputs.
Confidence calibration

Confidence separates routine from uncertain cases

Reliable BrainVLM concentrates correct answers in the highest-confidence range while reducing overconfident errors, creating a practical signal for when a case deserves closer review.

Vanilla and reliable BrainVLM confidence distributions
Vanilla versus reliable BrainVLM confidence distributions for correct and error answers.
Reliable inference

Correct and error distributions become visible

The revised confidence analysis contrasts vanilla and reliable BrainVLM across confidence bands. This makes calibration failures easier to identify and supports a confidence-triggered Top-2 review when the first prediction is uncertain.

73%Correct diagnoses above 90% confidence
70%Incorrect diagnoses in the 50–85% range
0.82 → 0.86Primary F1 with confidence-triggered Top-2 review
ReviewConfidence is a triage signal, not a replacement for judgment