Skip to content
Open access

Severity Matters: Risk-Calibrated Reinforcement Fine-Tuning for Clinically Aligned Medical Vision–Language Models

2026 · IEEE Access · Vol 14, pp. 115480-115505 · 0 citations · 76 references

Abstract

In clinical vision-language model (VLM) post-training, the relevant failure mode is not only how often a model is wrong, but whether errors concentrate in cases where missed findings carry greater clinical cost. Risk-Calibrated Reinforcement Fine-Tuning (RCRFT) addresses this mismatch by making severity an explicit signal in reinforcement learning rather than an after-the-fact evaluation stratum. The framework combines a severity-weighted multimodal reward, slot-aware post-rollout edits that respect clinical report structure, and a unified clinical utility spanning report quality, grounding fidelity, calibration, and safety penalties. A formal analysis decomposes the risk-weighted objective into mean utility, a global severity-scaling term, and a severity–reward covariance term; under severity centering only the mean and the covariance remain, indicating that the method favors positive association between utility and case severity rather than simple reward amplification. On MIMIC-CXR, this shift in optimization improves the safety-critical region first: relative to a matched RL baseline, severe-case false negatives fall from 9.8% to 5.6%, and high-risk utility rises from 0.65 to 0.74. The same training signal also lifts general report quality, raising ROUGE-L from 29.9 to 30.8 and RadGraph F1 from 25.8 to 27.3 while lowering RadCliQ from 0.86 to 0.72. Ablations show that each component contributes, with severity weighting accounting for the largest share of the improvement.

Read PDF