Comparative assessment of DeepSeek-R1 and GPT-4 for structured ultrasound reporting of adnexal masses
Objective This study primarily evaluated the ability of two large language models (DeepSeek-R1 and GPT-4) to generate structured ultrasound reports from free-text adnexal mass reports. Secondarily, we assessed their accuracy in O-RADS classification and management recommendations, with an exploratory analysis of their...