Verification of the impact of differences between objective and subjective evaluation methods on the interpretation of artificial intelligence systems generating mammogram reports using vision-language model.
Authors
Affiliations (7)
Affiliations (7)
- Department of Intelligent Information Engineering, Research Promotion Unit, School of Medical Sciences, Fujita Health University, 1-98 Dengakugakubo, Kutsukake-Cho, Toyoake-City, Aichi, 470-1192, Japan. [email protected].
- Department of Intelligent Information Engineering, Research Promotion Unit, School of Medical Sciences, Fujita Health University, 1-98 Dengakugakubo, Kutsukake-Cho, Toyoake-City, Aichi, 470-1192, Japan.
- The Asahi Shimbun Company, 5-3-2 Tsukiji, Chuo-Ku, Tokyo, 104-8011, Japan.
- TOITU CO., LTD., 1-5-10 Ebisu-Nishi, Shibuya-Ku, Tokyo, 150-0021, Japan.
- Graduate School of Engineering, Muroran Institute of Technology, 27-1 Mizumoto-Cho, Muroran City, Hokkaido, 050-8585, Japan.
- Ohtsuka Breastcare Clinic, 5-7-3 Takenotsuka, Adachi-Ku, Tokyo, 121-0813, Japan.
- Department of Radiological Technology, Faculty of Medical Technology, Niigata University of Health and Welfare, 1398 Shimami-Cho, Kita-Ku, Niigata City, Niigata, 950-3198, Japan.
Abstract
Research on vision-language models (VLMs) in the medical field has recently increased. However, while multifaceted evaluation is necessary to avoid the high risks associated with misdiagnosis, artificial intelligence (AI)-assisted mammogram report generation remains insufficient, with no studies on objective and subjective generation. We aimed to develop an AI system that generates mammogram reports and to verify the impact of differences between objective and subjective evaluation methods on the interpretation of this AI system. We used a public dataset consisting of mammograms and their reports, preparing question prompts and performing low-rank adaptation tuning on Qwen2.5(7B). We analyzed the Breast Imaging Reporting and Data System (BI-RADS) and findings agreement rate, Recall-Oriented Understudy for Gisting Evaluation (ROUGE), and Bilingual Evaluation Understudy (BLEU) for the objective evaluation. A breast clinician performed score-based evaluations of generated reports as subjective assessments. Finally, we analyzed samples of inconsistent objective and subjective results. The BI-RADS agreement rate was 58.1%. Findings were accurately included in generated reports at 76.7% for mass and 81.4% for calcification. ROUGE-L F1 and overall BLEU were 0.672 and 0.542, respectively. Although ROUGE-L F1 or overall BLEU was above average, two samples received low scores in the generated report evaluation; these were considered over- and under-estimation. We developed an AI model that generates mammogram reports. In assessing this model, we found that multifaceted objective and subjective medical VLM evaluations are necessary for determining over- and under-estimation.