Computer Vision System and Transformer-Based Multimodal Fusion of Medical Images for Enhanced Diagnostic Decision Making

Authors

  • soltan Abdullah Tikrit University
  • Kahlan F . Aljobory

DOI:

https://doi.org/10.31185/wjcms.543

Keywords:

Deep learning , computer vision , Medical Images , Diagnostic Decision

Abstract

The rise of artificial intelligence technologies and their increasing reliance in numerous fields, most notably medicine, is significant. AI plays a crucial role in assisting physicians with accurate patient diagnosis, enabling them to make informed decisions. This article presents a method for integrating computer vision technologies into the medical field, where multimodal medical imaging integrates structural and textual information to support clinical decision-making. However, most image fusion frameworks based on transformers remain narrowly designed for a single media combination and a single subsequent task. This paper proposes the Multimodal Adaptive Fusion Transformer (CMAFT), a unified architecture that combines Swin transformer encoders for each media with a Catalyst Cross Attention Fusion Module (GCAF) to dynamically weight each media's contribution based on its real-time reliability. CMAFT was evaluated against two structurally different multimodal reference standards: the BraTS 2021 dataset, in which four MRI sequences (T1, T1ce, T2, FLAIR) were combined to segment brain tumor volumes, and the MIMIC-CXR dataset, in which chest X-ray images were combined with duplex radiology report inclusions to classify multimodal pathology and generate reports. The proposed GCAF module provides a learnable reliability gateway that curbs degraded or lost data streams, while a multiscale hierarchical integration pyramid preserves both fine detail and overall semantic context. Experiments show that CMAFT achieves average Dice scores of 92.4%, 87.9%, and 84.6% for whole tumor, tumor nucleus, and enhanced tumor, respectively, on the BraTS 2021 dataset, and a multi-class average area under the curve (AUC) of 0.876 with an overall F1 value of 0.658 on the MIMIC-CXR dataset, outperforming modern hybrid CNN-Transformer models while using fewer parameters compared to similar dual-code baseline models. Resection studies confirm that both reliability gate and multi-scale integration contribute to measurable gains, and that CMAFT degrades smoothly when either an MRI sequence or radiology report is unavailable. These results suggest that a single adaptive mutual attention integration design can be generalized across volumetric and text-based multimodal diagnostic pathways.

Downloads

Download data is not yet available.

Downloads

Published

2026-09-30

Issue

Section

Computer

How to Cite

[1]
soltan Abdullah and K. F . Aljobory, “Computer Vision System and Transformer-Based Multimodal Fusion of Medical Images for Enhanced Diagnostic Decision Making”, WJCMS, vol. 5, no. 3, pp. 12–25, Sep. 2026, doi: 10.31185/wjcms.543.