Patient-specific multimodal learning with multi-view contrastive alignment for chest X-ray report generation
The proposed EVOKE surpasses recent state-of-the-art methods across multiple datasets, and introduces a multi-view contrastive learning method that captures semantic correspondences both among multi-view radiographs within a study and between these radiographs and their associated report, thereby improving visual repre...