Master of Science in Data Science and Analytics
Permanent URI for this collectionhttps://hdl.handle.net/20.500.11951/1204
Browse
Browsing Master of Science in Data Science and Analytics by Subject "attention maps"
Now showing 1 - 1 of 1
- Results Per Page
- Sort Options
Item An Explainable Transformer-Based Vision-Language Model For Multimodal Consistency Verification(none, 2026-09-23) Mugimba Kakure JudeThe rapid growth of multimodal data has underscored the necessity for robust methods to verify semantic alignment between images and their associated textual descriptions. This problem, termed multimodal consistency verification, asks whether a given image and text pair are semantically compatible or contradictory. Unlike conventional image–text retrieval, which ranks many candidates for a query, consistency verification is a pairwise decision on a single pair. While large-scale vision-language models demonstrate strong performance in retrieval and zero-shot tasks, explicit consistency verification, particularly with integrated explainability, remains underexplored. We present an explainable dual-encoder vision-language architecture that utilizes a Vision Transformer (ViT) and Bidirectional Encoder Representations from Transformers (BERT) to encode images and multi-caption documents, respectively. These representations are projected into a shared embedding space and aligned via a symmetric InfoNCE contrastive objective. We quantify consistency using cosine similarity, employing a decision threshold to distinguish between consistent and inconsistent pairs. The model incorporates visual attention maps from the ViT and token importance scores from the BERT encoder that provide qualitative evidence for its classification decisions. For evaluation, a proxy benchmark dataset was constructed from the Flickr30K dataset by aggregating the five captions associated with each image into a single multi-caption document and generating consistent, randomly mismatched, and Facebook AI Similarity Search (FAISS) retrieved hard-negative pairs. The inconsistent pairs were generated algorithmically rather than through extensive human annotation of semantic contradiction type. The model was tested across balanced, imbalanced, and hard-negative scenarios and compared with Contrastive Language–Image Pre-training (CLIP) as a baseline in all scenarios. To isolate the contributions of the transformer-based encoders, the Vision Transformer was replaced by Residual Network-50 (ResNet-50) to evaluate the impact of using a convolutional rather than an attention-based image encoder, whereas BERT was replaced by Term Frequency–Inverse Document Frequency (TF-IDF) representations to assess the value of contextual language modelling against a non contextual text baseline. The results show that the proposed model excels at identifying random image and textual mismatches on this proxy benchmark, achieving an ROC-AUC between 0.989 and 0.999, and maintains strong ranking quality under an imbalanced 90/10 class distribution (ROC-AUC 0.9894). However, performance declines substantially when faced with hard negatives (ROC-AUC approximately 0.70), indicating that semantically similar mismatches remain a challenge for dual-encoder architectures. The ablation results confirmed that both the transformer-based vision and language encoders are important for performance relative to the convolutional and TF-IDF alternatives. Qualitative analysis showed that attention maps and token rankings highlight influential image regions and tokens; these visualisations were not evaluated with quantitative faithfulness tests and are not claimed as causal explanations. The study is therefore limited to a public multi-caption photograph benchmark with algorithmically constructed negatives, and the findings should not be generalised to domains whose images, documents, or contradiction types differ from this setting.
