The Rahman Lab develops machine learning, computer vision, and multimodal AI methods for biomedical image understanding and healthcare applications. Our research spans medical image captioning, visual question answering, object detection and localization, dermatology and skin cancer analysis, explainable AI, privacy-preserving medical image generation, and trustworthy generative AI.

Medical Image Captioning & Vision-Language Models

Our lab develops vision-language models for generating clinically meaningful descriptions from medical images. This work investigates multimodal learning, parameter-efficient fine-tuning, modality-aware modeling, visual grounding, and explainability for radiology and biomedical image interpretation.

Model-intrinsic explainability pipeline for radiology image captioning

Published Work

  • CausalRAG-AD: Multimodal MRI Classification and Guideline-Compliant MRI Captioning for Alzheimer’s Diagnosis
    R. Farha, M. M. Rahman, F. Khalifa
    2026 IEEE 5th International Conference on Computing and Machine Intelligence
  • An Empirical Evaluation of Low-Rank Adapted Vision–Language Models for Radiology Image Captioning
    M. Hoque, R. N. Chowdhury, M. R. Hasan, O. O. E. Peter, F. Khalifa, M. M. Rahman
    Bioengineering, 2025, 12(12), 1330
  • Modality-Guided Radiology Caption Prediction with Small Vision-Language Models and Image Classifier
    R. N. Chowdhury, M. Hoque, M. R. Hasan, O. O. E. Peter, M. M. Rahman
    CLEF 2025 Working Notes
  • Comparative Analysis of Fine-Tuned Multimodal Models in Radiology Image Captioning
    M. Hoque, M. R. Hasan, M. I. S. Emon, E. P. O. Oluwafemi, M. M. Rahman, F. Khalifa
    2025 IEEE 4th International Conference on Computing and Machine Intelligence
  • Medical Image Interpretation with Large Multimodal Models
    M. Hoque, M. Hasan, M. Emon, F. Khalifa, M. Rahman
    CEUR Workshop Proceedings, 2024

Upcoming / Accepted Work

  • Model-Intrinsic Attention as Explanation for Radiology Image Captioning
    M. Hoque, R. N. Chowdhury, E. P. O. Oluwafemi, O. U. Islam, R. Hoque, M. M. Rahman
    Accepted for CLEF 2026 Working Notes

Medical Visual Question Answering (VQA)

Our lab develops multimodal and vision-language models for medical image interpretation and automated caption generation. This research explores parameter-efficient fine-tuning, modality-aware modeling, knowledge-guided reasoning, multimodal retrieval, and explainability across radiology and other medical imaging domains. The goal is to create efficient and clinically meaningful AI systems that can interpret medical images and generate useful textual descriptions.

Parameter-efficient GI endoscopy VQA and synthetic image generation framework

Published Work

  • Parameter-Efficient VLMs for Gastrointestinal Endoscopy: Medical Image Generation and Clinical Visual Question Answering
    O. O. E. Peter, F. Akor Ejiga, F. Khalifa, M. M. Rahman
    2025 IEEE EMBS International Conference on Biomedical and Health Informatics
  • Solving Medical Data Limitations Through AI: Multi-Modal Vision-Language Learning for Gastrointestinal VQA and Synthetic Training Data Generation
    E. P. O. Oluwafemi, M. Hoque, E. F. Akor, R. N. Chowdhury, et al.
    CLEF 2025 Working Notes
  • Advancing AI-Powered Medical Image Synthesis: Insights from MedVQA-GI Challenge Using CLIP, Fine-Tuned Stable Diffusion, and DreamBooth+LoRA
    O. O. E. Peter, M. M. Rahman, F. Khalifa
    2025

Upcoming / Accepted Work

  • Parameter-Efficient Vision–Language Fine-Tuning for Gastrointestinal Endoscopy VQA with Structured Explainability and Safety-Oriented Reasoning
    R. N. Chowdhury, O. U. Islam, E. P. O. Oluwafemi, M. Hoque, E. F. Akor, M. M. Rahman
    Accepted for CLEF 2026 Working Notes

Object Detection, Segmentation & Medical Image Analysis

Our lab develops machine learning and deep learning methods for detecting, localizing, classifying, and segmenting clinically relevant structures in biomedical images. This research includes brain tumor detection, polyp segmentation, colonoscopy image analysis, and multimodal approaches that combine vision-language models with conventional detection architectures.

DB- and LoRA-Based Colonoscopy Image Generation

Published Work

  • An Attention-Guided Deep Learning Framework for Brain Tumor Detection Using Multimodal MR Imaging
    I. Abdelhaliem, O. Akinniyi, J. Aina, J. Dixon, S. Mehravaran, M. M. Rahman, et al.
    2026 IEEE 5th International Conference on Computing and Machine Intelligence
  • Comparative Analysis of Knowledge-Guided Few-Shot Brain Tumor Detection Using Efficient Vision–Language Models and a Single-Shot Detector
    O. U. Islam, J. Muchangi, A. Bonojo, F. Khalifa, M. M. Rahman
    2026 IEEE 5th International Conference on Computing and Machine Intelligence
  • An Interpretable Framework for Brain Tumor Detection Integrating a Single-Shot Detector and Vision-Language Model (VLM)
    O. U. Islam, A. Ayomide, M. M. Rahman
    EMBC 2026
  • Synthetic Data-Driven Multi-Architecture Framework for Automated Polyp Segmentation Through Integrated Detection and Mask Generation
    O. O. E. Peter, A. Oluwapemiisin, A. Chetachi, A. Opeyemi, F. Khalifa, et al.
    Medical Imaging 2025: Clinical and Biomedical Imaging
  • Text-Guided Synthesis in Medical Multimedia Retrieval: A Framework for Enhanced Colonoscopy Image Classification and Segmentation
    O. O. Ejiga Peter, O. T. Adeniran, A. M. G. John-Otumu, F. Khalifa, M. M. Rahman
    Algorithms, 2025, 18(3), 155
  • A Hybrid Learning-Architecture for Improved Brain Tumor Recognition
    J. Dixon, O. Akinniyi, A. Abdelhamid, G. A. Saleh, M. M. Rahman, F. Khalifa
    Algorithms, 2024, 17(6), 221

Skin Cancer & Dermatology AI

Our lab is exploring knowledge-guided and resource-efficient vision-language models for dermatology image understanding. This work focuses on generating clinically meaningful descriptions from dermatology images while improving grounding, efficiency, and cross-domain applicability of multimodal AI systems.

Knowledge-Guided Dermatology Image Captioning with Fairness-Aware Vision-Language Models

Published Work

  • Automated Melanoma Recognition in Dermoscopic Images Based on Extreme Learning Machine (ELM)
    M. M. Rahman, N. Alpaslan
    SPIE Medical Imaging, 2017
  • Developing a Retrieval-Based Diagnostic Aid for Automated Melanoma Recognition of Dermoscopic Images
    M. M. Rahman, N. Alpaslan, P. Bhattacharya
    IEEE Applied Imagery Pattern Recognition Workshop, 2016

Current / Submitted Work

  • Knowledge-Guided Dermatology Image Captioning Using Resource-Efficient Vision-Language Models
    R. N. Chowdhury, M. Hoque, E. P. O. Oluwafemi, F. Khalifa, B. Ojeme, M. M. Rahman
    Manuscript submitted to Computer Methods and Programs in Biomedicine

Deepfake Detection & Generative AI

Our lab investigates multimodal generative AI and deepfake analysis across image and audio domains. This work focuses on both synthetic media generation and robust detection, combining transformer-based vision models, audio representation learning, and multimodal reasoning.


Multimodal Deepfake Detection Framework for Image and Audio Analysis

Upcoming / Accepted Work

  • Cross-Architecture Detection and Constraint-Guided Generation for Audio and Visual Deepfake Evaluation: Lessons from ImageCLEF 2026
    D. Emakporuena, R. N. Chowdhury, S. S. Alam, E. P. O. Oluwafemi, M. Hoque, O. G. Akingbola, M. M. Rahman
    Accepted for CLEF 2026 Working Notes

Privacy-Preserving Medical Image Generation

Our lab investigates generative AI methods for creating realistic synthetic medical images while reducing privacy risks associated with sensitive clinical data. This work explores privacy–utility tradeoffs, synthetic CT generation, and evaluation strategies for determining whether generated medical images preserve useful clinical information without memorizing patient-specific data.

Privacy-Oriented Synthetic CT Image Generation

Upcoming / Accepted Work

  • Empirical Privacy Without Differential Privacy: A Privacy–Utility Frontier Analysis of Synthetic CT Generation
    M. R. Shaharear, E. P. O. Oluwafemi, M. Hoque, D. Briggs, M. M. Rahman
    Accepted for CLEF 2026 Working Notes