Hybrid GAN-DIFFUSION Framework for Ultra sound to CT Translation: A Residual Refinement Approach
Loading...
Date
Journal Title
Journal ISSN
Volume Title
Publisher
Department of Electrical and Electronic Engineering (EEE), Islamic University of Technology(IUT), Board Bazar, Gazipur-1704, Bangladesh
Abstract
The translation of 3D ultrasound (US) to computed tomography (CT) is a significant challenge in
medical imaging due to the vast domain gap and inherent US artifacts. This thesis introduces a
novel, two-stage hybrid framework to address this, synergizing a custom Generative Adversarial
Network (GAN) with a Denoising Diffusion Probabilistic Model (DDPM). The first stage
confronts GAN instability and mode collapse with our proposed UNetSelfCritique3D generator,
which uses an internal feature-matching loss to enforce structural consistency and produce a
coarse but anatomically faithful CT volume. The second stage employs a conditional DDPM
for residual refinement, using the coarse GAN output as a strong prior to synthesize missing
high-frequency details. Validated on the clinical TRUSTED dataset, our baseline 3D Pix2Pix
model achieved a Peak Signal-to-Noise Ratio (PSNR) of 11.34 dB and a Structural Similarity
Index (SSIM) of 0.181. The UNetSelfCritique3D significantly improved performance to 20.57
dB PSNR and 0.468 SSIM, demonstrating effective mitigation of mode collapse. The final hybrid
GAN-Diffusion model achieved the best results with 21.10 dB PSNR, 0.456 SSIM, and 0.547
Normalized Cross-Correlation (NCC), representing a substantial 86% improvement in PSNR over
the baseline. This work establishes the first crucial performance benchmark for 3D US-to-CT
translation and validates the feasibility of our hybrid approach. The proposed framework and
code are made publicly available to foster reproducibility and accelerate future research.
Description
Supervised by
Dr. Khondokar Habibul Kabir,
Professor,
Department of Electrical and Electronic Engineering (EEE)
Islamic University of Technology (IUT)
Board Bazar, Gazipur, Bangladesh
This thesis is submitted in partial fulfillment of the requirement for the degree of Bachelor of Science in Electrical and Electronic Engineering, 2025
Citation
[1] I. Goodfellow et al., “Generative adversarial nets,” in Advances in Neural Information Processing Systems 27, 2014, pp. 2672–2680. [2] J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in Advances in Neural Information Processing Systems 33, 2020, pp. 6840–6851. [3] G. N. Hounsfield, “Computerized transverse axial scanning (tomography): Part 1. De scription of system,” The British Journal of Radiology, vol. 46, no. 552, pp. 1016–1022, 1973. [4] M.Arjovsky, S. Chintala, and L. Bottou, “Wasserstein generative adversarial networks,” in International Conference on Machine Learning, PMLR, 2017, pp. 214–223. [5] United Nations, Transforming our world: the 2030 Agenda for Sustainable Development, General Assembly, A/RES/70/1, 2015. [6] G. Litjens et al., “A survey on deep learning in medical image analysis,” Medical Image Analysis, vol. 42, pp. 60–88, 2017. [7] W. Ndzimbong et al., “TRUSTED: The Paired 3D Transabdominal Ultrasound and CT Human Data for Kidney Segmentation and Registration Research,” Scientific Data, vol. 12, no. 1, p. 615, 2025. [8] P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros, “Image-to-image translation with conditional adversarial networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 1125–1134. [9] J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros, “Unpaired image-to-image translation us ing cycle-consistent adversarial networks,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2017, pp. 2223–2232. [10] S. Hu, Y. Shen, S. Wang, B. Lei, et al., “Bidirectional mapping generative adversarial networks for brain mr to pet synthesis,” IEEE Transactions on Medical Imaging, vol. 41, no. 1, pp. 145–157, 2022. DOI: 10.1109/TMI.2021.3107013 [11] K. Armanious, C. He, S. Fischer, and T. Jiang, “MedGAN: Medical image translation using GANs,” Computerized Medical Imaging and Graphics, vol. 79, p. 101684, 2020. [12] H. Shan et al., “3D TomoGAN for high-resolution 3D medical image synthesis,” Medical Image Analysis, vol. 58, p. 101529, 2019
