Hybrid GAN-DIFFUSION Framework for Ultra sound to CT Translation: A Residual Refinement Approach
| dc.contributor.author | Kulumba , Sserujja Abdallah | |
| dc.contributor.author | Bargami, Oumar Ahmad | |
| dc.contributor.author | Ibrahim, Zakariaou | |
| dc.date.accessioned | 2026-07-01T04:51:23Z | |
| dc.date.issued | 2025-10-25 | |
| dc.description | Supervised by Dr. Khondokar Habibul Kabir, Professor, Department of Electrical and Electronic Engineering (EEE) Islamic University of Technology (IUT) Board Bazar, Gazipur, Bangladesh This thesis is submitted in partial fulfillment of the requirement for the degree of Bachelor of Science in Electrical and Electronic Engineering, 2025 | |
| dc.description.abstract | The translation of 3D ultrasound (US) to computed tomography (CT) is a significant challenge in medical imaging due to the vast domain gap and inherent US artifacts. This thesis introduces a novel, two-stage hybrid framework to address this, synergizing a custom Generative Adversarial Network (GAN) with a Denoising Diffusion Probabilistic Model (DDPM). The first stage confronts GAN instability and mode collapse with our proposed UNetSelfCritique3D generator, which uses an internal feature-matching loss to enforce structural consistency and produce a coarse but anatomically faithful CT volume. The second stage employs a conditional DDPM for residual refinement, using the coarse GAN output as a strong prior to synthesize missing high-frequency details. Validated on the clinical TRUSTED dataset, our baseline 3D Pix2Pix model achieved a Peak Signal-to-Noise Ratio (PSNR) of 11.34 dB and a Structural Similarity Index (SSIM) of 0.181. The UNetSelfCritique3D significantly improved performance to 20.57 dB PSNR and 0.468 SSIM, demonstrating effective mitigation of mode collapse. The final hybrid GAN-Diffusion model achieved the best results with 21.10 dB PSNR, 0.456 SSIM, and 0.547 Normalized Cross-Correlation (NCC), representing a substantial 86% improvement in PSNR over the baseline. This work establishes the first crucial performance benchmark for 3D US-to-CT translation and validates the feasibility of our hybrid approach. The proposed framework and code are made publicly available to foster reproducibility and accelerate future research. | |
| dc.identifier.citation | [1] I. Goodfellow et al., “Generative adversarial nets,” in Advances in Neural Information Processing Systems 27, 2014, pp. 2672–2680. [2] J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” in Advances in Neural Information Processing Systems 33, 2020, pp. 6840–6851. [3] G. N. Hounsfield, “Computerized transverse axial scanning (tomography): Part 1. De scription of system,” The British Journal of Radiology, vol. 46, no. 552, pp. 1016–1022, 1973. [4] M.Arjovsky, S. Chintala, and L. Bottou, “Wasserstein generative adversarial networks,” in International Conference on Machine Learning, PMLR, 2017, pp. 214–223. [5] United Nations, Transforming our world: the 2030 Agenda for Sustainable Development, General Assembly, A/RES/70/1, 2015. [6] G. Litjens et al., “A survey on deep learning in medical image analysis,” Medical Image Analysis, vol. 42, pp. 60–88, 2017. [7] W. Ndzimbong et al., “TRUSTED: The Paired 3D Transabdominal Ultrasound and CT Human Data for Kidney Segmentation and Registration Research,” Scientific Data, vol. 12, no. 1, p. 615, 2025. [8] P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros, “Image-to-image translation with conditional adversarial networks,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 1125–1134. [9] J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros, “Unpaired image-to-image translation us ing cycle-consistent adversarial networks,” in Proceedings of the IEEE International Conference on Computer Vision (ICCV), 2017, pp. 2223–2232. [10] S. Hu, Y. Shen, S. Wang, B. Lei, et al., “Bidirectional mapping generative adversarial networks for brain mr to pet synthesis,” IEEE Transactions on Medical Imaging, vol. 41, no. 1, pp. 145–157, 2022. DOI: 10.1109/TMI.2021.3107013 [11] K. Armanious, C. He, S. Fischer, and T. Jiang, “MedGAN: Medical image translation using GANs,” Computerized Medical Imaging and Graphics, vol. 79, p. 101684, 2020. [12] H. Shan et al., “3D TomoGAN for high-resolution 3D medical image synthesis,” Medical Image Analysis, vol. 58, p. 101529, 2019 | |
| dc.identifier.uri | https://repository.iutoic-dhaka.edu/handle/123456789/2641 | |
| dc.language.iso | en | |
| dc.publisher | Department of Electrical and Electronic Engineering (EEE), Islamic University of Technology(IUT), Board Bazar, Gazipur-1704, Bangladesh | |
| dc.subject | Medical Image Translation | |
| dc.subject | Ultrasound | |
| dc.subject | Computed Tomography | |
| dc.subject | Generative Adver sarial Networks | |
| dc.subject | Diffusion Models | |
| dc.subject | Deep Learning | |
| dc.subject | Hybrid Models | |
| dc.subject | Residual Learning. | |
| dc.title | Hybrid GAN-DIFFUSION Framework for Ultra sound to CT Translation: A Residual Refinement Approach | |
| dc.type | Thesis |
Files
License bundle
1 - 1 of 1
Loading...
- Name:
- license.txt
- Size:
- 1.71 KB
- Format:
- Item-specific license agreed upon to submission
- Description:
