Wavelet-guided Vision Transformer with Structural Gradient Loss for Denoising Low-Field Brain MRI and Downstream Segmentation
DOI:
https://doi.org/10.70917/ijcisim-2026-4479Keywords:
Brain MRI Denoising, Vision Transformer (ViT), Discrete Wavelet Transform (DWT), Rician Noise, Structural Gradient Loss, Low-Field NeuroimagingAbstract
Low-field Magnetic Resonance Imaging (MRI) is a cost-effective, accessible point-of-care alternative for neuroimaging. However, lower magnetic field strengths (<1.5T) inherently suffer from low Signal-to-Noise Ratios and widespread, signal-dependent Rician noise interference. Standard spatial domain denoising filters result in substantial blurring at critical structural boundaries, whereas typical Convolutional Neural Networks (CNNs) cannot capture the global anatomical spatial context. We propose a new Wavelet-Informed Vision Transformer (W-ViT) framework for restoring low-field brain MRI in this paper. The network integrates multi-scale frequency decomposition of Discrete Wavelet Transform (DWT) and multi-head self-attention mechanisms to separate high-frequency noise components in orthogonal wavelet sub-bands while explicitly preserving fine structural brain morphology. The architecture is optimized via a multi-component composite loss function combining pixel-wise reconstruction loss with an explicit Structural Gradient Penalty. Crucially, while previous restoration works prioritize general contrast enhancement, W-ViT specifically target-preserves tissue boundary gradients to directly prevent downstream automated tissue segmentation failures. By ensuring high structural fidelity at low operational costs, W-ViT bridges the gap between low-field hardware signal degradation and downstream diagnostic segmentation requirements. Comprehensive quantitative and qualitative assessments performed on public benchmark datasets reveal that the suggested model attains a Peak Signal-to-Noise Ratio (PSNR) of 34.12 dB and a Structural Similarity Index (SSIM) of 0.924 in both simulated and actual low-field Rician noise conditions. Comparative evaluations show that W-ViT significantly outperforms standard DnCNN, BM4D, and Non-Local Means baselines while maintaining low inference latency suitable for real-time neuroimaging pipeline.