Speech Denoising Using Diffusion Model Technique

dc.contributor.authorLaouar Meriem
dc.contributor.authorZahouani Sarra
dc.date.accessioned2026-09-04T16:56:58Z
dc.date.issued2026
dc.descriptionSpecialty: Intelligent Systems for Knowledge Extraction (SIEC) Adjila Abderrahmane/encadreur
dc.description.abstractSpeech is the primary medium for human communication, yet it is almost always degraded by background noise in real-world environments. This noise reduces intelligibility and harms the performance of critical applications such as automatic speech recognition, telecommunications, and audio streaming. Speech denoising, therefore, plays a vital role in restoring clean speech and ensuring reliable human-machine interaction. This thesis explores diffusion-based generative models for speech denoising, focusing on the DiffWave architecture. The objective is to build a system that produces high-quality enhanced speech while operating fast enough for real-world deployment. To this end, two diffusion step configurations are compared: T=50 and T=200. Although T=200 provides a finer discretization, T=50 achieves better generalization on unseen data (SI-SDR: 10.93 dB vs 9.94 dB), with an optimal inference step of t=5. This is primarily because T=50 was trained for significantly longer (240,000 iterations compared to 65,000 for T=200) due to GPU constraints, allowing it to converge more effectively. To overcome inference latency, Denoising Diffusion Implicit Models (DDIM) are integrated, reducing sampling steps from 50 to just 3. This achieves a real- time factor (RTF) of 0.0732, corresponding to 13.7× real-time performance, with a minimal quality loss of only -0.64 dB. A key contribution is the evaluation on Arabic speech. Although trained exclusively on English digits, the model achieves its best performance on the Arabic dataset (SI-SDR = 3.94 dB), with the majority of test samples showing improvement after processing, confirming strong cross-lingual generalization. This is attributed to the model’s ability to learn latent acoustic representations that transcend language boundaries. The model was also evaluated on French speech using the Common Voice dataset, demonstrating clear cross-lingual generalization capabilities with promis- ing results that support its effectiveness in diverse linguistic environments. The model is also validated on real-world recordings from office, street, and cafeteria environments, confirming its practical applicability beyond controlled laboratory conditions. These results demonstrate that the proposed approach offers a practical, fast, and language-robust solution for speech denoising, suitable for real-time applications.
dc.identifier.urihttps://dspace.univ-ghardaia.edu.dz/handle/123456789/10757
dc.publisheruniversity of ghardaia
dc.subjectSpeech Denoising
dc.subjectSpeech Processing
dc.subjectAudio Signal Processing
dc.subjectDeep Learning
dc.subjectDiffusion Models
dc.subjectSpeech Enhancement
dc.subjectSignal-to-Noise Ratio (SNR)
dc.subjectDDPM
dc.subjectDDIM.
dc.titleSpeech Denoising Using Diffusion Model Technique
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
No Thumbnail Available
Name:
Memoire - Meriem Laouar.pdf
Size:
5.98 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
No Thumbnail Available
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed upon to submission
Description: