Speech Denoising Using Diffusion Model Technique

No Thumbnail Available

Date

2026

Journal Title

Journal ISSN

Volume Title

Publisher

university of ghardaia

Abstract

Speech is the primary medium for human communication, yet it is almost always degraded by background noise in real-world environments. This noise reduces intelligibility and harms the performance of critical applications such as automatic speech recognition, telecommunications, and audio streaming. Speech denoising, therefore, plays a vital role in restoring clean speech and ensuring reliable human-machine interaction. This thesis explores diffusion-based generative models for speech denoising, focusing on the DiffWave architecture. The objective is to build a system that produces high-quality enhanced speech while operating fast enough for real-world deployment. To this end, two diffusion step configurations are compared: T=50 and T=200. Although T=200 provides a finer discretization, T=50 achieves better generalization on unseen data (SI-SDR: 10.93 dB vs 9.94 dB), with an optimal inference step of t=5. This is primarily because T=50 was trained for significantly longer (240,000 iterations compared to 65,000 for T=200) due to GPU constraints, allowing it to converge more effectively. To overcome inference latency, Denoising Diffusion Implicit Models (DDIM) are integrated, reducing sampling steps from 50 to just 3. This achieves a real- time factor (RTF) of 0.0732, corresponding to 13.7× real-time performance, with a minimal quality loss of only -0.64 dB. A key contribution is the evaluation on Arabic speech. Although trained exclusively on English digits, the model achieves its best performance on the Arabic dataset (SI-SDR = 3.94 dB), with the majority of test samples showing improvement after processing, confirming strong cross-lingual generalization. This is attributed to the model’s ability to learn latent acoustic representations that transcend language boundaries. The model was also evaluated on French speech using the Common Voice dataset, demonstrating clear cross-lingual generalization capabilities with promis- ing results that support its effectiveness in diverse linguistic environments. The model is also validated on real-world recordings from office, street, and cafeteria environments, confirming its practical applicability beyond controlled laboratory conditions. These results demonstrate that the proposed approach offers a practical, fast, and language-robust solution for speech denoising, suitable for real-time applications.

Description

Specialty: Intelligent Systems for Knowledge Extraction (SIEC) Adjila Abderrahmane/encadreur

Keywords

Speech Denoising, Speech Processing, Audio Signal Processing, Deep Learning, Diffusion Models, Speech Enhancement, Signal-to-Noise Ratio (SNR), DDPM, DDIM.

Citation

Endorsement

Review

Supplemented By

Referenced By