Speech Denoising Using Diffusion Model Technique
No Thumbnail Available
Date
2026
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
university of ghardaia
Abstract
Speech is the primary medium for human communication, yet it is almost
always degraded by background noise in real-world environments. This noise
reduces intelligibility and harms the performance of critical applications such as
automatic speech recognition, telecommunications, and audio streaming. Speech
denoising, therefore, plays a vital role in restoring clean speech and ensuring
reliable human-machine interaction.
This thesis explores diffusion-based generative models for speech denoising,
focusing on the DiffWave architecture. The objective is to build a system that
produces high-quality enhanced speech while operating fast enough for real-world
deployment. To this end, two diffusion step configurations are compared: T=50
and T=200. Although T=200 provides a finer discretization, T=50 achieves better
generalization on unseen data (SI-SDR: 10.93 dB vs 9.94 dB), with an optimal
inference step of t=5. This is primarily because T=50 was trained for significantly
longer (240,000 iterations compared to 65,000 for T=200) due to GPU constraints,
allowing it to converge more effectively.
To overcome inference latency, Denoising Diffusion Implicit Models (DDIM)
are integrated, reducing sampling steps from 50 to just 3. This achieves a real-
time factor (RTF) of 0.0732, corresponding to 13.7× real-time performance, with
a minimal quality loss of only -0.64 dB.
A key contribution is the evaluation on Arabic speech. Although trained
exclusively on English digits, the model achieves its best performance on the
Arabic dataset (SI-SDR = 3.94 dB), with the majority of test samples showing
improvement after processing, confirming strong cross-lingual generalization. This
is attributed to the model’s ability to learn latent acoustic representations that
transcend language boundaries.
The model was also evaluated on French speech using the Common Voice
dataset, demonstrating clear cross-lingual generalization capabilities with promis-
ing results that support its effectiveness in diverse linguistic environments.
The model is also validated on real-world recordings from office, street, and
cafeteria environments, confirming its practical applicability beyond controlled
laboratory conditions. These results demonstrate that the proposed approach
offers a practical, fast, and language-robust solution for speech denoising, suitable
for real-time applications.
Description
Specialty: Intelligent Systems for Knowledge Extraction (SIEC)
Adjila Abderrahmane/encadreur
Keywords
Speech Denoising, Speech Processing, Audio Signal Processing, Deep Learning, Diffusion Models, Speech Enhancement, Signal-to-Noise Ratio (SNR), DDPM, DDIM.
