Joint Audio Separation and Noise Suppression
- •A joint network separates overlapping speakers and suppresses changing noise from one shared audio representation.
- •On 3,000 mixtures, the model scored 15.60 ± 0.10 dB SI-SNR improvement, 3.02 PESQ, and 0.917 STOI at 0 dB.
- •TF-GridNet and MossFormer2 scored 0.50 to 0.80 dB higher at seven to eleven times the arithmetic cost.
Researchers at Hunan University of Information Technology developed a deep learning network that separates overlapping speakers and suppresses changing background noise together. The system processes both tasks from one shared audio representation, using a learnable encoder, a multi-scale separator with channel and time-frequency attention, a noise-suppression module for each separated stream, and a reconstructor. It is trained end to end with a three-part objective combining scale-invariant signal-to-noise ratio, spectral magnitude, and perceptual weighting. Training includes audio with input signal-to-noise ratios as low as −5 dB.
Tests used LibriMix, WHAM! and WHAMR! for source separation, VoiceBank-DEMAND and MUSAN for noise suppression, and MUSDB18-HQ to check full-band audio beyond speech. Across 3,000 mixtures and three random seeds, and against seven baselines retrained with the same recipe, the model achieved a 15.60 ± 0.10 dB improvement in scale-invariant signal-to-noise ratio, a PESQ score of 3.02, and a STOI score of 0.917 at 0 dB input SNR. It has 3.9 million parameters and outperformed Conv-TasNet, DPRNN, SepFormer, a diffusion-based cascade, and a pipeline that separates audio before denoising.
TF-GridNet and MossFormer2 scored 0.50 to 0.80 dB higher, but required seven to eleven times the arithmetic cost. The proposed system used 6.8 G multiply–accumulate operations per second of audio, reached a real-time factor of 0.51 on one CPU core, and consumed 0.41 J per second of processed audio. A causal version for streaming gave up 1.70 dB. Ablation tests attributed 4.50% and 7.10% to the two attention paths and 10.30% to the joint denoising route. Cross-corpus, unseen-language, and reverberant tests found transfer losses of 0.70 to 2.20 dB; improvement fell to 7.80 dB when reverberation time exceeded 0.90 s. Code, configurations, and per-utterance records are available with the study.