Music Source Separation (MSS) models based on Deep Neural Networks have reached state-of-the-art performance in splitting stereo mixtures into separate sources. However, little attention has been paid to how well these models preserve the spatial properties of each separated source. Motivated by applications in virtual/augmented reality and accessibility, we propose Spatially-Aware HT-Demucs (SA-HTDemucs), a novel deep learning framework that extends HT-Demucs with a per-source SpatialCueModule, which corrects the Interaural Level Difference (ILD) in the time-frequency domain to preserve the azimuth panning of each estimated source. The model is trained and evaluated on binauralMUSDB18-HQ, a novel binaural dataset synthesized by convolving each source stem in MUSDB18-HQ with open-source Head-Related Transfer Functions (HRTFs) at randomized horizontal positions. Results show that SA-HTDemucs consistently improves spatial location preservation across all sources compared to the HT-Demucs baseline, opening promising directions at the intersection of MSS and immersive audio.
Preserving Spatial Information in Music Source Separation
M. Acerbi;M. Pezzoli;F. Antonacci;
2026-01-01
Abstract
Music Source Separation (MSS) models based on Deep Neural Networks have reached state-of-the-art performance in splitting stereo mixtures into separate sources. However, little attention has been paid to how well these models preserve the spatial properties of each separated source. Motivated by applications in virtual/augmented reality and accessibility, we propose Spatially-Aware HT-Demucs (SA-HTDemucs), a novel deep learning framework that extends HT-Demucs with a per-source SpatialCueModule, which corrects the Interaural Level Difference (ILD) in the time-frequency domain to preserve the azimuth panning of each estimated source. The model is trained and evaluated on binauralMUSDB18-HQ, a novel binaural dataset synthesized by convolving each source stem in MUSDB18-HQ with open-source Head-Related Transfer Functions (HRTFs) at randomized horizontal positions. Results show that SA-HTDemucs consistently improves spatial location preservation across all sources compared to the HT-Demucs baseline, opening promising directions at the intersection of MSS and immersive audio.| File | Dimensione | Formato | |
|---|---|---|---|
|
Preserving_Spatial_Information_in_Music_Source_Separation.pdf
accesso aperto
:
Post-Print (DRAFT o Author’s Accepted Manuscript-AAM)
Dimensione
305.61 kB
Formato
Adobe PDF
|
305.61 kB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



