Music Source Separation (MSS) models based on Deep Neural Networks have reached state-of-the-art performance in splitting stereo mixtures into separate sources. However, little attention has been paid to how well these models preserve the spatial properties of each separated source. Motivated by applications in virtual/augmented reality and accessibility, we propose Spatially-Aware HT-Demucs (SA-HTDemucs), a novel deep learning framework that extends HT-Demucs with a per-source SpatialCueModule, which corrects the Interaural Level Difference (ILD) in the time-frequency domain to preserve the azimuth panning of each estimated source. The model is trained and evaluated on binauralMUSDB18-HQ, a novel binaural dataset synthesized by convolving each source stem in MUSDB18-HQ with open-source Head-Related Transfer Functions (HRTFs) at randomized horizontal positions. Results show that SA-HTDemucs consistently improves spatial location preservation across all sources compared to the HT-Demucs baseline, opening promising directions at the intersection of MSS and immersive audio.

Preserving Spatial Information in Music Source Separation

M. Acerbi;M. Pezzoli;F. Antonacci;
2026-01-01

Abstract

Music Source Separation (MSS) models based on Deep Neural Networks have reached state-of-the-art performance in splitting stereo mixtures into separate sources. However, little attention has been paid to how well these models preserve the spatial properties of each separated source. Motivated by applications in virtual/augmented reality and accessibility, we propose Spatially-Aware HT-Demucs (SA-HTDemucs), a novel deep learning framework that extends HT-Demucs with a per-source SpatialCueModule, which corrects the Interaural Level Difference (ILD) in the time-frequency domain to preserve the azimuth panning of each estimated source. The model is trained and evaluated on binauralMUSDB18-HQ, a novel binaural dataset synthesized by convolving each source stem in MUSDB18-HQ with open-source Head-Related Transfer Functions (HRTFs) at randomized horizontal positions. Results show that SA-HTDemucs consistently improves spatial location preservation across all sources compared to the HT-Demucs baseline, opening promising directions at the intersection of MSS and immersive audio.
2026
International Workshop on Acoustic Signal Enhancement (IWAENC 2026)
Music Source Separation, Binaural Audio, Interaural Level Difference, Spatial Cue Preservation, Deep Neural Networks
File in questo prodotto:
File Dimensione Formato  
Preserving_Spatial_Information_in_Music_Source_Separation.pdf

accesso aperto

: Post-Print (DRAFT o Author’s Accepted Manuscript-AAM)
Dimensione 305.61 kB
Formato Adobe PDF
305.61 kB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11311/1324389
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
  • OpenAlex ND
social impact