Contribution

Embedding-Based Intrusive Evaluation Metrics for Musical Source Separation Using MERT Representations

* Presenting author
Day / Time: 26.03.2026, 16:40-17:00
Manuscript: PDF-Download
Type: Vortrag (strukturierte Sitzung)
Abstract ID: DAGA2026/429
Abstract: Objective evaluation of musical source separation (MSS) has traditionally relied on Blind Source Separation Evaluation (BSS-Eval) metrics. However, recent work in singing voice separation suggests that BSS-Eval metrics exhibit shortcomings for evaluating generative models. A reasonable alternative is offered by embedding-based intrusive metrics that leverage latent embeddings of large autoencoder models trained in a self-supervised manner, e.g. the Music undERstanding model with large-scale self-supervised Training (MERT). These embedding-based metrics exhibit higher correlation with human perceptual quality scores for evaluating both discriminative and generative singing voice separation models. In this study, we investigate whether the mean square error of separated and reference signal in the MERT-L12 embedding space (MERT-L12 MSE) serves as a robust evaluation metric for musical source separation data, beyond vocal-only separation. We evaluate MERT-L12 MSE on an existing dataset consisting of separated musical stems (vocals, drums, bass, and other instruments) generated by various discriminative four-stem MSS models, accompanied by perceptual audio quality ratings. We analyse correlations between the embedding-based metric and perceptual audio quality ratings. The findings of this study will demonstrate whether MERT-L12 MSE encapsulates separation quality beyond separated vocals and can serve as a reliable perceptually aligned evaluation metric for four-stem musical source separation models.