TR2026-112

Technical Report for MERL’s Real-TSE Challenge Submission


    •  Klement, Dominik, Masuyama, Yoshiki, Boeddeker, Christoph, Saijo, Kohei, Richter, Julius, Wichern, Gordon, Le Roux, Jonathan, "Technical Report for MERL’s Real-TSE Challenge Submission", Tech. Rep. TR2026-112, Mitsubishi Electric Research Laboratories, Cambridge, MA, July 2026.
      BibTeX TR2026-112 PDF
      • @techreport{MERL_TR2026-112,
      • author = {Klement, Dominik; Masuyama, Yoshiki; Boeddeker, Christoph; Saijo, Kohei; Richter, Julius; Wichern, Gordon; Le Roux, Jonathan},
      • title = {Technical Report for MERL’s Real-TSE Challenge Submission},
      • institution = {MERL - Mitsubishi Electric Research Laboratories},
      • address = {Cambridge, MA 02139},
      • number = {TR2026-112},
      • month = jul,
      • year = 2026,
      • url = {https://www.merl.com/publications/TR2026-112/}
      • }
  • MERL Contacts:
  • Research Areas:

    Artificial Intelligence, Machine Learning, Speech & Audio

Abstract:

Target speech extraction (TSE) has largely been dominated by neural network-based approaches trained and evaluated on synthetic fully overlapped data. The Real-TSE
Challenge aims to advance performance on real-world farfield noisy and reverberant recordings. This technical report describes MERL’s submission to the Real-TSE Challenge. Rather than proposing a novel model architecture, we built upon the baseline model and focused primarily on data preparation and cleaning. Our system was trained in four stages, beginning with pre-training on fully overlapped mixtures and simulated multi-talker conversations with noise and reverberation applied to both the mixture and the enrollment utterances. We then adapted the model to real-world conditions using noisy farfield recordings with pseudo-targets derived from processed close-talk microphone signals. Our submission achieved first place in the second track, demonstrating the critical importance of high-quality data preparation. Furthermore, we observed that DNSMOS and speaker similarity are susceptible to overoptimization, motivating an investigation of their robustness using adversarial attacks. The results show that both metrics can be driven to extreme values without degrading the token error rate or the VAD-based F1 score.