With the growing adoption of voice assistants and Spoken Language Understanding (SLU) systems, the need for machine unlearning, i.e., the ability to remove the influence of specific training data without full retraining, has become increasingly critical to ensure user privacy and regulatory compliance. In this work, we present UnSLU-BENCH+, an extended benchmark for systematically evaluating unlearning methods in SLU. Our benchmark spans six datasets in four languages, each paired with three self-supervised models, and covers eleven state-of-the-art unlearning techniques. These methods are evaluated based on three foundational principles of unlearning: utility, i.e., the performance of the unlearned model, efficacy, i.e., the actual removal of the information to be forgotten, and efficiency, i.e., the computational cost.To support this evaluation, we introduce CGUM, a classification-aware extension of the Global Unlearning Metric (GUM), which provides a unified, interpretable score that jointly captures these three dimensions. Our extensive experiments demonstrate that unlearning performance varies significantly across models and datasets, with no single method universally dominating. Our results establish a solid foundation for evaluating and developing future machine unlearning techniques in speech applications while promoting CGUM as a robust evaluation metric for every classification-based unlearning scenario.

UnSLU-BENCH+: Extended Machine Unlearning Benchmark for Spoken Language Understanding / Savelli, C., Koudounas, A., Giobergia, F., Baralis, E.M.. - In: IEEE TRANSACTIONS ON AUDIO, SPEECH, AND LANGUAGE PROCESSING. - ISSN 2998-4173. - 34:(2026), pp. 1892-1902. [10.1109/TASLPRO.2026.3675768]

UnSLU-BENCH+: Extended Machine Unlearning Benchmark for Spoken Language Understanding

Savelli Claudio;Koudounas Alkis;Giobergia Flavio;Baralis Elena
2026

Abstract

With the growing adoption of voice assistants and Spoken Language Understanding (SLU) systems, the need for machine unlearning, i.e., the ability to remove the influence of specific training data without full retraining, has become increasingly critical to ensure user privacy and regulatory compliance. In this work, we present UnSLU-BENCH+, an extended benchmark for systematically evaluating unlearning methods in SLU. Our benchmark spans six datasets in four languages, each paired with three self-supervised models, and covers eleven state-of-the-art unlearning techniques. These methods are evaluated based on three foundational principles of unlearning: utility, i.e., the performance of the unlearned model, efficacy, i.e., the actual removal of the information to be forgotten, and efficiency, i.e., the computational cost.To support this evaluation, we introduce CGUM, a classification-aware extension of the Global Unlearning Metric (GUM), which provides a unified, interpretable score that jointly captures these three dimensions. Our extensive experiments demonstrate that unlearning performance varies significantly across models and datasets, with no single method universally dominating. Our results establish a solid foundation for evaluating and developing future machine unlearning techniques in speech applications while promoting CGUM as a robust evaluation metric for every classification-based unlearning scenario.
File in questo prodotto:
File Dimensione Formato  
UnSLU-BENCH_Extended_Machine_Unlearning_Benchmark_for_Spoken_Language_Understanding (1).pdf

accesso aperto

Tipologia: 2a Post-print versione editoriale / Version of Record
Licenza: Creative commons
Dimensione 1.77 MB
Formato Adobe PDF
1.77 MB Adobe PDF Visualizza/Apri
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11583/3015357