With the growing adoption of voice assistants and Spoken Language Understanding (SLU) systems, the need for machine unlearning, i.e., the ability to remove the influence of specific training data without full retraining, has become increasingly critical to ensure user privacy and regulatory compliance. In this work, we present UnSLU-BENCH+, an extended benchmark for systematically evaluating unlearning methods in SLU. Our benchmark spans six datasets in four languages, each paired with three self-supervised models, and covers eleven state-of-the-art unlearning techniques. These methods are evaluated based on three foundational principles of unlearning: utility, i.e., the performance of the unlearned model, efficacy, i.e., the actual removal of the information to be forgotten, and efficiency, i.e., the computational cost.To support this evaluation, we introduce CGUM, a classification-aware extension of the Global Unlearning Metric (GUM), which provides a unified, interpretable score that jointly captures these three dimensions. Our extensive experiments demonstrate that unlearning performance varies significantly across models and datasets, with no single method universally dominating. Our results establish a solid foundation for evaluating and developing future machine unlearning techniques in speech applications while promoting CGUM as a robust evaluation metric for every classification-based unlearning scenario.
UnSLU-BENCH+: Extended Machine Unlearning Benchmark for Spoken Language Understanding / Savelli, C., Koudounas, A., Giobergia, F., Baralis, E.M.. - In: IEEE TRANSACTIONS ON AUDIO, SPEECH, AND LANGUAGE PROCESSING. - ISSN 2998-4173. - 34:(2026), pp. 1892-1902. [10.1109/TASLPRO.2026.3675768]
UnSLU-BENCH+: Extended Machine Unlearning Benchmark for Spoken Language Understanding
Savelli Claudio;Koudounas Alkis;Giobergia Flavio;Baralis Elena
2026
Abstract
With the growing adoption of voice assistants and Spoken Language Understanding (SLU) systems, the need for machine unlearning, i.e., the ability to remove the influence of specific training data without full retraining, has become increasingly critical to ensure user privacy and regulatory compliance. In this work, we present UnSLU-BENCH+, an extended benchmark for systematically evaluating unlearning methods in SLU. Our benchmark spans six datasets in four languages, each paired with three self-supervised models, and covers eleven state-of-the-art unlearning techniques. These methods are evaluated based on three foundational principles of unlearning: utility, i.e., the performance of the unlearned model, efficacy, i.e., the actual removal of the information to be forgotten, and efficiency, i.e., the computational cost.To support this evaluation, we introduce CGUM, a classification-aware extension of the Global Unlearning Metric (GUM), which provides a unified, interpretable score that jointly captures these three dimensions. Our extensive experiments demonstrate that unlearning performance varies significantly across models and datasets, with no single method universally dominating. Our results establish a solid foundation for evaluating and developing future machine unlearning techniques in speech applications while promoting CGUM as a robust evaluation metric for every classification-based unlearning scenario.| File | Dimensione | Formato | |
|---|---|---|---|
|
UnSLU-BENCH_Extended_Machine_Unlearning_Benchmark_for_Spoken_Language_Understanding (1).pdf
accesso aperto
Tipologia:
2a Post-print versione editoriale / Version of Record
Licenza:
Creative commons
Dimensione
1.77 MB
Formato
Adobe PDF
|
1.77 MB | Adobe PDF | Visualizza/Apri |
Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/11583/3015357
