Despite the effectiveness of cosine scoring compared to classical Probabilistic Linear Discriminant Analysis (PLDA), structured generative models such as Toroidal Probabilistic Spherical Discriminant Analysis (T-PSDA) can still provide superior performance for DNN-based speaker verification. In this work we investigate the relationship between T-PSDA and an isotropic simplified PLDA (ISO-PLDA) model, and propose a novel approach that combines the ISO-PLDA Gaussian likelihood with the structured directional T-PSDA speaker characterization to obtain a Spherical-Gaussian (SG-TPSDA) model that is able to match and, in some cases, improve T-PSDA performance. Furthermore, we show how to extend both ISO-PLDA and SG-TPSDA to explicitly account for utterance duration. Our results on NIST and CN-Celeb data show that the proposed approach is effective and capable of leveraging duration information to further improve speaker verification accuracy.
Spherical-Gaussian TPSDA: combining PLDA, T-PSDA and duration models for speaker verification / Cumani, S.. - ELETTRONICO. - (2026), pp. 284-291. (Odyssey 2026 - The Speaker and Language Recognition Workshop Lisbon (POR) 23 - 26 June 2026) [10.21437/odyssey.2026-42].
Spherical-Gaussian TPSDA: combining PLDA, T-PSDA and duration models for speaker verification
Cumani, Sandro
2026
Abstract
Despite the effectiveness of cosine scoring compared to classical Probabilistic Linear Discriminant Analysis (PLDA), structured generative models such as Toroidal Probabilistic Spherical Discriminant Analysis (T-PSDA) can still provide superior performance for DNN-based speaker verification. In this work we investigate the relationship between T-PSDA and an isotropic simplified PLDA (ISO-PLDA) model, and propose a novel approach that combines the ISO-PLDA Gaussian likelihood with the structured directional T-PSDA speaker characterization to obtain a Spherical-Gaussian (SG-TPSDA) model that is able to match and, in some cases, improve T-PSDA performance. Furthermore, we show how to extend both ISO-PLDA and SG-TPSDA to explicitly account for utterance duration. Our results on NIST and CN-Celeb data show that the proposed approach is effective and capable of leveraging duration information to further improve speaker verification accuracy.| File | Dimensione | Formato | |
|---|---|---|---|
|
cumani26_odyssey.pdf
accesso aperto
Tipologia:
2a Post-print versione editoriale / Version of Record
Licenza:
Pubblico - Tutti i diritti riservati
Dimensione
314.02 kB
Formato
Adobe PDF
|
314.02 kB | Adobe PDF | Visualizza/Apri |
Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/11583/3013364
