Despite the effectiveness of cosine scoring compared to classical Probabilistic Linear Discriminant Analysis (PLDA), structured generative models such as Toroidal Probabilistic Spherical Discriminant Analysis (T-PSDA) can still provide superior performance for DNN-based speaker verification. In this work we investigate the relationship between T-PSDA and an isotropic simplified PLDA (ISO-PLDA) model, and propose a novel approach that combines the ISO-PLDA Gaussian likelihood with the structured directional T-PSDA speaker characterization to obtain a Spherical-Gaussian (SG-TPSDA) model that is able to match and, in some cases, improve T-PSDA performance. Furthermore, we show how to extend both ISO-PLDA and SG-TPSDA to explicitly account for utterance duration. Our results on NIST and CN-Celeb data show that the proposed approach is effective and capable of leveraging duration information to further improve speaker verification accuracy.

Spherical-Gaussian TPSDA: combining PLDA, T-PSDA and duration models for speaker verification / Cumani, S.. - ELETTRONICO. - (2026), pp. 284-291. (Odyssey 2026 - The Speaker and Language Recognition Workshop Lisbon (POR) 23 - 26 June 2026) [10.21437/odyssey.2026-42].

Spherical-Gaussian TPSDA: combining PLDA, T-PSDA and duration models for speaker verification

Cumani, Sandro
2026

Abstract

Despite the effectiveness of cosine scoring compared to classical Probabilistic Linear Discriminant Analysis (PLDA), structured generative models such as Toroidal Probabilistic Spherical Discriminant Analysis (T-PSDA) can still provide superior performance for DNN-based speaker verification. In this work we investigate the relationship between T-PSDA and an isotropic simplified PLDA (ISO-PLDA) model, and propose a novel approach that combines the ISO-PLDA Gaussian likelihood with the structured directional T-PSDA speaker characterization to obtain a Spherical-Gaussian (SG-TPSDA) model that is able to match and, in some cases, improve T-PSDA performance. Furthermore, we show how to extend both ISO-PLDA and SG-TPSDA to explicitly account for utterance duration. Our results on NIST and CN-Celeb data show that the proposed approach is effective and capable of leveraging duration information to further improve speaker verification accuracy.
File in questo prodotto:
File Dimensione Formato  
cumani26_odyssey.pdf

accesso aperto

Tipologia: 2a Post-print versione editoriale / Version of Record
Licenza: Pubblico - Tutti i diritti riservati
Dimensione 314.02 kB
Formato Adobe PDF
314.02 kB Adobe PDF Visualizza/Apri
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11583/3013364