Dependable Distributed Training of Compressed Machine Learning Models

Malandrino, F.; Di Giacomo, G.; Levorato, M.; Chiasserini, C. F.

The existing work on the distributed training of machine learning (ML) models has consistently overlooked the distribution of the achieved learning quality, focusing instead on its average value. This leads to a poor dependability of the resulting ML models, whose performance may be much worse than expected. We fill this gap by proposing DepL, a framework for dependable learning orchestration, able to make high-quality, efficient decisions on (i) the data to leverage for learning, (ii) the models to use and when to switch among them, and (iii) the clusters of nodes, and the resources thereof, to exploit. For concreteness, we consider as possible available models a full DNN and its compressed versions. Unlike previous studies, DepL guarantees that a target learning quality is reached with a target probability, while keeping the training cost at a minimum. We prove that DepL has constant competitive ratio and polynomial complexity, and show that it outperforms the state-of-the-art by over 27% and closely matches the optimum.

Dependable Distributed Training of Compressed Machine Learning Models / Malandrino, F.; Di Giacomo, G.; Levorato, M.; Chiasserini, C. F.. - ELETTRONICO. - (2024). (Intervento presentato al convegno IEEE WoWMoM 2024 tenutosi a Perth (Australia) nel June 2024).

Dependable Distributed Training of Compressed Machine Learning Models

F. Malandrino;G. Di Giacomo;M. Levorato;C. F. Chiasserini

2024

Abstract

Scheda breve

Scheda completa

Scheda completa (DC)

Anno del prodotto

2024

Appare nelle tipologie

4.1 Contributo in Atti di convegno

File in questo prodotto:

File	Dimensione	Formato
pipes-5.pdf non disponibili Tipologia: 2. Post-print / Author's Accepted Manuscript Licenza: Non Pubblico - Accesso privato/ristretto Dimensione 2.08 MB Formato Adobe PDF Visualizza/Apri Richiedi una copia	2.08 MB	Adobe PDF	Visualizza/Apri Richiedi una copia

Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11583/2986252

PORTO @ Archivio Istituzionale della Ricerca

Dependable Distributed Training of Compressed Machine Learning Models

F. Malandrino;G. Di Giacomo;M. Levorato;C. F. Chiasserini

2024

Abstract

Scheda breve Scheda completa Scheda completa (DC)

Pubblicazioni consigliate

Informazioni

Conferma cancellazione

Scheda breve

Scheda completa

Scheda completa (DC)