The rapid evolution of Machine Learning (ML) workloads, particularly Deep Neural Networks (DNNs) and Transformer-based models, has intensified demands on computing architectures, highlighting the limitations of traditional von Neumann systems due to the memory bottleneck. To address these challenges, this paper investigates the mapping of fundamental Machine Learning (ML) operations onto ARCANE, a Near-Memory Computing (NMC)-based architecture that integrates Vector Processing Units (VPUs) directly within the data cache. ARCANE offers a flexible ISA-extension (xmnmc) abstracting memory management, effectively reducing data movement and enhancing performance. We specifically explore the acceleration capabilities of ARCANE when executing fundamental Deep Neural Network (DNN) and Transformer-based operations. Experimental results show that, with a contained area overhead, ARCANE achieves consistent speedups, delivering up to 150x improvement in 2D convolution, 305x in Linear layer, and over 32x in Fused-Weight Self-Attention (FWSA), compared to conventional CPU approaches. These findings underline ARCANE's significant benefits in supporting efficient deployment of edge-oriented Machine Learning (ML) workloads.

Stop Wasting your Cache! Bringing Machine Learning into Cache Computing / Petrolo, V., Guella, F., Caon, M., Masera, G., Martina, M.. - 2:(2025), pp. 86-89. (22nd ACM International Conference on Computing Frontiers 2025, CF 2025 ita 2025) [10.1145/3706594.3726983].

Stop Wasting your Cache! Bringing Machine Learning into Cache Computing

Petrolo, Vincenzo;Guella, Flavia;Caon, Michele;Masera, Guido;Martina, Maurizio
2025

Abstract

The rapid evolution of Machine Learning (ML) workloads, particularly Deep Neural Networks (DNNs) and Transformer-based models, has intensified demands on computing architectures, highlighting the limitations of traditional von Neumann systems due to the memory bottleneck. To address these challenges, this paper investigates the mapping of fundamental Machine Learning (ML) operations onto ARCANE, a Near-Memory Computing (NMC)-based architecture that integrates Vector Processing Units (VPUs) directly within the data cache. ARCANE offers a flexible ISA-extension (xmnmc) abstracting memory management, effectively reducing data movement and enhancing performance. We specifically explore the acceleration capabilities of ARCANE when executing fundamental Deep Neural Network (DNN) and Transformer-based operations. Experimental results show that, with a contained area overhead, ARCANE achieves consistent speedups, delivering up to 150x improvement in 2D convolution, 305x in Linear layer, and over 32x in Fused-Weight Self-Attention (FWSA), compared to conventional CPU approaches. These findings underline ARCANE's significant benefits in supporting efficient deployment of edge-oriented Machine Learning (ML) workloads.
File in questo prodotto:
Non ci sono file associati a questo prodotto.
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11583/3015603
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo