The shift towards edge computing, coupled with the increasing adoption of Machine Learning (ML) for data-intensive tasks, imposes demanding performance requirements on embedded systems, where employing dedicated accelerators is often impractical due to tight power and area budgets. To address these challenges, we present ANT-V, a lightweight RISC-V vector processor designed to efficiently execute data-parallel workloads within the tight resource budgets of edge platforms. ANT-V integrates a selected subset of the Zve32x vector extension into the open-source CVE2 core, adopting a monolithic, zero-duplication approach. To minimize hardware overhead, vector instructions execute as controller-driven micro-sequences over the scalar datapath, avoiding duplication of functional units. To eliminate the cost of a dedicated Vector Register File, vector registers are mapped to a software-reconfigurable region of the on-chip SRAM, trading raw bandwidth for area efficiency and runtime flexibility. Post-layout evaluation in 65nm CMOS shows negligible system-level area overhead, with up to 17x higher throughput and 16x better energy efficiency compared to the scalar baseline on ML kernels. These results position ANT-V as an effective compromise, enabling faster edge ML inference on general-purpose embedded systems without the additional complexity of specialized hardware.
ANT-V: A Monolithic RISC-V Vector Processor for Machine Learning at the Edge / Guella, F., Caviglia, A., Caon, M., Carlo, S.D., Masera, G., Martina, M.. - In: IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS. II, EXPRESS BRIEFS. - ISSN 1549-7747. - (2026), pp. 1-1. [10.1109/tcsii.2026.3733550]
ANT-V: A Monolithic RISC-V Vector Processor for Machine Learning at the Edge
Guella, Flavia;Caviglia, Alessio;Caon, Michele;Carlo, Stefano Di;Masera, Guido;Martina, Maurizio
2026
Abstract
The shift towards edge computing, coupled with the increasing adoption of Machine Learning (ML) for data-intensive tasks, imposes demanding performance requirements on embedded systems, where employing dedicated accelerators is often impractical due to tight power and area budgets. To address these challenges, we present ANT-V, a lightweight RISC-V vector processor designed to efficiently execute data-parallel workloads within the tight resource budgets of edge platforms. ANT-V integrates a selected subset of the Zve32x vector extension into the open-source CVE2 core, adopting a monolithic, zero-duplication approach. To minimize hardware overhead, vector instructions execute as controller-driven micro-sequences over the scalar datapath, avoiding duplication of functional units. To eliminate the cost of a dedicated Vector Register File, vector registers are mapped to a software-reconfigurable region of the on-chip SRAM, trading raw bandwidth for area efficiency and runtime flexibility. Post-layout evaluation in 65nm CMOS shows negligible system-level area overhead, with up to 17x higher throughput and 16x better energy efficiency compared to the scalar baseline on ML kernels. These results position ANT-V as an effective compromise, enabling faster edge ML inference on general-purpose embedded systems without the additional complexity of specialized hardware.| File | Dimensione | Formato | |
|---|---|---|---|
|
ANT-V_A_Monolithic_RISC-V_Vector_Processor_for_Machine_Learning_at_the_Edge.pdf
accesso riservato
Tipologia:
2. Post-print / Author's Accepted Manuscript
Licenza:
Non Pubblico - Accesso privato/ristretto
Dimensione
671.33 kB
Formato
Adobe PDF
|
671.33 kB | Adobe PDF | Visualizza/Apri Richiedi una copia |
|
ANT_V_final (1) (1).pdf
accesso aperto
Descrizione: Pre-print version
Tipologia:
1. Preprint / submitted version [pre- review]
Licenza:
Pubblico - Tutti i diritti riservati
Dimensione
609.56 kB
Formato
Adobe PDF
|
609.56 kB | Adobe PDF | Visualizza/Apri |
Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/11583/3015893
