In a context of rising demand for computing infrastructures, the cost and efficiency of deployed systems become key factors. Application-specific architectures offer a good solution to this challenge, but they are constrained by memory limitations, which affect both performance and energy cost. To mitigate this issue, this work explores the implementation of a cache memory prefetcher to hide the latency of accessing external DRAM. To this end, we propose a hardware design of the Spatial Greedily Accurate SVM-based Prefetcher (SGASP), a member of the GASP family, utilizing High-Level Synthesis (HLS) to significantly simplify the transition from algorithmic definition to hardware logic. The design features a Support Vector Machine (SVM) model optimized for predicting data access patterns using efficient discrete operations. Experimental validation follows a novel two-stage methodology: a first stage to validate our hardware implementation by comparing its output with that executed in a simulator; and a second stage that measures the real performance reached through our prefetcher in an FPGA platform. First, we validate the SGASP prefetcher by comparing the output of the ChampSim implementation with the output of the synthesized hardware simulations, achieving a matching rate exceeding 75% across almost all traces. Second, we integrate a low-cost SGASP as an L2 data prefetcher within an embedded MicroBlaze-based system. When running the most memory-intensive benchmarks in CoreMark-PRO, SGASP with low- and high-cost configurations achieves average speedups of 6.45% and 7.43%, significantly outperforming traditional Next-Line (1.39%) and Stride (4.19%) prefetchers. With a hardware budget utilizing less than 6% of the available resources (specifically, ∼2800 LUTs, ∼ 2100 FFs and 5 BRAMs for low-cost configuration; ∼ 4300 LUTs, ∼ 3200 FFs and 5 BRAMs for high-cost configuration)a hardware budget utilizing under 6% of Look-Up-Tables (LUTs) and Flip-Flops (FFs) on the Ultra96-V2 platform, and consuming less than 3% of total power (0.043 and 0.054 watts for low- and high-cost configuration, resp.), SGASP stands out as a highly feasible, low-overhead solution for accelerating modern FPGA-based architectures.
From Simulation to Speedup: Hardware Implementation of the SGASP Cache Memory Prefetcher through HLS in FPGA Systems / Sánchez-Cuevas, P., Ríos-Navarro, A., Pérez-Peña, A.M., Barocci, M., Díaz-del-Río, F.. - In: JOURNAL OF SYSTEMS ARCHITECTURE. - ISSN 1383-7621. - 181:(2026). [10.1016/j.sysarc.2026.103977]
From Simulation to Speedup: Hardware Implementation of the SGASP Cache Memory Prefetcher through HLS in FPGA Systems
Michelangelo Barocci;
2026
Abstract
In a context of rising demand for computing infrastructures, the cost and efficiency of deployed systems become key factors. Application-specific architectures offer a good solution to this challenge, but they are constrained by memory limitations, which affect both performance and energy cost. To mitigate this issue, this work explores the implementation of a cache memory prefetcher to hide the latency of accessing external DRAM. To this end, we propose a hardware design of the Spatial Greedily Accurate SVM-based Prefetcher (SGASP), a member of the GASP family, utilizing High-Level Synthesis (HLS) to significantly simplify the transition from algorithmic definition to hardware logic. The design features a Support Vector Machine (SVM) model optimized for predicting data access patterns using efficient discrete operations. Experimental validation follows a novel two-stage methodology: a first stage to validate our hardware implementation by comparing its output with that executed in a simulator; and a second stage that measures the real performance reached through our prefetcher in an FPGA platform. First, we validate the SGASP prefetcher by comparing the output of the ChampSim implementation with the output of the synthesized hardware simulations, achieving a matching rate exceeding 75% across almost all traces. Second, we integrate a low-cost SGASP as an L2 data prefetcher within an embedded MicroBlaze-based system. When running the most memory-intensive benchmarks in CoreMark-PRO, SGASP with low- and high-cost configurations achieves average speedups of 6.45% and 7.43%, significantly outperforming traditional Next-Line (1.39%) and Stride (4.19%) prefetchers. With a hardware budget utilizing less than 6% of the available resources (specifically, ∼2800 LUTs, ∼ 2100 FFs and 5 BRAMs for low-cost configuration; ∼ 4300 LUTs, ∼ 3200 FFs and 5 BRAMs for high-cost configuration)a hardware budget utilizing under 6% of Look-Up-Tables (LUTs) and Flip-Flops (FFs) on the Ultra96-V2 platform, and consuming less than 3% of total power (0.043 and 0.054 watts for low- and high-cost configuration, resp.), SGASP stands out as a highly feasible, low-overhead solution for accelerating modern FPGA-based architectures.| File | Dimensione | Formato | |
|---|---|---|---|
|
1-s2.0-S138376212600295X-main.pdf
accesso riservato
Tipologia:
2a Post-print versione editoriale / Version of Record
Licenza:
Non Pubblico - Accesso privato/ristretto
Dimensione
1.77 MB
Formato
Adobe PDF
|
1.77 MB | Adobe PDF | Visualizza/Apri Richiedi una copia |
Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/11583/3015682
