Reconfigurable SoCs are widely adopted for deploying deep neural networks in embedded environments, due to their balance between performance, cost, and flexibility. In this context, AMD Deep Learning Processing Units (DPUs) provide efficient acceleration for a wide range of models, including safety-critical applications where reliable inference is essential. However, soft errors in FPGA configuration memory remain a concern, as radiation-induced faults may silently corrupt the hardware and alter model inference. Despite this vulnerability, most existing studies focus on application-level robustness, typically evaluated offline, while lightweight runtime detection solutions are still limited. This work proposes a runtime anomaly detection system for DPU-accelerated neural networks, based on monitoring model outputs using metrics calibrated through hardware-level fault injection. The approach combines per-inference decisions with a window-based verification mechanism, reducing false positives while maintaining sensitivity to persistent faults. Experimental results on multiple models show effective generalization, improving reliability of FPGA-based deployments with limited overhead, making it suitable for safety-critical embedded applications.
Hardware-Aware Runtime Detection of Soft-Error Anomalies in DPU-Accelerated Neural Networks / Buccellato, F., De Sio, C., Azimi, S., Sterpone, L.. - (In corso di stampa). (32nd IEEE International Symposium on On-Line Testing and Robust System Design Polignano a Mare, Italy 1-3 July 2026).
Hardware-Aware Runtime Detection of Soft-Error Anomalies in DPU-Accelerated Neural Networks
Federico Buccellato;Corrado De Sio;Sarah Azimi;Luca Sterpone
In corso di stampa
Abstract
Reconfigurable SoCs are widely adopted for deploying deep neural networks in embedded environments, due to their balance between performance, cost, and flexibility. In this context, AMD Deep Learning Processing Units (DPUs) provide efficient acceleration for a wide range of models, including safety-critical applications where reliable inference is essential. However, soft errors in FPGA configuration memory remain a concern, as radiation-induced faults may silently corrupt the hardware and alter model inference. Despite this vulnerability, most existing studies focus on application-level robustness, typically evaluated offline, while lightweight runtime detection solutions are still limited. This work proposes a runtime anomaly detection system for DPU-accelerated neural networks, based on monitoring model outputs using metrics calibrated through hardware-level fault injection. The approach combines per-inference decisions with a window-based verification mechanism, reducing false positives while maintaining sensitivity to persistent faults. Experimental results on multiple models show effective generalization, improving reliability of FPGA-based deployments with limited overhead, making it suitable for safety-critical embedded applications.| File | Dimensione | Formato | |
|---|---|---|---|
|
IOLTS_2026.pdf
accesso riservato
Tipologia:
2a Post-print versione editoriale / Version of Record
Licenza:
Non Pubblico - Accesso privato/ristretto
Dimensione
542.82 kB
Formato
Adobe PDF
|
542.82 kB | Adobe PDF | Visualizza/Apri Richiedi una copia |
Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/11583/3013188
