Reconfigurable SoCs are widely adopted for deploying deep neural networks in embedded environments, due to their balance between performance, cost, and flexibility. In this context, AMD Deep Learning Processing Units (DPUs) provide efficient acceleration for a wide range of models, including safety-critical applications where reliable inference is essential. However, soft errors in FPGA configuration memory remain a concern, as radiation-induced faults may silently corrupt the hardware and alter model inference. Despite this vulnerability, most existing studies focus on application-level robustness, typically evaluated offline, while lightweight runtime detection solutions are still limited. This work proposes a runtime anomaly detection system for DPU-accelerated neural networks, based on monitoring model outputs using metrics calibrated through hardware-level fault injection. The approach combines per-inference decisions with a window-based verification mechanism, reducing false positives while maintaining sensitivity to persistent faults. Experimental results on multiple models show effective generalization, improving reliability of FPGA-based deployments with limited overhead, making it suitable for safety-critical embedded applications.

Hardware-Aware Runtime Detection of Soft-Error Anomalies in DPU-Accelerated Neural Networks / Buccellato, F., De Sio, C., Azimi, S., Sterpone, L.. - (In corso di stampa). (32nd IEEE International Symposium on On-Line Testing and Robust System Design Polignano a Mare, Italy 1-3 July 2026).

Hardware-Aware Runtime Detection of Soft-Error Anomalies in DPU-Accelerated Neural Networks

Federico Buccellato;Corrado De Sio;Sarah Azimi;Luca Sterpone
In corso di stampa

Abstract

Reconfigurable SoCs are widely adopted for deploying deep neural networks in embedded environments, due to their balance between performance, cost, and flexibility. In this context, AMD Deep Learning Processing Units (DPUs) provide efficient acceleration for a wide range of models, including safety-critical applications where reliable inference is essential. However, soft errors in FPGA configuration memory remain a concern, as radiation-induced faults may silently corrupt the hardware and alter model inference. Despite this vulnerability, most existing studies focus on application-level robustness, typically evaluated offline, while lightweight runtime detection solutions are still limited. This work proposes a runtime anomaly detection system for DPU-accelerated neural networks, based on monitoring model outputs using metrics calibrated through hardware-level fault injection. The approach combines per-inference decisions with a window-based verification mechanism, reducing false positives while maintaining sensitivity to persistent faults. Experimental results on multiple models show effective generalization, improving reliability of FPGA-based deployments with limited overhead, making it suitable for safety-critical embedded applications.
In corso di stampa
File in questo prodotto:
File Dimensione Formato  
IOLTS_2026.pdf

accesso riservato

Tipologia: 2a Post-print versione editoriale / Version of Record
Licenza: Non Pubblico - Accesso privato/ristretto
Dimensione 542.82 kB
Formato Adobe PDF
542.82 kB Adobe PDF   Visualizza/Apri   Richiedi una copia
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11583/3013188