Convolutional Neural Networks (CNNs) are increasingly deployed in safety-critical embedded systems, where reliability against hardware faults is crucial. Traditional fault-tolerance techniques rely on redundancy or fault-aware training, both of which introduce computational or design overhead. This paper presents a zero-cost approach to improve CNN fault tolerance by exploiting implicit sparsity induced through training dynamics. We show that tuning the learning rate during training naturally promotes sparse internal representations that act as fault masks, whose masking probability can be modeled statistically, enhancing resilience without explicit fault injection. We derive and experimentally validate a quantitative model linking sparsity and fault masking, demonstrating a strong linear correlation (r = 0.91) across four CNN architectures and two standard datasets. Highly sparse models achieve up to 90% fault masking relative to standard training, while dynamic batch-size co-scheduling mitigates accuracy loss and preserves robustness. Beyond algorithmic effects, sparsity may reduce Multiply–Accumulate (MAC) operations and switching activity in hardware accelerators, potentially yielding energy efficiency benefits. This work reframes CNN reliability as a property of optimization dynamics rather than a post-training add-on.

Implicit Activation Sparsity for Fault-Tolerant CNNs: A Training-Dynamics Perspective / Bellarmino, N., Bosio, A., Ruospo, A., Ernesto, S., Cantoro, R.. - IEEE International Conference on Omni-Layer Intelligent Systems (COINS):(In corso di stampa). (IEEE International Conference on Omni-Layer Intelligent Systems (COINS) Bologna, Italy 7-9 September 2026).

Implicit Activation Sparsity for Fault-Tolerant CNNs: A Training-Dynamics Perspective

Nicolo, Bellarmino;Alberto, Bosio;Annachiara, Ruospo;Ernesto, Sanchez;Riccardo, Cantoro
In corso di stampa

Abstract

Convolutional Neural Networks (CNNs) are increasingly deployed in safety-critical embedded systems, where reliability against hardware faults is crucial. Traditional fault-tolerance techniques rely on redundancy or fault-aware training, both of which introduce computational or design overhead. This paper presents a zero-cost approach to improve CNN fault tolerance by exploiting implicit sparsity induced through training dynamics. We show that tuning the learning rate during training naturally promotes sparse internal representations that act as fault masks, whose masking probability can be modeled statistically, enhancing resilience without explicit fault injection. We derive and experimentally validate a quantitative model linking sparsity and fault masking, demonstrating a strong linear correlation (r = 0.91) across four CNN architectures and two standard datasets. Highly sparse models achieve up to 90% fault masking relative to standard training, while dynamic batch-size co-scheduling mitigates accuracy loss and preserves robustness. Beyond algorithmic effects, sparsity may reduce Multiply–Accumulate (MAC) operations and switching activity in hardware accelerators, potentially yielding energy efficiency benefits. This work reframes CNN reliability as a property of optimization dynamics rather than a post-training add-on.
In corso di stampa
File in questo prodotto:
File Dimensione Formato  
2026_COINS_LOL_TRAINING_DYNAMICS_CNN (7).pdf

accesso aperto

Tipologia: 2. Post-print / Author's Accepted Manuscript
Licenza: Pubblico - Tutti i diritti riservati
Dimensione 568.33 kB
Formato Adobe PDF
568.33 kB Adobe PDF Visualizza/Apri
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11583/3016154