The growing adoption of Artificial Intelligence is driving demand for accurate, robust models, often characterized by deep, large-scale neural networks. While these models achieve high performance, their computational cost leads to significant inference latency, particularly on embedded, resource-constrained platforms such as FPGA-based architectures. This challenge is especially relevant in the space domain, where on-board deep learning is required for data analysis and decision-making under strict constraints on latency and computational resources. To address this challenge, this work proposes a distributed inference architecture based on FPGA accelerators, combining Deep Processing Units with a cluster-oriented execution model over TCP. The approach leverages a high-level development flow based on Vitis AI, enabling deployment without low-level hardware design expertise. Inference is executed across multiple FPGA boards, overcoming the limitations of a single device while maintaining portability and ease of integration. The proposed solution is validated on a cloud-detection task using satellite imagery, representative of realistic on-board scenarios. Experimental results show that the proposed approach achieves a speedup of up to 1.93× over single-board execution, demonstrating its effectiveness for large neural networks and suitability for on-board processing in space applications.

Distributed Deep Learning Inference on DPU-Based FPGA Systems for On-Board Earth Observation / Buccellato, F., Mannini, L., De Sio, C.. - (2026), pp. 235-240. (23rd ACM International Conference on Computing Frontiers: Workshops and Special Sessions Catania (ITA) May 19 - 21, 2026) [10.1145/3801488.3808246].

Distributed Deep Learning Inference on DPU-Based FPGA Systems for On-Board Earth Observation

Federico Buccellato;Corrado De Sio
2026

Abstract

The growing adoption of Artificial Intelligence is driving demand for accurate, robust models, often characterized by deep, large-scale neural networks. While these models achieve high performance, their computational cost leads to significant inference latency, particularly on embedded, resource-constrained platforms such as FPGA-based architectures. This challenge is especially relevant in the space domain, where on-board deep learning is required for data analysis and decision-making under strict constraints on latency and computational resources. To address this challenge, this work proposes a distributed inference architecture based on FPGA accelerators, combining Deep Processing Units with a cluster-oriented execution model over TCP. The approach leverages a high-level development flow based on Vitis AI, enabling deployment without low-level hardware design expertise. Inference is executed across multiple FPGA boards, overcoming the limitations of a single device while maintaining portability and ease of integration. The proposed solution is validated on a cloud-detection task using satellite imagery, representative of realistic on-board scenarios. Experimental results show that the proposed approach achieves a speedup of up to 1.93× over single-board execution, demonstrating its effectiveness for large neural networks and suitability for on-board processing in space applications.
2026
979-8-4007-2569-2
File in questo prodotto:
File Dimensione Formato  
3801488.3808246.pdf

accesso aperto

Tipologia: 2a Post-print versione editoriale / Version of Record
Licenza: Creative commons
Dimensione 713.29 kB
Formato Adobe PDF
713.29 kB Adobe PDF Visualizza/Apri
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11583/3013175