Latest semiconductor technologies are increasingly sensitive to faults, which can compromise the reliability of Data Centers and High-Performance Computing (HPC) systems when running compute-intensive workloads. To excite and detect in-field permanent faults, stress-oriented routines are periodically executed before such faults manifest as critical failures during standard workload execution. However, state-of-the-art stress tests typically focus on available applications (e.g., matrix multiplication) and overlook latent faults that emerge only under specific computational patterns. Designing ad-hoc applications capable of triggering these faults requires substantial manual effort and deep knowledge of hardware architecture. To address these limitations, we present an empirical study that, for the first time, assesses the ability of Large Language Models (LLMs) to collaborate within a multi-agent iterative framework to automatically generate GPU-stress-testing routines. Experimental results show that four of the seven evaluated LLMs consistently produce compilable and executable code, enabling the automated development of stress-test applications without human intervention and significantly reducing the time needed to obtain a functional prototype compared to manual development. Among the tested LLMs, two, used as the core LLMs for all agents, generated two stress-testing applications that were 80.35% and 37.51% more compact than GPU-Burn (a state-of-the-art stress-testing tool for GPUs), yet still reached 89.9% of GPU-Burn's peak temperature and 77.2% of its energy consumption.

From Prompts to Pressure: Evaluating LLM-driven Agents for GPU Stress-code Generation / Gensale, A., Esposito, G., Guerrero-Balaguera, J., Condia, J.E.R., Cagliero, L., Reorda, M.S.. - (2026), pp. 1-7. (32nd International Symposium on On-Line Testing and Robust System Design, IOLTS 2026 Polignano a Mare (ITA) 01-03 July 2026) [10.1109/iolts69666.2026.11633814].

From Prompts to Pressure: Evaluating LLM-driven Agents for GPU Stress-code Generation

Gensale, Aurora;Esposito, Giuseppe;Guerrero-Balaguera, Juan-David;Cagliero, Luca;Reorda, Matteo Sonza
2026

Abstract

Latest semiconductor technologies are increasingly sensitive to faults, which can compromise the reliability of Data Centers and High-Performance Computing (HPC) systems when running compute-intensive workloads. To excite and detect in-field permanent faults, stress-oriented routines are periodically executed before such faults manifest as critical failures during standard workload execution. However, state-of-the-art stress tests typically focus on available applications (e.g., matrix multiplication) and overlook latent faults that emerge only under specific computational patterns. Designing ad-hoc applications capable of triggering these faults requires substantial manual effort and deep knowledge of hardware architecture. To address these limitations, we present an empirical study that, for the first time, assesses the ability of Large Language Models (LLMs) to collaborate within a multi-agent iterative framework to automatically generate GPU-stress-testing routines. Experimental results show that four of the seven evaluated LLMs consistently produce compilable and executable code, enabling the automated development of stress-test applications without human intervention and significantly reducing the time needed to obtain a functional prototype compared to manual development. Among the tested LLMs, two, used as the core LLMs for all agents, generated two stress-testing applications that were 80.35% and 37.51% more compact than GPU-Burn (a state-of-the-art stress-testing tool for GPUs), yet still reached 89.9% of GPU-Burn's peak temperature and 77.2% of its energy consumption.
2026
979-8-3315-4685-4
File in questo prodotto:
File Dimensione Formato  
From_Prompts_to_Pressure_Evaluating_LLM-driven_Agents_for_GPU_Stress-code_Generation.pdf

accesso riservato

Tipologia: 2a Post-print versione editoriale / Version of Record
Licenza: Non Pubblico - Accesso privato/ristretto
Dimensione 1.23 MB
Formato Adobe PDF
1.23 MB Adobe PDF   Visualizza/Apri   Richiedi una copia
_IOLTS2026_GPU_COMA-1.pdf

accesso aperto

Tipologia: 2. Post-print / Author's Accepted Manuscript
Licenza: Pubblico - Tutti i diritti riservati
Dimensione 290.4 kB
Formato Adobe PDF
290.4 kB Adobe PDF Visualizza/Apri
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11583/3015047