Few-shot learning (FSL) models for bearing fault diagnosis often require meta-training on dedicated datasets and are limited by the scarcity of fault data in real industrial scenarios. Multimodal Large Language Models (MLLMs) offer an alternative, enabling few-shot classification via prompting on time-frequency images of vibration signals. This work studies medium-to-large bearings tested on a dedicated test rig and analyzes (i) when MLLM performance is statistically superior to a Prototypical Network FSL baseline, and (ii) how prompt design affects few-shot classification. On a 4-way task with Continuous Wavelet Transform (CWT) images and 1-, 5-, and 10-shot configurations, we perform multiple repetitions and compare models using independent two-sample t-tests and effect sizes defined by Cohen's d. At a fixed rotation speed, advanced MLLMs significantly outperform the FSL baseline, whereas Prototypical Networks remain competitive under variable speeds. The prompt study shows that more verbose templates do not bring benefits. The concise prompt generally matches or outperforms the detailed one, which can even degrade performance. Overall, concise prompts are sufficient and often preferable, and MLLMs emerge as an effective complement to conventional FSL approaches in realistic industrial scenarios.
Multimodal Large Language Models for Few-Shot Bearing Fault Diagnosis: Impact of Prompt Design and Comparison with Prototypical Networks / Di Maggio, L.G., Delprete, C., Brusa, E.. - (2026), pp. 1-7. (2026 International Conference on Control, Automation and Diagnosis, ICCAD 2026 Lisbona 2026) [10.1109/ICCAD69956.2026.11643196].
Multimodal Large Language Models for Few-Shot Bearing Fault Diagnosis: Impact of Prompt Design and Comparison with Prototypical Networks
Di Maggio L. G.;Delprete C.;Brusa E.
2026
Abstract
Few-shot learning (FSL) models for bearing fault diagnosis often require meta-training on dedicated datasets and are limited by the scarcity of fault data in real industrial scenarios. Multimodal Large Language Models (MLLMs) offer an alternative, enabling few-shot classification via prompting on time-frequency images of vibration signals. This work studies medium-to-large bearings tested on a dedicated test rig and analyzes (i) when MLLM performance is statistically superior to a Prototypical Network FSL baseline, and (ii) how prompt design affects few-shot classification. On a 4-way task with Continuous Wavelet Transform (CWT) images and 1-, 5-, and 10-shot configurations, we perform multiple repetitions and compare models using independent two-sample t-tests and effect sizes defined by Cohen's d. At a fixed rotation speed, advanced MLLMs significantly outperform the FSL baseline, whereas Prototypical Networks remain competitive under variable speeds. The prompt study shows that more verbose templates do not bring benefits. The concise prompt generally matches or outperforms the detailed one, which can even degrade performance. Overall, concise prompts are sufficient and often preferable, and MLLMs emerge as an effective complement to conventional FSL approaches in realistic industrial scenarios.Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/11583/3016245
Attenzione
Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo
