Psychological stress is a major risk factor for several physical and mental health conditions, motivating the development of reliable and unobtrusive monitoring solutions. Wearable devices enable continuous and ecological acquisition of physiological and physical signals, providing a promising basis for automatic stress detection in real-world settings. In this context, machine learning (ML) and deep learning (DL) methods have shown strong potential for extracting stress-related patterns from multimodal wearable data. However, most studies evaluate performance on a single dataset, limiting insights into model robustness and generalization across populations and experimental conditions. Furthermore, there is currently no comprehensive comparison between feature-based ML methods and data-driven DL methods. Finally, the use of transfer learning to address differences across various datasets has not yet been thoroughly explored. In this work, we investigate automatic stress detection from smartwatch data by comparing feature- and data-driven models, focusing on cross-dataset generalization. Experiments were conducted on four public datasets collected with the same wearable device and sensing modalities, including photoplethysmography, acceleration, skin temperature, and electrodermal activity. We first evaluated model performance within each dataset to establish baseline results. We then assessed generalization using cross-dataset testing, where models trained on one dataset are evaluated on unseen datasets. Finally, we explored different transfer learning strategies to improve cross-dataset performance and promote generalization, especially in small datasets. Results show that DL methods (F1-score 0.69–0.96) consistently outperform ML models (F1-score 0.69–0.95) in within-dataset evaluations, although performance varies across datasets. Cross-dataset tests reveal significant performance degradation, particularly for DL models, where the F1-score showed an average drop of −21%. Transfer learning mitigates this gap, improving robustness across datasets. These findings highlight the importance of cross-dataset evaluation for realistic assessment of wearable-based stress detection systems.

Machine learning-based automatic stress detection: Performance and generalization across datasets / Calza-Metre, M., Borzì, L.. - In: SMART HEALTH. - ISSN 2352-6483. - ELETTRONICO. - 41:(2026). [10.1016/j.smhl.2026.100693]

Machine learning-based automatic stress detection: Performance and generalization across datasets

Luigi Borzì
2026

Abstract

Psychological stress is a major risk factor for several physical and mental health conditions, motivating the development of reliable and unobtrusive monitoring solutions. Wearable devices enable continuous and ecological acquisition of physiological and physical signals, providing a promising basis for automatic stress detection in real-world settings. In this context, machine learning (ML) and deep learning (DL) methods have shown strong potential for extracting stress-related patterns from multimodal wearable data. However, most studies evaluate performance on a single dataset, limiting insights into model robustness and generalization across populations and experimental conditions. Furthermore, there is currently no comprehensive comparison between feature-based ML methods and data-driven DL methods. Finally, the use of transfer learning to address differences across various datasets has not yet been thoroughly explored. In this work, we investigate automatic stress detection from smartwatch data by comparing feature- and data-driven models, focusing on cross-dataset generalization. Experiments were conducted on four public datasets collected with the same wearable device and sensing modalities, including photoplethysmography, acceleration, skin temperature, and electrodermal activity. We first evaluated model performance within each dataset to establish baseline results. We then assessed generalization using cross-dataset testing, where models trained on one dataset are evaluated on unseen datasets. Finally, we explored different transfer learning strategies to improve cross-dataset performance and promote generalization, especially in small datasets. Results show that DL methods (F1-score 0.69–0.96) consistently outperform ML models (F1-score 0.69–0.95) in within-dataset evaluations, although performance varies across datasets. Cross-dataset tests reveal significant performance degradation, particularly for DL models, where the F1-score showed an average drop of −21%. Transfer learning mitigates this gap, improving robustness across datasets. These findings highlight the importance of cross-dataset evaluation for realistic assessment of wearable-based stress detection systems.
2026
File in questo prodotto:
File Dimensione Formato  
1-s2.0-S2352648326000619-main.pdf

accesso aperto

Descrizione: Articolo
Tipologia: 2a Post-print versione editoriale / Version of Record
Licenza: Creative commons
Dimensione 4.77 MB
Formato Adobe PDF
4.77 MB Adobe PDF Visualizza/Apri
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11583/3013267