The growing use of large language models for code generation makes distinguishing machine-generated code from human-written code increasingly difficult, especially under distribution shifts in language, domain, and generator family. SemEval-2026 Task 13 targets this challenge through three subtasks: binary detection, multi-class authorship attribution, and hybrid/adversarial code detection.In this paper, we conduct an empirical study across all subtasks, comparing a variety of approaches: frozen encoder representations, feature-based classifiers, fine-tuned transformer models, post-hoc calibration, and probability-level ensembling. Our results show a consistent generalisation gap: strong in-domain validation scores substantially overestimate performance on shifted test conditions.The code is available at https://github.com/AlexandraElena-Holota/SemEval-2026-Task13.git

MINDS at SemEval-2026-Task 13: Robust Detection of Machine-Generated Code under Distribution Shift / Buccelli Giorgia, R., Coviello, A., Holota, A.E., Scaglione, M., Scalora, S., Savelli, C., Coppola, R., Giobergia, F.. - (2026), pp. 3080-3088. (20th International Workshop on Semantic Evaluation (SemEval-2026) San Diego (USA) July 2-7, 2026) [10.18653/v1/2026.semeval-1.386].

MINDS at SemEval-2026-Task 13: Robust Detection of Machine-Generated Code under Distribution Shift

Coviello Antonella;Holota Alexandra Elena;Scaglione Marco;Savelli Claudio;Coppola Riccardo;Giobergia Flavio
2026

Abstract

The growing use of large language models for code generation makes distinguishing machine-generated code from human-written code increasingly difficult, especially under distribution shifts in language, domain, and generator family. SemEval-2026 Task 13 targets this challenge through three subtasks: binary detection, multi-class authorship attribution, and hybrid/adversarial code detection.In this paper, we conduct an empirical study across all subtasks, comparing a variety of approaches: frozen encoder representations, feature-based classifiers, fine-tuned transformer models, post-hoc calibration, and probability-level ensembling. Our results show a consistent generalisation gap: strong in-domain validation scores substantially overestimate performance on shifted test conditions.The code is available at https://github.com/AlexandraElena-Holota/SemEval-2026-Task13.git
2026
979-8-89176-414-9
File in questo prodotto:
File Dimensione Formato  
2026.semeval-1.386.pdf

accesso aperto

Tipologia: 2a Post-print versione editoriale / Version of Record
Licenza: Creative commons
Dimensione 507.23 kB
Formato Adobe PDF
507.23 kB Adobe PDF Visualizza/Apri
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11583/3015350