The growing use of large language models for code generation makes distinguishing machine-generated code from human-written code increasingly difficult, especially under distribution shifts in language, domain, and generator family. SemEval-2026 Task 13 targets this challenge through three subtasks: binary detection, multi-class authorship attribution, and hybrid/adversarial code detection.In this paper, we conduct an empirical study across all subtasks, comparing a variety of approaches: frozen encoder representations, feature-based classifiers, fine-tuned transformer models, post-hoc calibration, and probability-level ensembling. Our results show a consistent generalisation gap: strong in-domain validation scores substantially overestimate performance on shifted test conditions.The code is available at https://github.com/AlexandraElena-Holota/SemEval-2026-Task13.git
MINDS at SemEval-2026-Task 13: Robust Detection of Machine-Generated Code under Distribution Shift / Buccelli Giorgia, R., Coviello, A., Holota, A.E., Scaglione, M., Scalora, S., Savelli, C., Coppola, R., Giobergia, F.. - (2026), pp. 3080-3088. (20th International Workshop on Semantic Evaluation (SemEval-2026) San Diego (USA) July 2-7, 2026) [10.18653/v1/2026.semeval-1.386].
MINDS at SemEval-2026-Task 13: Robust Detection of Machine-Generated Code under Distribution Shift
Coviello Antonella;Holota Alexandra Elena;Scaglione Marco;Savelli Claudio;Coppola Riccardo;Giobergia Flavio
2026
Abstract
The growing use of large language models for code generation makes distinguishing machine-generated code from human-written code increasingly difficult, especially under distribution shifts in language, domain, and generator family. SemEval-2026 Task 13 targets this challenge through three subtasks: binary detection, multi-class authorship attribution, and hybrid/adversarial code detection.In this paper, we conduct an empirical study across all subtasks, comparing a variety of approaches: frozen encoder representations, feature-based classifiers, fine-tuned transformer models, post-hoc calibration, and probability-level ensembling. Our results show a consistent generalisation gap: strong in-domain validation scores substantially overestimate performance on shifted test conditions.The code is available at https://github.com/AlexandraElena-Holota/SemEval-2026-Task13.git| File | Dimensione | Formato | |
|---|---|---|---|
|
2026.semeval-1.386.pdf
accesso aperto
Tipologia:
2a Post-print versione editoriale / Version of Record
Licenza:
Creative commons
Dimensione
507.23 kB
Formato
Adobe PDF
|
507.23 kB | Adobe PDF | Visualizza/Apri |
Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/11583/3015350
