Optimizing Datasets for Code Summarization: Is Code-Comment Coherence Enough?

Vitale, A.; Mastropaolo, A.; Oliveto, R.; Di Penta, M.; Scalabrino, S.

doi:10.1109/ICPC66645.2025.00033

Automated code summarization is a long-standing goal for code comprehension. This task automatically generates documentation using a given method. Deep Learning (DL) -based approaches have been proven beneficial for various software engineering (SE) tasks, including this one. Most state-of-the-art datasets for code summarization are automatically mined from GitHub and, thus, might contain erroneous or sub-optimal examples. Previous work showed that using a simple rule-based approach for removing noisy instances allows for a tangible reduction of the training set size while not reducing the effectiveness of the trained models. Motivated by this finding, we conjecture that it is possible to further reduce the dataset size by removing instances that contain different issues. In this paper, we explore the extent to which code-comment coherence, a specific quality attribute of code summaries, can be used to optimize code summarization datasets. Specifically, we hypothesize that removing incoherent code-comment pairs might positively impact the effectiveness of the models. To do this, we rely on SIDE, a recently introduced metric for code-summary coherence. We examine multiple selectivity levels of training instances from two state-of-the-art datasets (TL-CodeSum and Funcom) and evaluate the resulting models on three manually curated test sets. The results show that even halving the training set sizes does not significantly affect the model's ability to generate summaries. However, when comparing the most restrictive selection strategy with a simpler one that randomly selects the training instances, we observe that the resulting accuracy of the model also does not change. This result suggests that (i) current datasets contain many irrelevant examples, and (ii) different quality attributes should be explored for optimizing code summarization datasets.

Optimizing Datasets for Code Summarization: Is Code-Comment Coherence Enough? / Vitale, A.; Mastropaolo, A.; Oliveto, R.; Di Penta, M.; Scalabrino, S.. - (2025), pp. 237-249. ( 2025 IEEE/ACM 33rd International Conference on Program Comprehension (ICPC) Ottawa (CAN) 27-28 April 2025) [10.1109/ICPC66645.2025.00033].

Optimizing Datasets for Code Summarization: Is Code-Comment Coherence Enough?

Vitale A.;Mastropaolo A.;Oliveto R.;Di Penta M.;Scalabrino S.

2025

Abstract

Scheda breve

Scheda completa

Scheda completa (DC)

	Anno del prodotto
	
				2025
			
	Titolo della Serie/Collana
	
				PROCEEDINGS IEEE INTERNATIONAL CONFERENCE ON PROGRAM COMPREHENSION
			
	Codice ISBN
	
				979-8-3315-0223-2
			
	Appare nelle tipologie
	
				4.1 Contributo in Atti di convegno

File in questo prodotto:

File	Dimensione	Formato
Optimizing_Datasets_for_Code_Summarization_Is_Code-Comment_Coherence_Enough.pdf accesso riservato Tipologia: 2a Post-print versione editoriale / Version of Record Licenza: Non Pubblico - Accesso privato/ristretto Dimensione 566.03 kB Formato Adobe PDF Visualizza/Apri Richiedi una copia	566.03 kB	Adobe PDF	Visualizza/Apri Richiedi una copia
Optimizing_Datasets_for_Code_Summarization_Is_Code-Comment_Coherence_Enough.pdf accesso aperto Tipologia: 2. Post-print / Author's Accepted Manuscript Licenza: Pubblico - Tutti i diritti riservati Dimensione 462.79 kB Formato Adobe PDF Visualizza/Apri	462.79 kB	Adobe PDF	Visualizza/Apri

Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11583/3002676

PORTO @ Archivio Istituzionale della Ricerca

Optimizing Datasets for Code Summarization: Is Code-Comment Coherence Enough?

Vitale A.;Mastropaolo A.;Oliveto R.;Di Penta M.;Scalabrino S.

2025

Abstract

Scheda breve Scheda completa Scheda completa (DC)

Pubblicazioni consigliate

Informazioni

Conferma cancellazione

Scheda breve

Scheda completa

Scheda completa (DC)