Computational notebooks are widely used in data science workflows, yet theircode cells are often poorly documented, hindering comprehension, reproducibility and maintenance. Unlike traditional software artifacts, notebook cells frequently perform implicit datatransformations whose intent cannot be inferred from syntax alone, making automatic comment generation particularly challenging. In this paper, we investigate whether executionaware specialization can compensate for limited training data and smaller model scale innotebook comment generation. We propose an execution-aware encoder–decoder model thataugments static code representations with compact symbolic descriptions capturing high-leveldata transformations. The model is trained exclusively on a curated notebook dataset andcompared against lightweight code-oriented LLMs from the Qwen2.5-Coder family of comparable parameter scale. Empirical evaluation on a unified test set shows that the executionaware model achieves competitive performance. It outperforms similarly sized LLMs onROUGE-L and embedding-based semantic similarity metrics, and controlled data augmentation further improves robustness. Moreover, it substantially reduces inference time, energyconsumption, and estimated CO2 emissions under identical hardware settings. These resultssuggest that execution-aware architectures can provide a practical alternative to lightweightLLMs for notebook comment generation, especially when documentation support must account for runtime data transformations and operate under constrained computational budgets.
Execution-Aware Comment Generation for Computational Notebooks: An Empirical Comparison with Lightweight LLMs / Fantino, G., Vetro', A., Torchiano, M., Cappelluti, F.. - ELETTRONICO. - (In corso di stampa). (19th International Conference on the Quality of Information and Communications Technology Genova (ITA) September 9-11, 2026).
Execution-Aware Comment Generation for Computational Notebooks: An Empirical Comparison with Lightweight LLMs
Fantino, Giacomo;Vetro', Antonio;Torchiano, Marco;Cappelluti, Federica
In corso di stampa
Abstract
Computational notebooks are widely used in data science workflows, yet theircode cells are often poorly documented, hindering comprehension, reproducibility and maintenance. Unlike traditional software artifacts, notebook cells frequently perform implicit datatransformations whose intent cannot be inferred from syntax alone, making automatic comment generation particularly challenging. In this paper, we investigate whether executionaware specialization can compensate for limited training data and smaller model scale innotebook comment generation. We propose an execution-aware encoder–decoder model thataugments static code representations with compact symbolic descriptions capturing high-leveldata transformations. The model is trained exclusively on a curated notebook dataset andcompared against lightweight code-oriented LLMs from the Qwen2.5-Coder family of comparable parameter scale. Empirical evaluation on a unified test set shows that the executionaware model achieves competitive performance. It outperforms similarly sized LLMs onROUGE-L and embedding-based semantic similarity metrics, and controlled data augmentation further improves robustness. Moreover, it substantially reduces inference time, energyconsumption, and estimated CO2 emissions under identical hardware settings. These resultssuggest that execution-aware architectures can provide a practical alternative to lightweightLLMs for notebook comment generation, especially when documentation support must account for runtime data transformations and operate under constrained computational budgets.| File | Dimensione | Formato | |
|---|---|---|---|
|
QUATIC_2026_camera.pdf
accesso riservato
Tipologia:
2. Post-print / Author's Accepted Manuscript
Licenza:
Non Pubblico - Accesso privato/ristretto
Dimensione
418.33 kB
Formato
Adobe PDF
|
418.33 kB | Adobe PDF | Visualizza/Apri Richiedi una copia |
Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/11583/3015367
