Computational notebooks are widely used in data science workflows, yet theircode cells are often poorly documented, hindering comprehension, reproducibility and maintenance. Unlike traditional software artifacts, notebook cells frequently perform implicit datatransformations whose intent cannot be inferred from syntax alone, making automatic comment generation particularly challenging. In this paper, we investigate whether executionaware specialization can compensate for limited training data and smaller model scale innotebook comment generation. We propose an execution-aware encoder–decoder model thataugments static code representations with compact symbolic descriptions capturing high-leveldata transformations. The model is trained exclusively on a curated notebook dataset andcompared against lightweight code-oriented LLMs from the Qwen2.5-Coder family of comparable parameter scale. Empirical evaluation on a unified test set shows that the executionaware model achieves competitive performance. It outperforms similarly sized LLMs onROUGE-L and embedding-based semantic similarity metrics, and controlled data augmentation further improves robustness. Moreover, it substantially reduces inference time, energyconsumption, and estimated CO2 emissions under identical hardware settings. These resultssuggest that execution-aware architectures can provide a practical alternative to lightweightLLMs for notebook comment generation, especially when documentation support must account for runtime data transformations and operate under constrained computational budgets.
Execution-Aware Comment Generation for Computational Notebooks: An Empirical Comparison with Lightweight LLMs / Fantino, G., Vetro', A., Torchiano, M., Cappelluti, F.. - ELETTRONICO. - (2026). (19th International Conference on the Quality of Information and Communications Technology Genova (ITA) ).
Execution-Aware Comment Generation for Computational Notebooks: An Empirical Comparison with Lightweight LLMs
Fantino, Giacomo;Vetro', Antonio;Torchiano, Marco;Cappelluti, Federica
2026
Abstract
Computational notebooks are widely used in data science workflows, yet theircode cells are often poorly documented, hindering comprehension, reproducibility and maintenance. Unlike traditional software artifacts, notebook cells frequently perform implicit datatransformations whose intent cannot be inferred from syntax alone, making automatic comment generation particularly challenging. In this paper, we investigate whether executionaware specialization can compensate for limited training data and smaller model scale innotebook comment generation. We propose an execution-aware encoder–decoder model thataugments static code representations with compact symbolic descriptions capturing high-leveldata transformations. The model is trained exclusively on a curated notebook dataset andcompared against lightweight code-oriented LLMs from the Qwen2.5-Coder family of comparable parameter scale. Empirical evaluation on a unified test set shows that the executionaware model achieves competitive performance. It outperforms similarly sized LLMs onROUGE-L and embedding-based semantic similarity metrics, and controlled data augmentation further improves robustness. Moreover, it substantially reduces inference time, energyconsumption, and estimated CO2 emissions under identical hardware settings. These resultssuggest that execution-aware architectures can provide a practical alternative to lightweightLLMs for notebook comment generation, especially when documentation support must account for runtime data transformations and operate under constrained computational budgets.Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/11583/3015367
Attenzione
Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo
