Slopsquatting is an emerging software supply-chain attack in which adversaries exploit package-name hallucinations produced by Large Language Models (LLMs) during AI-assisted code generation. This paper presents Hal-Detector, an explainable AI pipeline for identifying hallucinated package names in LLM-generated JavaScript code. To distinguish real and hallucinated package names while providing token-level explanations of Hal-Detector's model decisions, the proposed approach combines LLM adaptation via probing, Shapley Value estimation, and topology regularisation. The LLM's decisions are explained by exposing the lexical patterns and package-name segments that drive its predictions. We evaluate Hal-Detector on a dataset of JavaScript package names and compare multiple combinations of black-box models, Shapley values estimators, and topological regularisation settings. Results show that the proposed framework can characterise hallucinated dependencies, provide consistent local and global explanations, and offer preliminary evidence that topological regularisation promotes more structured latent representations. These findings suggest that explainability-oriented hallucination detection can mitigate security concerns in LLM-assisted software development.

Hal-Detector: Is Explainable AI Helping Detect Slopsquatting? / Bonfanti, C., Colaiacomo, D., Cagliero, L., Basile, C.. - (In corso di stampa). (International Conference on Cybersecurity and AI-Based Systems Bucharest, Romania 22–25 September 2026).

Hal-Detector: Is Explainable AI Helping Detect Slopsquatting?

Bonfanti,Chiara;Colaiacomo,Davide;Cagliero, Luca;Basile, Cataldo
In corso di stampa

Abstract

Slopsquatting is an emerging software supply-chain attack in which adversaries exploit package-name hallucinations produced by Large Language Models (LLMs) during AI-assisted code generation. This paper presents Hal-Detector, an explainable AI pipeline for identifying hallucinated package names in LLM-generated JavaScript code. To distinguish real and hallucinated package names while providing token-level explanations of Hal-Detector's model decisions, the proposed approach combines LLM adaptation via probing, Shapley Value estimation, and topology regularisation. The LLM's decisions are explained by exposing the lexical patterns and package-name segments that drive its predictions. We evaluate Hal-Detector on a dataset of JavaScript package names and compare multiple combinations of black-box models, Shapley values estimators, and topological regularisation settings. Results show that the proposed framework can characterise hallucinated dependencies, provide consistent local and global explanations, and offer preliminary evidence that topological regularisation promotes more structured latent representations. These findings suggest that explainability-oriented hallucination detection can mitigate security concerns in LLM-assisted software development.
In corso di stampa
File in questo prodotto:
Non ci sono file associati a questo prodotto.
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11583/3015318