The pervasive spread of online hate speech on social media has fostered the development of counter-narrative generation strategies based on Large Language Models (LLMs). Previous approaches to mitigate hate speech content has either fine-tuned LLMs on human-curated datasets consisting of single-turn interactions, i.e., pairs of hate speech and counternarratives, or leveraged in-context learning strategies. However, these approaches fall short when contextspecific counter-narratives are unavailable or lack sufficient evidence. To overcome this issue, in this work we propose a novel method that combines an adversarial LLM-based approach to produce multi-turn debates with hate speech counter-narrative generation based on Retrieval-Augmented Generation (RAG). In the early adversarial learning stage, the LLM agent defends an unethical positions whereas the adversarial responds with ethical, evidence-based counter-arguments. A human evaluation of the generated counter-narratives confirmed their relevance and pertinence. Then, a RAG system retrieves the evidence-based counter-narratives and generates the mitigated version of the hate speech content. Experimental results show that the RAG approach outperforms state-of-the-art methods on the MultiTarget CONAN benchmark dataset. Implementation details, prompts, and further information for reproducibility available here: https://github.com/erfan-bayat13/debaterag-counterspeech Notice: this paper includes potentially offensive content for scientific purposes only.
Adversarial Generation of Multi-Turn Debates for Hate Speech Mitigation via Retrieval-Augmented Generation / Bayat, E., Gensale, A., Cagliero, L.. - (2026), pp. 1443-1452. (2026 IEEE 50th Annual Computers, Software, and Applications Conference (COMPSAC) Madrid (ESP) 07-10 July 2026) [10.1109/compsac69091.2026.00191].
Adversarial Generation of Multi-Turn Debates for Hate Speech Mitigation via Retrieval-Augmented Generation
Bayat, Erfan;Gensale, Aurora;Cagliero, Luca
2026
Abstract
The pervasive spread of online hate speech on social media has fostered the development of counter-narrative generation strategies based on Large Language Models (LLMs). Previous approaches to mitigate hate speech content has either fine-tuned LLMs on human-curated datasets consisting of single-turn interactions, i.e., pairs of hate speech and counternarratives, or leveraged in-context learning strategies. However, these approaches fall short when contextspecific counter-narratives are unavailable or lack sufficient evidence. To overcome this issue, in this work we propose a novel method that combines an adversarial LLM-based approach to produce multi-turn debates with hate speech counter-narrative generation based on Retrieval-Augmented Generation (RAG). In the early adversarial learning stage, the LLM agent defends an unethical positions whereas the adversarial responds with ethical, evidence-based counter-arguments. A human evaluation of the generated counter-narratives confirmed their relevance and pertinence. Then, a RAG system retrieves the evidence-based counter-narratives and generates the mitigated version of the hate speech content. Experimental results show that the RAG approach outperforms state-of-the-art methods on the MultiTarget CONAN benchmark dataset. Implementation details, prompts, and further information for reproducibility available here: https://github.com/erfan-bayat13/debaterag-counterspeech Notice: this paper includes potentially offensive content for scientific purposes only.| File | Dimensione | Formato | |
|---|---|---|---|
|
Adversarial_Generation_of_Multi-Turn_Debates_for_Hate_Speech_Mitigation_via_Retrieval-Augmented_Generation.pdf
accesso riservato
Tipologia:
2a Post-print versione editoriale / Version of Record
Licenza:
Non Pubblico - Accesso privato/ristretto
Dimensione
499.06 kB
Formato
Adobe PDF
|
499.06 kB | Adobe PDF | Visualizza/Apri Richiedi una copia |
|
COMPSAC_26___Adversarial_generation_of_multi_turn____-1.pdf
accesso aperto
Tipologia:
2. Post-print / Author's Accepted Manuscript
Licenza:
Pubblico - Tutti i diritti riservati
Dimensione
458.63 kB
Formato
Adobe PDF
|
458.63 kB | Adobe PDF | Visualizza/Apri |
Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/11583/3015048
