The pervasive spread of online hate speech on social media has fostered the development of counter-narrative generation strategies based on Large Language Models (LLMs). Previous approaches to mitigate hate speech content has either fine-tuned LLMs on human-curated datasets consisting of single-turn interactions, i.e., pairs of hate speech and counternarratives, or leveraged in-context learning strategies. However, these approaches fall short when contextspecific counter-narratives are unavailable or lack sufficient evidence. To overcome this issue, in this work we propose a novel method that combines an adversarial LLM-based approach to produce multi-turn debates with hate speech counter-narrative generation based on Retrieval-Augmented Generation (RAG). In the early adversarial learning stage, the LLM agent defends an unethical positions whereas the adversarial responds with ethical, evidence-based counter-arguments. A human evaluation of the generated counter-narratives confirmed their relevance and pertinence. Then, a RAG system retrieves the evidence-based counter-narratives and generates the mitigated version of the hate speech content. Experimental results show that the RAG approach outperforms state-of-the-art methods on the MultiTarget CONAN benchmark dataset. Implementation details, prompts, and further information for reproducibility available here: https://github.com/erfan-bayat13/debaterag-counterspeech Notice: this paper includes potentially offensive content for scientific purposes only.

Adversarial Generation of Multi-Turn Debates for Hate Speech Mitigation via Retrieval-Augmented Generation / Bayat, E., Gensale, A., Cagliero, L.. - (2026), pp. 1443-1452. (2026 IEEE 50th Annual Computers, Software, and Applications Conference (COMPSAC) Madrid (ESP) 07-10 July 2026) [10.1109/compsac69091.2026.00191].

Adversarial Generation of Multi-Turn Debates for Hate Speech Mitigation via Retrieval-Augmented Generation

Bayat, Erfan;Gensale, Aurora;Cagliero, Luca
2026

Abstract

The pervasive spread of online hate speech on social media has fostered the development of counter-narrative generation strategies based on Large Language Models (LLMs). Previous approaches to mitigate hate speech content has either fine-tuned LLMs on human-curated datasets consisting of single-turn interactions, i.e., pairs of hate speech and counternarratives, or leveraged in-context learning strategies. However, these approaches fall short when contextspecific counter-narratives are unavailable or lack sufficient evidence. To overcome this issue, in this work we propose a novel method that combines an adversarial LLM-based approach to produce multi-turn debates with hate speech counter-narrative generation based on Retrieval-Augmented Generation (RAG). In the early adversarial learning stage, the LLM agent defends an unethical positions whereas the adversarial responds with ethical, evidence-based counter-arguments. A human evaluation of the generated counter-narratives confirmed their relevance and pertinence. Then, a RAG system retrieves the evidence-based counter-narratives and generates the mitigated version of the hate speech content. Experimental results show that the RAG approach outperforms state-of-the-art methods on the MultiTarget CONAN benchmark dataset. Implementation details, prompts, and further information for reproducibility available here: https://github.com/erfan-bayat13/debaterag-counterspeech Notice: this paper includes potentially offensive content for scientific purposes only.
2026
979-8-3315-4497-3
File in questo prodotto:
File Dimensione Formato  
Adversarial_Generation_of_Multi-Turn_Debates_for_Hate_Speech_Mitigation_via_Retrieval-Augmented_Generation.pdf

accesso riservato

Tipologia: 2a Post-print versione editoriale / Version of Record
Licenza: Non Pubblico - Accesso privato/ristretto
Dimensione 499.06 kB
Formato Adobe PDF
499.06 kB Adobe PDF   Visualizza/Apri   Richiedi una copia
COMPSAC_26___Adversarial_generation_of_multi_turn____-1.pdf

accesso aperto

Tipologia: 2. Post-print / Author's Accepted Manuscript
Licenza: Pubblico - Tutti i diritti riservati
Dimensione 458.63 kB
Formato Adobe PDF
458.63 kB Adobe PDF Visualizza/Apri
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11583/3015048