Regulated AI systems need requirements that link testable SHALL statements to statutory source text, yet local large language models often invent obligations or misstate modality when elicitation is under-specified. We ask whether a sequential Analyst–Engineer–QA multi-agent workflow (B5) outperforms fair baselines when every role shares one Llama 3 8B model via Ollama on EU AI Act and GDPR chunks. Using quote-validated expert gold, the structured single-prompt baseline B1 achieves the highest F1 on the EU AI Act test (0.382) and GDPR generalization (0.337), ahead of B5 (0.320,0.322) and unstructured B0 (0.269, 0.275). B5 exhibits higher hallucination than B1 on both splits (e.g., 0.134 vs 0.009 onEU test) despite three LLM calls per chunk versus one for B1. Ablations B0–B5 indicate that precision gains, not added agents, drive quality on consumer hardware. We answer the primary research question in the negative for matched prompt content: role-separated inference (B5) does not beat the singlecall structured prompt (B1) on F1. For trustworthy, privacypreserving regulatory requirements engineering, practitioners should prefer single-call structured prompts and deploy multiagent pipelines only when intermediate artifacts are required for human audit.

Trustworthy Regulatory Requirements Extraction using Local LLM: A Multi-Agent vs. Single-Prompt Benchmark on the EU AI Act and GDPR / Ahmed, M., Naz, M.A.. - ELETTRONICO. - (2026), pp. 1-5. (2026 IEEE 2nd International Conference on AI and Emerging Technology for Sustainable Future Catania (Italy) 24-25 July 2026).

Trustworthy Regulatory Requirements Extraction using Local LLM: A Multi-Agent vs. Single-Prompt Benchmark on the EU AI Act and GDPR

Naz, Muhammad Ajmal
2026

Abstract

Regulated AI systems need requirements that link testable SHALL statements to statutory source text, yet local large language models often invent obligations or misstate modality when elicitation is under-specified. We ask whether a sequential Analyst–Engineer–QA multi-agent workflow (B5) outperforms fair baselines when every role shares one Llama 3 8B model via Ollama on EU AI Act and GDPR chunks. Using quote-validated expert gold, the structured single-prompt baseline B1 achieves the highest F1 on the EU AI Act test (0.382) and GDPR generalization (0.337), ahead of B5 (0.320,0.322) and unstructured B0 (0.269, 0.275). B5 exhibits higher hallucination than B1 on both splits (e.g., 0.134 vs 0.009 onEU test) despite three LLM calls per chunk versus one for B1. Ablations B0–B5 indicate that precision gains, not added agents, drive quality on consumer hardware. We answer the primary research question in the negative for matched prompt content: role-separated inference (B5) does not beat the singlecall structured prompt (B1) on F1. For trustworthy, privacypreserving regulatory requirements engineering, practitioners should prefer single-call structured prompts and deploy multiagent pipelines only when intermediate artifacts are required for human audit.
File in questo prodotto:
Non ci sono file associati a questo prodotto.
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11583/3015996
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo