Transforming cultural heritage knowledge into compelling narratives is a challenging task, requiring the integration of historical accuracy, clear narrative structure, and audience engagement. Museum catalogues contain rich but loosely structured information that is difficult to convert into coherent and effective stories without compromising fidelity to documented events. Although large language models (LLMs) enable fluent story generation, fluency alone is insufficient to address these challenges. To this end, we present a structured and interpretable pipeline for knowledge-grounded storytelling in cultural heritage domains, explicitly designed to integrate source knowledge with narrative structure. The approach separates narrative event extraction, narratological structuring, and constrained story generation. Central to the pipeline is an explicit narrative scaffold that models conflict structures and role relations, constraining how events are organized into multi-scene stories and ensuring that catalogue-derived knowledge plays a functional role in the causal and dramatic progression of the narrative. We further address the challenge of evaluating narrative generation systems by introducing a rubric-based evaluation protocol that aligns historical and dramatic dimensions of narrative quality. The same evaluation framework is applied by human experts (curators and drama experts) and by automated LLM-based judges, enabling both fine-grained human assessment and scalable automated analysis. Rather than assuming the validity of LLM-as-a-judge, we explicitly study its reliability, agreement with expert judgments, and sources of bias. To validate the proposed pipeline and evaluation framework and assess its viability within museum workflows, we apply it to a real-world case study based on the catalogue of the Egyptian Museum of Turin. Through quantitative agreement analysis and qualitative inspection, we provide empirical insight into the suitability and limitations of LLM-based judging for structured narrative evaluation.

From a True Story: Leveraging Museum Catalogue Data for LLM-Driven Narrative Generation / Mensa, E., Pecora, A.E., Fulfaro, C., Pizzo, A., Ferraris, E., Bottino, A., Damiano, R.. - In: ACM TRANSACTIONS ON INTELLIGENT SYSTEMS AND TECHNOLOGY. - ISSN 2157-6904. - (In corso di stampa).

From a True Story: Leveraging Museum Catalogue Data for LLM-Driven Narrative Generation

Pecora, Alessandro Emmanuel;Bottino, Andrea;
In corso di stampa

Abstract

Transforming cultural heritage knowledge into compelling narratives is a challenging task, requiring the integration of historical accuracy, clear narrative structure, and audience engagement. Museum catalogues contain rich but loosely structured information that is difficult to convert into coherent and effective stories without compromising fidelity to documented events. Although large language models (LLMs) enable fluent story generation, fluency alone is insufficient to address these challenges. To this end, we present a structured and interpretable pipeline for knowledge-grounded storytelling in cultural heritage domains, explicitly designed to integrate source knowledge with narrative structure. The approach separates narrative event extraction, narratological structuring, and constrained story generation. Central to the pipeline is an explicit narrative scaffold that models conflict structures and role relations, constraining how events are organized into multi-scene stories and ensuring that catalogue-derived knowledge plays a functional role in the causal and dramatic progression of the narrative. We further address the challenge of evaluating narrative generation systems by introducing a rubric-based evaluation protocol that aligns historical and dramatic dimensions of narrative quality. The same evaluation framework is applied by human experts (curators and drama experts) and by automated LLM-based judges, enabling both fine-grained human assessment and scalable automated analysis. Rather than assuming the validity of LLM-as-a-judge, we explicitly study its reliability, agreement with expert judgments, and sources of bias. To validate the proposed pipeline and evaluation framework and assess its viability within museum workflows, we apply it to a real-world case study based on the catalogue of the Egyptian Museum of Turin. Through quantitative agreement analysis and qualitative inspection, we provide empirical insight into the suitability and limitations of LLM-based judging for structured narrative evaluation.
In corso di stampa
File in questo prodotto:
File Dimensione Formato  
251219___TIST26_SI_Storytelling__POLI_hosting_-11.pdf

accesso riservato

Tipologia: 1. Preprint / submitted version [pre- review]
Licenza: Non Pubblico - Accesso privato/ristretto
Dimensione 2.02 MB
Formato Adobe PDF
2.02 MB Adobe PDF   Visualizza/Apri   Richiedi una copia
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11583/3015267