Tracking Knowledge Propagation Across Wikipedia Languages

Valentim, Rodolfo Vieira; Comarela, Giovanni; Park, Souneil; Sáez-Trumper, Diego

In this paper, we present a dataset of inter-language knowledge propagation in Wikipedia. Covering the entire 309 language editions and 33M articles, the dataset aims to track the full propagation history of Wikipedia concepts, and allow follow-up research on building predictive models of them. For this purpose, we align all the Wikipedia articles in a language-agnostic manner according to the concept they cover, which results in 13M propagation instances. To the best of our knowledge, this dataset is the first to explore the full inter-language propagation at a large scale. Together with the dataset, a holistic overview of the propagation and key insights about the underlying structural factors are provided to aid future research. For example, we find that although long cascades are unusual, the propagation tends to continue further once it reaches more than four language editions. We also find that the size of language editions is associated with the speed of propagation. We believe the dataset not only contributes to the prior literature on Wikipedia growth but also enables new use cases such as edit recommendation for addressing knowledge gaps, detection of disinformation, and cultural relationship analysis.

Tracking Knowledge Propagation Across Wikipedia Languages / Valentim, R.V., Comarela, G., Park, S., Sáez-Trumper, D.. - ELETTRONICO. - 15:(2021), pp. 1046-1052. (Fifteenth International AAAI Conference on Web and Social Media virtuale June 7-10, 2021).

Tracking Knowledge Propagation Across Wikipedia Languages

Valentim, Rodolfo Vieira;Comarela, Giovanni;Park, Souneil;Sáez-Trumper, Diego

2021

Abstract

Scheda breve

Scheda completa

Scheda completa (DC)

	Anno del prodotto
	
				2021
			
	Codice ISBN
	
				978-1-57735-869-5
			
	Appare nelle tipologie
	
				4.1 Contributo in Atti di convegno

File in questo prodotto:

File	Dimensione	Formato
18128-Article Text-21623-1-2-20210521.pdf accesso riservato Tipologia: 2a Post-print versione editoriale / Version of Record Licenza: Non Pubblico - Accesso privato/ristretto Dimensione 3.11 MB Formato Adobe PDF Visualizza/Apri Richiedi una copia	3.11 MB	Adobe PDF	Visualizza/Apri Richiedi una copia
WikidataItemsCascades_dataset_camera_ready.pdf accesso aperto Tipologia: 2. Post-print / Author's Accepted Manuscript Licenza: Pubblico - Tutti i diritti riservati Dimensione 3.17 MB Formato Adobe PDF Visualizza/Apri	3.17 MB	Adobe PDF	Visualizza/Apri

Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11583/2920734

PORTO @ Archivio Istituzionale della Ricerca

Tracking Knowledge Propagation Across Wikipedia Languages

Valentim, Rodolfo Vieira;Comarela, Giovanni;Park, Souneil;Sáez-Trumper, Diego

2021

Abstract

Scheda breve Scheda completa Scheda completa (DC)

Pubblicazioni consigliate

Informazioni

Conferma cancellazione

Scheda breve

Scheda completa

Scheda completa (DC)