Benchmarking unsupervised near-duplicate image detection

Morra, Lia; Lamberti, Fabrizio

doi:10.1016/j.eswa.2019.05.002

Unsupervised near-duplicate detection has many practical applications ranging from social media analysis and web-scale retrieval, to digital image forensics. It entails running a threshold-limited query on a set of descriptors extracted from the images, with the goal of identifying all possible near-duplicates, while limiting the false positives due to visually similar images. Since the rate of false alarms grows with the dataset size, a very high specificity is thus required, up to 1-10^-9 for realistic use cases; this important requirement, however, is often overlooked in literature. In recent years, descriptors based on deep convolutional neural networks have matched or surpassed traditional feature extraction methods in content-based image retrieval tasks. To the best of our knowledge, ours is the first attempt to establish the performance range of deep learning-based descriptors for unsupervised near-duplicate detection on a range of datasets, encompassing a broad spectrum of near-duplicate definitions. We leverage both established and new benchmarks, such as the Mir-Flick Near-Duplicate (MFND) dataset, in which a known ground truth is provided for all possible pairs over a general, large scale image collection. To compare the specificity of different descriptors, we reduce the problem of unsupervised detection to that of binary classification of near-duplicate vs. not-near-duplicate images. The latter can be conveniently characterized using Receiver Operating Curve (ROC). Our findings in general favor the choice of fine-tuning deep convolutional networks, as opposed to using off-the-shelf features, but differences at high specificity settings depend on the dataset and are often small. The best performance was observed on the MFND benchmark, achieving 96% sensitivity at a false positive rate of 1.43x10^-6.

Benchmarking unsupervised near-duplicate image detection / Morra, Lia; Lamberti, Fabrizio. - In: EXPERT SYSTEMS WITH APPLICATIONS. - ISSN 0957-4174. - STAMPA. - 135:(2019), pp. 313-326. [10.1016/j.eswa.2019.05.002]

Benchmarking unsupervised near-duplicate image detection

Lia Morra;Fabrizio Lamberti

2019

Abstract

Scheda breve

Scheda completa

Scheda completa (DC)

	Anno del prodotto
	
				2019
			
	Codice DOI
	
				https://dx.doi.org/10.1016/j.eswa.2019.05.002
			
	Titolo della Rivista
	
				EXPERT SYSTEMS WITH APPLICATIONS
			
	Appare nelle tipologie
	
				1.1 Articolo in rivista

File in questo prodotto:

File	Dimensione	Formato
preprint_lq.pdf Open Access dal 09/05/2021 Descrizione: Postprint Tipologia: 2. Post-print / Author's Accepted Manuscript Licenza: Creative commons Dimensione 9.51 MB Formato Adobe PDF Visualizza/Apri	9.51 MB	Adobe PDF	Visualizza/Apri
1-s2.0-S095741741930315X-main.pdf accesso riservato Descrizione: Post-print versione editoriale Tipologia: 2a Post-print versione editoriale / Version of Record Licenza: Non Pubblico - Accesso privato/ristretto Dimensione 3.47 MB Formato Adobe PDF Visualizza/Apri Richiedi una copia	3.47 MB	Adobe PDF	Visualizza/Apri Richiedi una copia

Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11583/2733022

PORTO @ Archivio Istituzionale della Ricerca

Benchmarking unsupervised near-duplicate image detection

Lia Morra;Fabrizio Lamberti

2019

Abstract

Scheda breve Scheda completa Scheda completa (DC)

Pubblicazioni consigliate

Informazioni

Conferma cancellazione

Scheda breve

Scheda completa

Scheda completa (DC)