Mind the Error! Detection and Localization of Instruction Errors in Vision-and-Language Navigation

Taioli, Francesco; Rosa, Stefano; Castellini, Alberto; Natale, Lorenzo; Del Bue, Alessio; Farinelli, Alessandro; Cristani, Marco; Wang, Yiming

doi:10.1109/iros58592.2024.10801822

Vision-and-Language Navigation in Continuous Environments (VLN-CE) is one of the most intuitive yet challenging embodied AI tasks. Agents are tasked to navigate towards a target goal by executing a set of low-level actions, following a series of natural language instructions. All VLN-CE methods in the literature assume that language instructions are exact. However, in practice, instructions given by humans can contain errors when describing a spatial environment due to inaccurate memory or confusion. Current VLN-CE benchmarks do not address this scenario, making the state-of-the-art methods in VLN-CE fragile in the presence of erroneous instructions from human users. For the first time, we propose a novel benchmark dataset that introduces various types of instruction errors considering potential human causes. This benchmark provides valuable insight into the robustness of VLN systems in continuous environments. We observe a noticeable performance drop (up to −25%) in Success Rate when evaluating the state-of-the-art VLN-CE methods on our benchmark. Moreover, we formally define the task of Instruction Error Detection and Localization, and establish an evaluation protocol on top of our benchmark dataset. We also propose an effective method, based on a cross-modal transformer architecture, that achieves the best performance in error detection and localization, compared to baselines. Surprisingly, our proposed method has revealed errors in the validation set of the two commonly used datasets for VLN-CE, i.e., R2R-CE and RxR-CE, demonstrating the utility of our technique in other tasks.

Mind the Error! Detection and Localization of Instruction Errors in Vision-and-Language Navigation / Taioli, Francesco; Rosa, Stefano; Castellini, Alberto; Natale, Lorenzo; Del Bue, Alessio; Farinelli, Alessandro; Cristani, Marco; Wang, Yiming. - ELETTRONICO. - (2024), pp. 12993-13000. (Intervento presentato al convegno 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) tenutosi a Abu Dhabi (ARE) nel 14-18 October 2024) [10.1109/iros58592.2024.10801822].

Mind the Error! Detection and Localization of Instruction Errors in Vision-and-Language Navigation

Taioli, Francesco;Rosa, Stefano;Castellini, Alberto;Natale, Lorenzo;Del Bue, Alessio;Farinelli, Alessandro;Cristani, Marco;Wang, Yiming

2024

Abstract

Scheda breve

Scheda completa

Scheda completa (DC)

	Anno del prodotto
	
				2024
			
	Codice ISBN
	
				979-8-3503-7770-5
			
	Appare nelle tipologie
	
				4.1 Contributo in Atti di convegno

File in questo prodotto:

File	Dimensione	Formato
Mind_the_Error_Detection_and_Localization_of_Instruction_Errors_in_Vision-and-Language_Navigation.pdf accesso riservato Tipologia: 2a Post-print versione editoriale / Version of Record Licenza: Non Pubblico - Accesso privato/ristretto Dimensione 2.19 MB Formato Adobe PDF Visualizza/Apri Richiedi una copia	2.19 MB	Adobe PDF	Visualizza/Apri Richiedi una copia

Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11583/2996560

PORTO @ Archivio Istituzionale della Ricerca

Mind the Error! Detection and Localization of Instruction Errors in Vision-and-Language Navigation

Taioli, Francesco;Rosa, Stefano;Castellini, Alberto;Natale, Lorenzo;Del Bue, Alessio;Farinelli, Alessandro;Cristani, Marco;Wang, Yiming

2024

Abstract

Scheda breve Scheda completa Scheda completa (DC)

Pubblicazioni consigliate

Informazioni

Conferma cancellazione

Scheda breve

Scheda completa

Scheda completa (DC)