Accurate anatomical landmark identification is important for safe laparoscopic navigation, yet limited view and strong tissue similarity make multi-label organ classification diffi- cult. Existing vision-language models mainly rely on appearance and overlook the spatial structure of surgical scenes (Zhang et al., 2025). We propose SpatialContext, a multi- modal framework that injects scene geometry into classification through natural language prompts derived from segmentation masks, together with a context-conditional training strategy centered on the primary surgical target. Results on DSAD (Carstens et al., 2023) and Endoscapes (Mascagni et al., 2025) show improved recognition of scene-defining and off-target anatomy, suggesting that explicit spatial semantics can improve surgical scene understanding.

Injecting Textual Spatial Context into Vision-Language Models for Surgical Scene Understanding / Martignano, A., Guo, L., Camerota, C., Sacco, A., Esposito, F.. - ELETTRONICO. - (2026), pp. 1-5. (Medical Imaging with Deep Learning (MIDL) 2025 Taipei, Taiwan July 8–10, 2026).

Injecting Textual Spatial Context into Vision-Language Models for Surgical Scene Understanding

Alessio Sacco;
2026

Abstract

Accurate anatomical landmark identification is important for safe laparoscopic navigation, yet limited view and strong tissue similarity make multi-label organ classification diffi- cult. Existing vision-language models mainly rely on appearance and overlook the spatial structure of surgical scenes (Zhang et al., 2025). We propose SpatialContext, a multi- modal framework that injects scene geometry into classification through natural language prompts derived from segmentation masks, together with a context-conditional training strategy centered on the primary surgical target. Results on DSAD (Carstens et al., 2023) and Endoscapes (Mascagni et al., 2025) show improved recognition of scene-defining and off-target anatomy, suggesting that explicit spatial semantics can improve surgical scene understanding.
File in questo prodotto:
File Dimensione Formato  
109_Injecting_Textual_Spatial_.pdf

accesso aperto

Tipologia: 2. Post-print / Author's Accepted Manuscript
Licenza: Creative commons
Dimensione 554.72 kB
Formato Adobe PDF
554.72 kB Adobe PDF Visualizza/Apri
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11583/3013598