Accurate anatomical landmark identification is important for safe laparoscopic navigation, yet limited view and strong tissue similarity make multi-label organ classification diffi- cult. Existing vision-language models mainly rely on appearance and overlook the spatial structure of surgical scenes (Zhang et al., 2025). We propose SpatialContext, a multi- modal framework that injects scene geometry into classification through natural language prompts derived from segmentation masks, together with a context-conditional training strategy centered on the primary surgical target. Results on DSAD (Carstens et al., 2023) and Endoscapes (Mascagni et al., 2025) show improved recognition of scene-defining and off-target anatomy, suggesting that explicit spatial semantics can improve surgical scene understanding.
Injecting Textual Spatial Context into Vision-Language Models for Surgical Scene Understanding / Martignano, A., Guo, L., Camerota, C., Sacco, A., Esposito, F.. - ELETTRONICO. - (2026), pp. 1-5. (Medical Imaging with Deep Learning (MIDL) 2025 Taipei, Taiwan July 8–10, 2026).
Injecting Textual Spatial Context into Vision-Language Models for Surgical Scene Understanding
Alessio Sacco;
2026
Abstract
Accurate anatomical landmark identification is important for safe laparoscopic navigation, yet limited view and strong tissue similarity make multi-label organ classification diffi- cult. Existing vision-language models mainly rely on appearance and overlook the spatial structure of surgical scenes (Zhang et al., 2025). We propose SpatialContext, a multi- modal framework that injects scene geometry into classification through natural language prompts derived from segmentation masks, together with a context-conditional training strategy centered on the primary surgical target. Results on DSAD (Carstens et al., 2023) and Endoscapes (Mascagni et al., 2025) show improved recognition of scene-defining and off-target anatomy, suggesting that explicit spatial semantics can improve surgical scene understanding.| File | Dimensione | Formato | |
|---|---|---|---|
|
109_Injecting_Textual_Spatial_.pdf
accesso aperto
Tipologia:
2. Post-print / Author's Accepted Manuscript
Licenza:
Creative commons
Dimensione
554.72 kB
Formato
Adobe PDF
|
554.72 kB | Adobe PDF | Visualizza/Apri |
Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/11583/3013598
