Large Language Models (LLMs) are increasingly adopted for legal document understanding by attorneys and legal consultants. Despite advances in adapting LLMs to their legal terminology and domain-specific linguistic nuances, the LLMs’ ability to reason about temporal relations in legal documents remains largely underexplored. In this work, we explore the capabilities of LLMs to verify the correctness of a legal temporal ordering clause and to classify the type of temporal relationships between two legal entities. The results achieved on a public Englishwritten benchmark show that (1) instruction-based models generally perform better than the corresponding chat versions; (2) LLMs reasoning capabilities are, typically, marginally useful to address the specific temporal reasoning tasks; (3) LLMs under a Few-Shot Learning (FSL) setting turn out to be the most effective, with Grok 4 surpassing the state of the art.
Exploring In-Context Learning Strategies for Temporal Ordering of Legal Events using Large Language Models / Cacioli, A., Cagliero, L., Tarasconi, F.. - 4192:(2026). (EDBT/ICDT-WS 2026 EDBT/ICDT 2026 Workshops Tampere (Finland) March 24, 2026.).
Exploring In-Context Learning Strategies for Temporal Ordering of Legal Events using Large Language Models.
Cacioli, Andrea;Cagliero, Luca;
2026
Abstract
Large Language Models (LLMs) are increasingly adopted for legal document understanding by attorneys and legal consultants. Despite advances in adapting LLMs to their legal terminology and domain-specific linguistic nuances, the LLMs’ ability to reason about temporal relations in legal documents remains largely underexplored. In this work, we explore the capabilities of LLMs to verify the correctness of a legal temporal ordering clause and to classify the type of temporal relationships between two legal entities. The results achieved on a public Englishwritten benchmark show that (1) instruction-based models generally perform better than the corresponding chat versions; (2) LLMs reasoning capabilities are, typically, marginally useful to address the specific temporal reasoning tasks; (3) LLMs under a Few-Shot Learning (FSL) setting turn out to be the most effective, with Grok 4 surpassing the state of the art.| File | Dimensione | Formato | |
|---|---|---|---|
|
DARLIAP-paper8.pdf
accesso aperto
Tipologia:
2a Post-print versione editoriale / Version of Record
Licenza:
Creative commons
Dimensione
298.97 kB
Formato
Adobe PDF
|
298.97 kB | Adobe PDF | Visualizza/Apri |
Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/11583/3015321
