Entropy regularization is a recurring mechanism in reinforcement learning (RL), but its meaning changes across algorithmic settings. In classical online RL, entropy encourages exploration and smooths policy improvement; in inverse RL and imitation learning, maximum-entropy resolves ambiguity among expert-consistent behaviors; in offline RL, entropy must be balanced against data support; in generative policies, entropy becomes a tractability problem; and in reinforcement learning with verifiable rewards (RLVR) for large language models (LLMs), token entropy is tied to reasoning diversity, calibration, and collapse. This review organizes these developments into a unified taxonomy. We first summarize the mathematical foundations of maximum-entropy RL, soft Bellman equations, policy-gradient entropy dynamics, and Kullback–Leibler (KL)-constrained mirror descent. We then review entropy in imitation learning, offline RL, intrinsic motivation, diffusion and flow-based policy classes, and RLVR. Particular attention is given to recent work on entropy collapse in reasoning LLMs, entropy-based advantage shaping, covariance-based control, positive-advantage reweighting, and ordinary differential equation (ODE)-based flow-matching policies with tractable entropy. The review emphasizes that entropy is not universally beneficial: useful exploration, support preservation, multimodality, calibration, and reasoning diversity require different entropy objects and different control mechanisms.
Entropy Regularization in Deep Reinforcement Learning: A Structured Review Across Classical Control, Generative Policies, and Reasoning Language Models / Taricco, G.. - In: ENTROPY. - ISSN 1099-4300. - 28:7(2026). [10.3390/e28070811]
Entropy Regularization in Deep Reinforcement Learning: A Structured Review Across Classical Control, Generative Policies, and Reasoning Language Models
Giorgio Taricco
2026
Abstract
Entropy regularization is a recurring mechanism in reinforcement learning (RL), but its meaning changes across algorithmic settings. In classical online RL, entropy encourages exploration and smooths policy improvement; in inverse RL and imitation learning, maximum-entropy resolves ambiguity among expert-consistent behaviors; in offline RL, entropy must be balanced against data support; in generative policies, entropy becomes a tractability problem; and in reinforcement learning with verifiable rewards (RLVR) for large language models (LLMs), token entropy is tied to reasoning diversity, calibration, and collapse. This review organizes these developments into a unified taxonomy. We first summarize the mathematical foundations of maximum-entropy RL, soft Bellman equations, policy-gradient entropy dynamics, and Kullback–Leibler (KL)-constrained mirror descent. We then review entropy in imitation learning, offline RL, intrinsic motivation, diffusion and flow-based policy classes, and RLVR. Particular attention is given to recent work on entropy collapse in reasoning LLMs, entropy-based advantage shaping, covariance-based control, positive-advantage reweighting, and ordinary differential equation (ODE)-based flow-matching policies with tractable entropy. The review emphasizes that entropy is not universally beneficial: useful exploration, support preservation, multimodality, calibration, and reasoning diversity require different entropy objects and different control mechanisms.| File | Dimensione | Formato | |
|---|---|---|---|
|
entropy-28-00811.pdf
accesso aperto
Tipologia:
2a Post-print versione editoriale / Version of Record
Licenza:
Creative commons
Dimensione
567.47 kB
Formato
Adobe PDF
|
567.47 kB | Adobe PDF | Visualizza/Apri |
Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/11583/3013230
