Long-memory correlation among individual processes and components yield information about how complex systems evolve. Well-established methods to discriminate correlated and anticorrelated sequences (e.g. boxcounting (BC), rescaled range (R/S), detrended fluctuation (DFA), and detrended moving average (DMA) analysis) mainly operate via linear regressions of log-log data plots. The Kullback-Leibler cluster entropy (KLCE), a measure of divergence between probability distributions, has been recently proposed to estimate the scaling exponent of long-range correlated sequences. Here, the KLCE is illustrated on the 24 chromosomes of the T2T-CHM13+Y human reference genome, including centromeric regions and short arms last remaining gaps in the GRCh38 reference assembly. It is shown that the Kullback-Leibler cluster entropy can accurately detect such structurally complex nucleotide regions through the local variability of the correlation exponents. To test performance and prove accuracy in detecting the newly added strands, the outcomes are mapped against nucleotide positions of the GRCh38 assembly. Statistical tests are also included to prove KLCE robustness and accuracy against DFA and DMA performances. Potential applications of the Kullback-Leibler cluster divergence in machine-learning and deep-learning pipelines for genomic data are outlined.
Kullback–Leibler cluster entropy: a proxy of human chromosome complexity / Gandino, F., Ferrero, R., Carbone, A.. - In: JOURNAL OF COMPUTATIONAL SCIENCE. - ISSN 1877-7511. - STAMPA. - 100:(2026). [10.1016/j.jocs.2026.102964]
Kullback–Leibler cluster entropy: a proxy of human chromosome complexity
Filippo Gandino;Renato Ferrero;Anna Carbone
2026
Abstract
Long-memory correlation among individual processes and components yield information about how complex systems evolve. Well-established methods to discriminate correlated and anticorrelated sequences (e.g. boxcounting (BC), rescaled range (R/S), detrended fluctuation (DFA), and detrended moving average (DMA) analysis) mainly operate via linear regressions of log-log data plots. The Kullback-Leibler cluster entropy (KLCE), a measure of divergence between probability distributions, has been recently proposed to estimate the scaling exponent of long-range correlated sequences. Here, the KLCE is illustrated on the 24 chromosomes of the T2T-CHM13+Y human reference genome, including centromeric regions and short arms last remaining gaps in the GRCh38 reference assembly. It is shown that the Kullback-Leibler cluster entropy can accurately detect such structurally complex nucleotide regions through the local variability of the correlation exponents. To test performance and prove accuracy in detecting the newly added strands, the outcomes are mapped against nucleotide positions of the GRCh38 assembly. Statistical tests are also included to prove KLCE robustness and accuracy against DFA and DMA performances. Potential applications of the Kullback-Leibler cluster divergence in machine-learning and deep-learning pipelines for genomic data are outlined.| File | Dimensione | Formato | |
|---|---|---|---|
|
Kullback–Leibler Cluster Entropy A proxy of human chromosome complexity.pdf
accesso aperto
Descrizione: versione pubblicata - repository istituzionale
Tipologia:
2a Post-print versione editoriale / Version of Record
Licenza:
Creative commons
Dimensione
1.72 MB
Formato
Adobe PDF
|
1.72 MB | Adobe PDF | Visualizza/Apri |
Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/11583/3015024
