Large-scale genomic and population-genetic datasets offer unprecedented opportunities to study population structure and uncover the genetic basis of complex traits and diseases. Existing analytical tools, however, are characterized by format incompatibilities, limited functionality, and computational inefficiencies, forcing researchers to construct pipelines chaining together fragmented command-line utilities and ad hoc scripts. These are difficult to maintain, scale, and reproduce. We present snputils, an open-source Python library for high-performance processing and analysis of genotype, ancestry, phenotype, and identity-by-descent data within a single framework suitable for biobank-scale research. snputils provides unified data containers and efficient routines for genotype quality control, filtering, merging, and computing classical population-genetic statistics, with optional ancestry-specific masking. An identity-by-descent module supports reading multiple formats, filtering, and ancestry-restricted segment trimming for relatedness and demographic inference. snputils also incorporates ancestry masking and multi-array functionalities for dimensionality reduction methods, as well as efficient implementations of admixture simulation, admixture mapping, genome-wide association testing, and advanced visualization capabilities. With support for the most commonly used file formats, snputils natively supports human and non-human diploid datasets represented in standard formats. At the same time, its modular and optimized design reduces technical overhead, facilitating reproducible workflows that accelerate discoveries in population genetics, genomic research, and precision medicine. Benchmarks show faster I/O and concordant PCA, allele-frequency statistics, and admixture mapping results relative to established reference tools, with substantial runtime reductions. snputils is available at https://github.com/AI-sandbox/snputils, with documentation and tutorials at docs.snputils.org.
snputils: A High-Performance Python Library for Genetic Variation and Population Structure / Bonet, D., Comajoan Cara, M., Barrabés, M., Smeriglio, R., Geleta, M., Dominguez Mantes, A., Thomassin, C., Agrawal, D., Shanks, C., Huang, E.C., Aounallah, K., Franquesa Monés, M., Luis, A., Saurina, J., Perera, M., López, C., Jaras, A., Oriol Sabat, B., Abante, J., Moreno-Grau, S., et al.. - In: MOLECULAR BIOLOGY AND EVOLUTION. - ISSN 1537-1719. - (In corso di stampa).
snputils: A High-Performance Python Library for Genetic Variation and Population Structure
Smeriglio, Riccardo;
In corso di stampa
Abstract
Large-scale genomic and population-genetic datasets offer unprecedented opportunities to study population structure and uncover the genetic basis of complex traits and diseases. Existing analytical tools, however, are characterized by format incompatibilities, limited functionality, and computational inefficiencies, forcing researchers to construct pipelines chaining together fragmented command-line utilities and ad hoc scripts. These are difficult to maintain, scale, and reproduce. We present snputils, an open-source Python library for high-performance processing and analysis of genotype, ancestry, phenotype, and identity-by-descent data within a single framework suitable for biobank-scale research. snputils provides unified data containers and efficient routines for genotype quality control, filtering, merging, and computing classical population-genetic statistics, with optional ancestry-specific masking. An identity-by-descent module supports reading multiple formats, filtering, and ancestry-restricted segment trimming for relatedness and demographic inference. snputils also incorporates ancestry masking and multi-array functionalities for dimensionality reduction methods, as well as efficient implementations of admixture simulation, admixture mapping, genome-wide association testing, and advanced visualization capabilities. With support for the most commonly used file formats, snputils natively supports human and non-human diploid datasets represented in standard formats. At the same time, its modular and optimized design reduces technical overhead, facilitating reproducible workflows that accelerate discoveries in population genetics, genomic research, and precision medicine. Benchmarks show faster I/O and concordant PCA, allele-frequency statistics, and admixture mapping results relative to established reference tools, with substantial runtime reductions. snputils is available at https://github.com/AI-sandbox/snputils, with documentation and tutorials at docs.snputils.org.| File | Dimensione | Formato | |
|---|---|---|---|
|
snputils.pdf
accesso aperto
Descrizione: Accepted Manuscript
Tipologia:
2. Post-print / Author's Accepted Manuscript
Licenza:
Creative commons
Dimensione
1.01 MB
Formato
Adobe PDF
|
1.01 MB | Adobe PDF | Visualizza/Apri |
Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.
https://hdl.handle.net/11583/3016081
