Skip to content
Open access

Regional diversity meets global representation in large-scale human pangenomics

· The Innovation Life · 0 citations · 10 references

Abstract

Human pangenomics is moving beyond the single linear reference paradigm towards population-scale representations that better capture complex genomic variation. Many current human pangenomes include both global and regional initiatives, but their limited sample sizes constrain the representation of population diversity. Recently, the establishment of the 1000 Chinese Pangenome (1KCP) project represents a major milestone in large-scale regional pangenomics. By combining high-coverage de novo assemblies with pangenome-informed assembly of modest-coverage samples, 1KCP expanded the resource to 1,116 diploid genome assemblies while substantially reducing sequencing costs relative to uniformly high-coverage de novo assembly. This large-scale framework not only broadens the catalog of non-reference sequences and genetic variants, but also enables complex variants, including structural variants and tandem repeats, to be linked to gene regulation and human phenotypes, revealing classes of variation incompletely captured by conventional small-variant-centered analyses. Nevertheless, some challenges remain, including reduced assembly accuracy in repetitive and structurally complex genomic regions when using modest-coverage data. Broader and more balanced sampling will also be needed to better represent population diversity. Building on these insights, we propose a hierarchical framework for the next-generation global pangenome, where large-scale regional pangenomes guide diversity-informed selection of representative individuals for high-quality genome assembly, while global allele frequencies are integrated with the resulting pangenome graph.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.