Skip to content

Author

Daofeng Li

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jun 2026

Unlocking cis-regulatory landscapes across 500 million years of evolution and disease mechanisms

Abstract Genomic DNA encodes regulatory information that determines where, when, and to what extent genes are expressed. Theoretically, we should be able to identify these transcriptional “instructions” by examining genomic DNA sequence alone, yet this has remained challenging. Here we present the Vertebrate Regulatory MOdule Detector (VRMOD), a method that accurately predicts gene regulatory sequences using only the query genomic sequences. We applied VRMOD to 309 Ensembl genomes, generating a compendium of high-resolution, genome-position-fixed cis-regulatory modules without parameter tuning. We performed extensive computational evaluation and experimental validation of VRMOD predictions. Notably, VRMOD predicted three sub-enhancers within the human hs52 enhancer at the FTO locus from the VISTA database, including one missed by existing methods. Using a chicken embryo system and 3D tissue imaging, we showed that each sub-enhancer exhibits restricted spatiotemporal activity within specific subsets of tissues where the full enhancer is active. We further demonstrated VRMOD’s utility for identifying evolutionarily non-conserved enhancers, annotating regulatory sequences in non-model organisms, and identifying candidate disease-causal variants. Collectively, VRMOD provides a universal coordinate reference system for regulatory sequences across 309 vertebrate genomes and enables genome-wide annotation of non-coding regulatory elements in any vertebrate species using genomic sequence alone.

Tássia Mangetti Gonçalves, Casey L. Stewart, Samantha D. Baxley et al. · 0 citations
Open access Jul 2026

HPRC2: A human pangenome reference with near-complete coverage of common genetic variation

A pangenome reference overcomes the inherent limitation of any individual reference genome by integrating the variation present in a population. We present the Human Pangenome Reference Consortium’s (HPRC) Release 2 (HPRC2), an openly available, second phase pangenome that is an approximately fivefold expansion in genome number over HPRC Release 1 (HPRC1) and measurable improvement in genome completeness, contiguity, and accuracy. Selecting samples with a principled algorithm prioritising common variant coverage, HPRC2 contributes 460 haplotypes that together capture over 99% of common variation observed in the All of Us Research Program v8 cohort. Combining high-coverage long and ultra-long reads with modern assemblers and polishers, we produce thousands of telomere-to-telomere (T2T) chromosomes, and relative to HPRC1 halve the number of structurally unreliable regions as well as individual base errors per haplotype. We complement the assemblies with whole genome multiple alignments and gene annotations, and derive formal pangenome coordinate systems for addressing off-reference variation, demonstrating that individual human genomes contain more than one hundred thousand variants not succinctly described with respect to existing reference genomes. We also present the first matched long-read backed pantranscriptome and panepigenome at this scale, provide continuous local-ancestry estimates spanning every genome, and outline a host of new tools and applications that leverage the pangenome resource for improved genomics analysis.

Julian K. Lucas, Prajna Hebbar, Wen-Wei Liao et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.