bioRxiv Science⌕ Search

Biology subjects

luo, y.

Publications and source records attributed to luo, y..

2 recordsLinked to original sources

AutoHiC: a deep-learning method for automatic and accurate chromosome-level genome assembly

An accurate genome at the chromosome level is the key to unraveling the mysteries of gene function and unlocking the mechanisms of disease. Irrespective of the sequencing methodology adopted, Hi-C aided scaffolding serves as a principal avenue for generating genome assemblies at the chromosomal level. However, the results of such scaffolding are often flawed and require extensive manual refinement. In this paper, we introduce AutoHiC, an innovative deep learning-based tool designed to identify and rectify genome assembly errors. Diverging from conventional approaches, AutoHiC harnesses the power of high-dimensional Hi-C data to enhance genome continuity and accuracy through a fully automated workflow and iterative error correction mechanism. AutoHiC was trained on Hi-C data from more than 300 species (approximately five hundred thousand interaction maps) in DNA Zoo and NCBI. Its confusion matrix results show that the average error detection accuracy is over 90%, and the area under the precision-recall curve is close to 1, making it a powerful error detection capability. The benchmarking results demonstrate AutoHiCs ability to substantially enhance genome continuity and significantly reduce error rates, providing a more reliable foundation for genomics research. Furthermore, AutoHiC generates comprehensive result reports, offering users insights into the assembly process and outcomes. In summary, AutoHiC represents a breakthrough in automated error detection and correction for genome assembly, effectively promoting more accurate and comprehensive genome assemblies.

bioinformatics↗

PlantCADB: A comprehensive plant chromatin accessibility database

Chromatin accessibility landscapes are essential for detecting regulatory elements, illustrating the corresponding regulatory networks, and, ultimately, understanding the molecular bases underlying key biological processes. With the advancement of sequencing technologies, a large volume of chromatin accessibility data has been accumulated and integrated in humans and other mammals. These data have greatly advanced the study of disease pathogenesis, cancer survival prognosis, and tissue development. To advance the understanding of molecular mechanisms regulating plant key traits and biological processes, we developed a comprehensive plant chromatin accessibility database (PlantCADB, https://bioinfor.nefu.edu.cn/PlantCADB/) from 649 samples of 37 species. Among these samples, 159 are abiotic stress-related (including heat, cold, drought, salt, etc.), 232 are development-related and 376 are tissue-specific. Overall, 18,339,426 accessible chromatin regions (ACRs) were compiled. These ACRs were annotated with genomic information, associated genes, transcription factors footprint, motif, and SNPs. Additionally, PlantCADB provides various tools to visualize ACRs and corresponding annotations. It thus forms an integrated, annotated, and analyzed plant-related chromatin accessibility information which can aid to better understand genetic regulatory networks underlying development, important traits, stress adaptions, and evolution.

genomics↗