bioRxiv ScienceSearch

Biology subjects

Leung, A. W.-S.

Publications and source records attributed to Leung, A. W.-S..

3 recordsLinked to original sources

HKG: An open genetic variant database of 205 Hong Kong Cantonese exomes

HKG is the first fully accessible variant database for Hong Kong Cantonese, constructed from 205 novel whole-exome sequencing data. There has long been a research gap in the understanding of the genetic architecture of southern Chinese subgroups, including Hong Kong Cantonese. HKG detected 196,325 high-quality variants with 5.93% being novel, and 25,472 variants were found to be unique in HKG compared to other Chinese populations (CHN). PCA illustrates the uniqueness of HKG in CHN, and IBD analysis revealed that it is related mostly to southern Chinese with a similar effective population size. An admixture study estimated the ancestral composition of HKG and CHN, with a gradient change from north to south, consistent with their geological distribution. ClinVar, CIViC and PharmGKB annotated 599 clinically significant variants and 360 putative loss-of-function variants, substantiating our understanding of population characteristics for future medical development. Among the novel variants, 96.57% were singleton and 6.85% were of high impact. With a good representation of Hong Kong Cantonese, we demonstrated better variant imputation using reference with the addition of HKG data, thus successfully filling the data gap in southern Chinese to facilitate the regional and global development of population genetics.

bioinformatics

DNA methylation affects pre-mRNA transcriptional initiation and processing in Arabidopsis

BackgroundDNA methylation may regulate pre-mRNA transcriptional initiation and processing, thus affecting gene expression. Unlike animal cells, plants, especially Arabidopsis thaliana, have relatively low DNA methylation levels, limiting our ability to observe any correlation between DNA methylation and pre-mRNA processing using typical short-read sequencing. However, with newly developed long-read sequencing technologies, such as Oxford Nanopore Technology Direct RNA sequencing (ONT DRS), combined with whole-genome bisulfite sequencing, we were able to precisely analyze the relationship between DNA methylation and pre-mRNA transcriptional initiation and processing using DNA methylation-related mutants. ResultsUsing ONT DRS, we generated more than 2 million high-quality full-length long reads of native mRNA for each of the wild type Col-0 and mutants defective in DNA methylation, identifying a total of 117,474 isoforms. We found that low DNA methylation levels around splicing sites tended to prevent splicing events from occurring. The lengths of the poly(A) tail of mRNAs were positively correlated with DNA methylation. DNA methylation before transcription start sites or around transcription termination sites tended to result in gene-silencing or read-through events. Furthermore, using ONT DRS, we identified novel transcripts that we could not have otherwise, since transcripts with intron retention and fusion transcripts containing the uncut intergenic sequence tend not to be exported to the cytoplasm. Using the met1-3 mutant with activated constitutive heterochromatin regions, we confirmed the effects of DNA methylation on pre-mRNA processing. ConclusionThe combination of ONT DRS with whole-genome bisulfite sequencing was a powerful tool for studying the effects of DNA methylation on splicing site selection and pre-mRNA processing, and therefore regulation of gene expression.

plant biology

ECNano: A Cost-Effective Workflow for Target Enrichment Sequencing and Accurate Variant Calling on 4,800 Clinically Significant Genes Using a Single MinION Flowcell

BackgroundThe application of long-read sequencing using the Oxford Nanopore Technologies (ONT) MinION sequencer is getting more diverse in the medical field. Having a high sequencing error of ONT and limited throughput from a single MinION flowcell, however, limits its applicability for accurate variant detection. Medical exome sequencing (MES) targets clinically significant exon regions, allowing rapid and comprehensive screening of pathogenic variants. By applying MES with MinION sequencing, the technology can achieve a more uniform capture of the target regions, shorter turnaround time, and lower sequencing cost per sample. MethodWe introduced a cost-effective optimized workflow, ECNano, comprising a wet-lab protocol and bioinformatics analysis, for accurate variant detection at 4,800 clinically important genes and regions using a single MinION flowcell. The ECNano wet-lab protocol was optimized to perform long-read target enrichment and ONT library preparation to stably generate high-quality MES data with adequate coverage. The subsequent variant-calling workflow, Clair-ensemble, adopted a fast RNN-based variant caller, Clair, and was optimized for target enrichment data. To evaluate its performance and practicality, ECNano was tested on both reference DNA samples and patient samples. ResultsECNano achieved deep on-target depth of coverage (DoC) at average >100x and >98% uniformity using one MinION flowcell. For accurate ONT variant calling, the generated reads sufficiently covered 98.9% of pathogenic positions listed in ClinVar, with 98.96% having at least 30x DoC. ECNano obtained an average read length of 1,000 bp. The long reads of ECNano also covered the adjacent splice sites well, with 98.5% of positions having [≥] 30x DoC. Clair-ensemble achieved >99% recall and accuracy for SNV calling. The whole workflow from wet-lab protocol to variant detection was completed within three days. ConclusionWe presented ECNano, an out-of-the-box workflow comprising (1) a wet-lab protocol for ONT target enrichment sequencing and (2) a downstream variant detection workflow, Clair-ensemble. The workflow is cost-effective, with a short turnaround time for high accuracy variant calling in 4,800 clinically significant genes and regions using a single MinION flowcell. The long-read exon captured data has potential for further development, promoting the application of long-read sequencing in personalized disease treatment and risk prediction.

bioinformatics