bioRxiv ScienceSearch

Biology subjects

Liu, S.

Publications and source records attributed to Liu, S..

80 records · Page 5Linked to original sources

Tunable Genotyping-By-Sequencing (tGBS(R)) Enables Reliable Genotyping of Heterozygous Loci

Most Genotyping-by-Sequencing (GBS) strategies suffer from high rates of missing data and high error rates, particularly at heterozygous sites. Tunable genotyping-by-sequencing (tGBS(R)), a novel genome reduction method, consists of the ligation of single-strand oligos to restriction enzyme fragments. DNA barcodes are added during PCR amplification; additional (selective) nucleotides included at the 3-end of the PCR primers result in more genome reduction as compared to conventional GBS methods. By adjusting the number of selective bases different numbers of genomic sites can be targeted for sequencing. Because this genome reduction strategy concentrates sequencing reads on fewer sites, SNP calls are based on more reads than conventional GBS, resulting in higher SNP calling accuracy (>97-99%) even for heterozygous sites and less missing data per marker. tGBS genotyping is expected to be particularly useful for genomic selection, which requires the ability to genotype populations of individuals that are heterozygous at many loci.

genomics

Structural basis of Mycobacterium tuberculosis transcription and transcription inhibition

One Sentence SummaryStructures of Mycobacterium tuberculosis RNA polymerase reveal taxon-specific properties and binding sites of known and new antituberculosis agents\n\nAbstractMycobacterium tuberculosis (Mtb) is the causative agent of tuberculosis, which kills 1.8 million annually. Mtb RNA polymerase (RNAP) is the target of the first-line antituberculosis drug rifampin (Rif). We report crystal structures of Mtb RNAP, alone and in complex with Rif. The results identify an Mtb-specific structural module of Mtb RNAP and establish that Rif functions by a steric-occlusion mechanism that prevents extension of RNA. We also report novel non-Rif-related compounds-N-aroyl-N-aryl-phenylalaninamides (AAPs)-that potently and selectively inhibit Mtb RNAP and Mtb growth, and we report crystal structures of Mtb RNAP in complex with AAPs. AAPs bind to a different site on Mtb RNAP than Rif, exhibit no cross-resistance with Rif, function additively when co-administered with Rif, and suppress resistance emergence when co-administered with Rif.

molecular biology

Genome sequence of a diabetes-prone desert rodent reveals a mutation hotspot around the ParaHox gene cluster

The sand rat Psammomys obesus is a gerbil native to deserts of North Africa and the Middle East1. Sand rats survive with low caloric intake and when given high carbohydrate diets can become obese and develop type II diabetes2 which, in extreme cases, leads to pancreatic failure and death3,4. Previous studies have reported inability to detect the Pdx1 gene or protein in gerbils5-7, suggesting that absence of this key insulin-regulating homeobox gene might underlie diabetes susceptibility. Here we report sequencing of the sand rat genome and discovery of an extensive, mutationally-biased GC-rich genomic domain encompassing many essential genes, including the elusive Pdx1. The sequence of Pdx1 has been grossly affected by GC-biased mutation leading to the highest divergence observed in the animal kingdom. In addition to molecular insights into restricted caloric intake in a desert species, the discovery that specific chromosomal regions can be subject to elevated mutation rate has widespread significance to evolution.

evolutionary biology

Chimeras Link to Tandem Repeats and Transposable Elements in Tetraploid Hybrid Fish

The formation of the allotetraploid hybrid lineage (4nAT) encompasses both distant hybridization and polyploidization processes. The allotetraploid offspring have two sets of sub-genomes inherited from both parental species and therefore it is important to explore its genetic structure. Herein, we construct a bacterial artificial chromosome library of allotetraploids, and then sequence and analyze the full-length sequences of 19 bacterial artificial chromosomes. Sixty-eight DNA chimeras are identified, which are divided into four models according to the distribution of the genomic DNA derived from the parents. Among the 68 genetic chimeras, 44 (64.71%) are linked to tandem repeats (TRs) and 23 (33.82%) are linked to transposable elements (TEs). The chimeras linked to TRs are related to slipped-strand mispairing and double-strand break repair while the chimeras linked to TEs are benefit from the intervention of recombinases. In addition, TRs and TEs are linked not only with the recombinations, but also with the insertions/deletions of DNA segments. We conclude that DNA chimeras accompanied by TRs and TEs coordinate a balance between the sub-genomes derived from the parents which reduces the genomic shock effects and favors the evolutionary and adaptive capacity of the allotetraploidization. It is the first report on the relationship between formation of the DNA chimeras and TRs and TEs in the polyploid animals.

genetics

Folding Principle of Chromosome Emerges from Mapping of Genome Features onto its 3D Structure

How chromosomes fold into 3D structures and how genome functions are affected or even controlled by their spatial organization remain challenging questions. Hi-C experiment has provided important structural insights for chromosome, and Hi-C data are used here to construct the 3D chromatin structure which are characterized by two spatially segregated chromatin compartments A and B. By mapping a plethora of genome features onto the constructed 3D chromatin model, we show vividly the close connection between genome properties and the spatial organization of chromatin. We are able to dissect the whole chromatin into two types of chromatin domains which have clearly different Hi-C contact patterns as well as different sizes of chromatin loops. The two chromatin types can be respectively regarded as the basic units of chromatin compartments A and B, and also spatially segregate from each other as the two chromatin compartments. Therefore, the chromatin loops segregate in the space according to their sizes, suggesting the excluded volume or entropic effect in chromatin compartmentalization as well as chromosome positioning. Taken together, these results provide clues to the folding principles of chromosomes, their spatial organization, and the resulted clustering of many genome features in the 3D space.

biophysics

Accurate prediction of human essential genes using only nucleotide composition and association information

Three groups recently identified essential genes in human cancer cell lines using wet experiments, and these genes are of high values. Herein, we improved the widely used Z curve method by creating a {lambda}-interval Z curve, which considered interval association information. With this method and recursive feature elimination technology, a computational model was developed to predict human gene essentiality. The 5-fold cross-validation test based on our benchmark dataset obtained an area under the receiver operating characteristic curve (AUC) of 0.8814. For the rigorous jackknife test, the AUC score was 0.8854. These results demonstrated that the essentiality of human genes could be reliably reflected by only sequence information. However, previous classifiers in three eukaryotes can gave satisfactory prediction only combining sequence with other features. It is also demonstrated that although the information contributed by interval association is less than adjacent nucleotides, this information can still play an independent role. Integrating the interval information into adjacent ones can significantly improve our classifiers prediction capacity. We re-predicted the benchmark negative dataset by Pheg server (https://cefg.uestc.edu.cn/Pheg), and 118 genes were additionally predicted as essential. Among them, 21 were found to be homologues in mouse essential genes, indicating that at least a part of the 118 genes were indeed essential, however previous experiments overlooked them. As the first available server, Pheg could predict essentiality for anonymous gene sequences of human. It is also hoped the {lambda}-interval Z curve method could be effectively extended to classification issues of other DNA elements.

bioinformatics

Chimeric Genes Revealed in the Polyploidy Fish Hybrids of Carassius cuvieri (Female) x Megalobrama amblycephala (Male)

The genomes of newly formed natural or artificial polyploids may experience rapid gene loss and genome restructuring. In this study, we obtained tetraploid hybrids (4n=148, 4nJB) and triploid hybrids (3n=124, 3nJB) derived from the hybridization of two different subfamily species Carassius cuvieri ([female], 2n = 100, JCC) and Megalobrama amblycephala ([male], 2n = 48, BSB). Some significant morphological and physiological differences were detected in the polyploidy hybrids compared with their parents. To reveal the molecular traits of the polyploids, we compared the liver transcriptomes of 4nJB, 3nJB and their parents. The results indicated high proportion chimeric genes (31 > %) and mutated orthologous genes (17 > %) both in 4nJB and 3nJB. We classified 10 gene patterns within three categories in 4nJB and 3nJB orthologous gene, and characterized 30 randomly chosen genes using genomic DNA to confirm the chimera or mutant. Moreover, we mapped chimeric genes involved pathways and discussed that the phenotypic novelty of the hybrids may relate to some chimeric genes. For example, we found there is an intragenic insertion in the K+ channel kcnk5b, which may be related to the novel presence of the barbels in 4nJB. Our results indicated that the genomes of newly formed polyploids experienced rapid restructuring post-polyploidization, which may results in the phenotypic and phenotypic changes among the polyploidy hybrid offspring. The formation of the 4nJB and 3nJB provided new insights into the genotypic and phenotypic diversity of hybrid fish resulting from distant hybridization between subfamilies.

developmental biology

DNA Methylation Landscape Reflects the Spatial Organization of Chromatin in Different Cells

The relation between DNA methylation and chromatin structure is still largely unknown. By analyzing a large set of sequencing data, we observed a long-range power law correlation of DNA methylation with cell-class-specific scaling exponents in the range of thousands to millions of base pairs. We showed such cell-class-specific scaling exponents are caused by different patchiness of DNA methylation in different cells. By modeling the chromatin structure using Hi-C data and mapping the methylation level onto the modeled structure, we demonstrated the patchiness of DNA methylation is related to chromatin structure. The scaling exponents of the power law correlation is thus a display of the spatial organization of chromatin. Besides, the local correlation of DNA methylation is associated with nucleosome positioning and different between partially-methylated-domain and non-partially-methylated-domain, suggesting their different chromatin structures at several nucleosomes level. Our study provides a novel view of the spatial organization of chromatin structure from a perspective of DNA methylation, in which both long-range and local correlations of DNA methylation along the genome reflect the spatial organization of chromatin.

biophysics