bioRxiv ScienceSearch

Biology subjects

Bajic, V.

Publications and source records attributed to Bajic, V..

2 recordsLinked to original sources

Characterization and identification of long non-coding RNAs based on feature relationship

The significance of long non-coding RNAs (lncRNAs) in many biological processes and diseases has gained intense interests over the past several years. However, computational identification of lncRNAs in a wide range of species remains challenging; it requires prior knowledge of well-established sequences and annotations or species-specific training data, but the reality is that only a limited number of species have high-quality sequences and annotations. Here we first characterize lncRNAs by contrast to protein-coding RNAs based on feature relationship and find that the feature relationship between ORF (open reading frame) length and GC content presents universally substantial divergence in lncRNAs and protein-coding RNAs, as observed in a broad variety of species. Based on the feature relationship, accordingly, we further present LGC, a novel algorithm for identifying lncRNAs that is able to accurately distinguish lncRNAs from protein-coding RNAs in a cross-species manner without any prior knowledge. As validated on large-scale empirical datasets, comparative results show that LGC outperforms existing algorithms by achieving higher accuracy, well-balanced sensitivity and specificity, and is robustly effective (>90% accuracy) in discriminating lncRNAs from protein-coding RNAs across diverse species that range from plants to mammals. To our knowledge, this study, for the first time, differentially characterizes lncRNAs and protein-coding RNAs based on feature relationship, which is further applied in computational identification of lncRNAs. Taken together, our study represents a significant advance in characterization and identification of lncRNAs and LGC thus bears broad potential utility for computational analysis of lncRNAs in a wide range of species.

bioinformatics

Genetic structure and sex-biased gene flow in the history of southern African populations

ObjectivesWe investigated the genetic history of southern African populations with a special focus on their paternal history. We reexamined previous claims that the Y-chromosome haplogroup E1b1b was brought to southern Africa by pastoralists from eastern Africa, and investigated patterns of sex-biased gene flow in southern Africa.\n\nMaterial and MethodsWe analyzed previously published complete mtDNA genome sequences and ~900 kb of NRY sequences from 23 populations from Namibia, Botswana and Zambia, as well as haplogroup frequencies from a large sample of southern African populations and 23 newly genotyped Y-linked STR loci for samples assigned to haplogroup E1b1b.\n\nResultsOur results support an eastern African origin for Y-chromosome haplogroup E1b1b; however, its current distribution in southern Africa is not strongly associated with pastoralism, suggesting a more complex origin for pastoralism in this region. We confirm that the Bantu expansion had a notable genetic impact in southern Africa, and that in this region it was probably a rapid, male-dominated expansion. Furthermore, we find a significant increase in the intensity of sex-biased gene flow from north to south, which may reflect changes in the social dynamics between Khoisan and Bantu groups over time.\n\nConclusionsOur study shows that the population history of southern Africa has been very complex, with different immigrating groups mixing to different degrees with the autochthonous populations. The Bantu expansion led to heavily sex-biased admixture as a result of interactions between Khoisan females and Bantu males, with a geographic gradient which may reflect changes in the social dynamics between Khoisan and Bantu groups over time.

genetics