bioRxiv ScienceSearch

Biology subjects

Sese, J.

Publications and source records attributed to Sese, J..

4 recordsLinked to original sources

Homeolog expression quantification methods for allopolyploids

Genome duplication with hybridization, or allopolyploidization, occurs in animals, fungi, and plants, and is especially common in crop plants. There is increasing interest in the study of allopolyploids due to advances in polyploid genome assembly, however the high level of sequence similarity in duplicated gene copies (homeologs) pose many challenges. Here we compared standard RNA-seq expression quantification approaches used currently for diploid species against subgenome-classification approaches which maps reads to each subgenome separately. We examined mapping error using our previous and new RNA-seq data in which a subgenome is experimentally added (synthetic allotetraploid Arabidopsis kamchatica) or reduced (allohexaploid wheat Triticum aestivum versus extracted allotetraploid) as ground truth. The error rates in the two species were very similar. The standard approaches showed higher error rates (> 10% using pseudo-alignment with Kallisto) while subgenome-classification approaches showed much lower error rates (< 1% using EAGLE-RC, < 2% using HomeoRoq). Although downstream analysis may partly mitigate mapping errors, the difference in methods was substantial in hexaploid wheat, where Kallisto appeared to have systematic differences relative to other methods. Only approximately half of the differentially expressed homeologs detected using Kallisto overlapped with those by any other method. In general, disagreement in low expression genes was responsible for most of the discordance between methods, which is consistent with known biases in Kallisto. We also observed that there exist uncertainties in genome sequences and annotation which can affect each method differently. Overall, subgenome-classification approaches tend to perform better than standard approaches with EAGLE-RC having the highest precision.

bioinformatics

reactIDR: Evaluation of the statistical reproducibility of high-throughput structural analyses for a robust RNA reactivity classification

MotivationRecently, next-generation sequencing techniques have been applied for the detection of RNA secondary structures called high-throughput RNA structural (HTS) analy- sis, and dozens of different protocols were used to detect comprehensive RNA structures at single-nucleotide resolution. However, the existing computational analyses heavily depend on experimental data generation methodology, which results in many difficulties associated with statistically sound comparisons or combining the results obtained using different HTS methods.\n\nResultsHere, we introduced a statistical framework, reactIDR, which is applicable to the experimental data obtained using multiple HTS methodologies, and it classifies the nucleotides into three structural categories, stem, loop, and unmapped. reactIDR uses the irreproducible discovery rate (IDR) with a hidden Markov model (HMM) to discriminate accurately between the true and spurious signals obtained in the replicated HTS experiments. In reactIDR, IDR and HMM parameters are efficiently optimized by using an expectation-maximization algorithm. Furthermore, if known reference structures are given, a supervised learning can be applicable in a semi-supervised manner. The results of our analyses for real HTS data showed that reactIDR achieved the highest accuracy in the classification problem of stem/loop structures of rRNA using both individual and integrated HTS datasets as well as the best correspondence with the three-dimensional structure. Because reactIDR is the first method to compare HTS datasets obtained from multiple sources in a single unified model, it has a great potential to increase the accuracy of RNA secondary structure prediction at transcriptome-wide level with further experiments performed.\n\nAvailabilityreactIDR is implemented in Python. Source code is publicly available at https://github.com/carushi/reactIDRhttps://github.com/carushi/reactIDR.\n\nContactkawaguchi-rs@aist.go.jp\n\nSupplementary informationSupplementary data are available at online.

bioinformatics

Integrative analysis of transcription factor occupancy at enhancers and disease risk loci in noncoding genomic regions

Noncoding regions of the human genome possess enhancer activity and harbor risk loci for heritable diseases. Whereas the binding profiles of multiple transcription factors (TFs) have been investigated, integrative analysis with the large body of public data available so as to provide an overview of the function of such noncoding regions has remained a challenge. Here we have fully integrated public ChIP-seq and DNase-seq data (n ~ 70,000), including those for 743 human transcription factors (TFs) with 97 million binding sites, and have devised a data- mining platform --designated ChIP-Atlas--to identify significant TF-genome, TF-gene, and TF-TF interactions. Using this platform, we found that TFs enriched at macrophage or T-cell enhancers also accumulated around risk loci for autoimmune diseases, whereas those enriched at hepatocyte or macrophage enhancers were preferentially detected at loci associated with HDL-cholesterol levels. Of note, we identified \"hotspots\" around such risk loci that accumulated multiple TFs and are therefore candidates for causal variants. Integrative analysis of public chromatin-profiling data is thus able to identify TFs and tissues associated with heritable disorders.

genomics

Patterns of polymorphism, selection and linkage disequilibrium in the subgenomes of the allopolyploid Arabidopsis kamchatica

Although genome duplication is widespread in wild and crop plants, little is known about genome-wide selection due to the complexity of polyploid genomes. In allopolyploid species, the patterns of purifying selection and adaptive substitutions would be affected by masking owing to duplicated genes or homeologs as well as by effective population size. We resequenced 25 distribution-wide accessions of the allotetraploid Arabidopsis kamchatica, which has a relatively small genome size (450 Mb) derived from the diploid species A. halleri and A. lyrata. The level of nucleotide polymorphism and linkage disequilibrium decay were comparable to A. thaliana, indicating the feasibility of association studies. A reduction in purifying selection compared with parental species was observed. Interestingly, the proportion of adaptive substitutions () was significantly positive in contrast to the majority of plant species. A recurrent pattern observed in both frequency and divergence-based neutrality tests is that the genome-wide distributions of both subgenomes were similar, but the correlation between homeologous pairs was low. This may increase the opportunity of different evolutionary trajectories such as in the HMA4 gene involved in heavy metal hyperaccumulation.

evolutionary biology