bioRxiv · 10.1101/2021.06.09.447735
Optimized Splitting of RNA Sequencing Data by Species
Abstract
Gene expression studies using chimeric xenograft transplants or co-culture systems have proven to be valuable to uncover cellular dynamics and interactions during development or in disease models. However, the mRNA sequence similarities among species presents a challenge for accurate transcript quantification. To identify optimal strategies for analyzing mixed-species RNA sequencing data, we evaluate both alignment-dependent and alignment-independent methods. Alignment of reads to a pooled reference index is effective, particularly if optimal alignments are used to classify sequencing reads by species, which are re-aligned with individual genomes, generating >97% accuracy across a range of species ratios. Alignment-independent methods, such as Convolutional Neural Networks, which extract the conserved patterns of sequences from two species, classify RNA sequencing reads with over 85% accuracy. Importantly, both methods perform well with different ratios of human and mouse reads. Our evaluation identifies valuable and effective strategies to dissect species composition of RNA sequencing data from mixed populations.
Source connections
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Song, X., Gao, H. Y., Herrup, K., Hart, R. P.. 2021-06-10. Optimized Splitting of RNA Sequencing Data by Species. https://doi.org/10.1101/2021.06.09.447735
Cite the original work for its findings. Save a collection to share your selection of sources.