bioRxiv ScienceSearch

Biology subjects

Saif, R.

Publications and source records attributed to Saif, R..

3 recordsLinked to original sources

Differential Gene Expression Pipeline for Whole Transcriptome RNA-Seq Data using Personal Computer

Advances in the next generation sequencing (NGS) technologies, their cost effectiveness and well-developed pipelines using computational tools/softwares has allowed researchers to reveal ground-breaking discoveries in multi-omics data analysis. However, there is still uncertainty due to massive upsurge in parallel tools and difficulty in choosing best practiced pipeline for expression profiling of RNA sequenced (RNA-seq) data. Here, we detail the optimized pipeline that works at a fast pace with enhanced accuracy on personal computer rather than using cloud or high-performance computing clusters (HPC). The steps include quality check, base filtration, quasi-mapping, quantification of samples, estimation and counting of transcript/gene expression abundances, identification and clustering of differentially expressed features and visualization of the data. The tools FastQC, Trimmomatic, Salmon and some other scripts in Trinity toolkit were applied on two paired-end datasets. An extension of this pipeline may also be formulated in future for the gene ontology enrichment analysis and functional annotation of the differential expression matrix to make this data biologically more significant.

bioinformatics

Whole Genome Selective Sweeps Analysis in Pakistani Kamori Goat

Natural and artificial selection fix certain genomic regions of reduce heterozygosity which is an initial process in breed development. Primary goal of the current study is to identify these genomic selection signatures under positive selection and harbor genes in Pakistani Kamori goat breed. High throughput whole genome pooled-seq of Kamori (n = 12) and Bezoar (n = 8) was carried out. Raw fastq files were undergone quality checks, trimming and mapping process against ARS1 reference followed by calling variant allele frequencies. Selection sweeps were identified by applying pooled heterozygosity (Hp) and Tajimas D (TD) on Kamori while regions under divergent selection between Kamori & Bezoar were observed by Fixation Index (FST) analysis. Genome sequencing yielded 619,031,812 reads of which, 616,624,284 were successfully mapped. Total 98,574 autosomal selection signals were detected; 32,838 from Hp and 32,868 from each FST & TD statistics. Annotation of the regions with threshold (-ZHp [≥] 5, TD [≤] -2.72 & FST [≤] 0.09) detected 60 candidate genes. The top hits harbor Chr.1, 6, 8 & 21 having genes associated with body weight (GLIS3, ASTE1), coat color (DOCK8, MIPOL1) & body height (SLC25A21). Other significant windows harbor milk production, wool production, immunity, adaptation and reproduction trait related genes. Current finding highlighted the under-selection genomic regions of Kamori breed and likely to be associated with its vested traits and further useful in breed improvement, and may be also propagated to other undefined goat breeds by adopting targeted breeding policies to improve the genetic potential of this valued species.

genomics

Whole Genome Comparison of Pakistani Corona Virus with Chinese and US Strains along with its Predictive Severity of COVID-19

Recently submitted 784 SARS-nCoV2 whole genome sequences from NCBI Virus database were taken for constructing phylogenetic tree to look into their similarities. Pakistani strain MT240479 (Gilgit1-Pak) was found in close proximity to MT184913 (CruiseA-USA), while the second Pakistani strain MT262993 (Manga-Pak) was neighboring to MT039887 (WI-USA) strain in the constructed cladogram in this article. Afterward, four whole genome SARS-nCoV2 strain sequences were taken for variant calling analysis, those who appeared nearest relative in the earlier cladogram constructed a week time ago. Among those two Pakistani strains each of 29,836 bases were compared against MT263429 from (WI-USA) of 29,889 bases and MT259229 (Wuhan-China) of 29,864 bases. We identified 31 variants in both Pakistani strains, (Manga-Pak vs USA=2del+7SNPs, Manga-Pak vs Chinese=2del+2SNPs, Gilgit1-Pak vs USA=10SNPs, Gilgit1-Pak vs Chinese=8SNPs), which caused alteration in ORF1ab, ORF1a and N genes with having functions of viral replication and translation, host innate immunity and viral capsid formation respectively. These novel variants are assumed to be liable for low mortality rate in Pakistan with 385 as compared to USA with 63,871 and China with 4,633 deaths by May 01, 2020. However functional effects of these variants need further confirmatory studies. Moreover, mutated N & ORF1a proteins in Pakistani strains were also analyzed by 3D structure modelling, which give another dimension of comparing these alterations at amino acid level. In a nutshell, these novel variants are assumed to be linked with reduced mortality of COVID-19 in Pakistan along with other influencing factors, these novel variants would also be useful to understand the virulence of this virus and to develop indigenous vaccines and therapeutics.

genomics