bioRxiv Science⌕ Search

Biology subjects

Farh, K.

Publications and source records attributed to Farh, K..

5 recordsLinked to original sources

Phylogenetic signal in primate tooth enamel proteins and its relevance for paleoproteomics

Ancient tooth enamel, and to some extent dentin and bone, contain characteristic peptides that persist for long periods of time. In particular, peptides from the enamel proteome (enamelome) have been used to reconstruct the phylogenetic relationships of fossil specimens and to estimate divergence times. However, the enamelome is based on only about 10 genes, whose protein products undergo fragmentation post mortem. Moreover, some of the enamelome genes are paralogous or may coevolve. This raises the question as to whether the enamelome provides enough information for reliable phylogenetic inference. We address these considerations on a selection of enamel-associated proteins that has been computationally predicted from genomic data from 232 primate species. We created multiple sequence alignments (MSAs) for each protein and estimated the evolutionary rate for each site and examined which sites overlap with the parts of the protein sequences that are typically isolated from fossils. Based on this, we simulated ancient data with different degrees of sequence fragmentation, followed by phylogenetic analysis. We compared these trees to a reference species tree. Up to a degree of fragmentation that is similar to that of fossil samples from 1-2 million years ago, the phylogenetic placements of most nodes at family level are consistent with the reference species tree. We found that the composition of the proteome influences the phylogenetic placement of Tarsiiformes. For the inference of molecular phylogenies based on paleoproteomic data, we recommend characterizing the evolution of the proteomes from the closest extant relatives to maximize the reliability of phylogenetic inference.

evolutionary biology↗

Single cell sequencing as a general variant interpretation assay

The human genome contains [~]70 million possible protein-altering variants, the vast majority of which are of uncertain clinical significance. Closing this gap is essential for accurate diagnosis of disease-causing variants and understanding their mechanisms of action. Towards this goal, we developed a pooled perturbation approach combining saturation mutagenesis with single cell RNA sequencing to map the effects of every single nucleotide variant in a gene. We sequenced [~]440,000 cells expressing variants in CDKN2A (p16INK4a), TP53, and SOD1, observing almost all possible protein-coding variants, with a mean of 61 cells per variant. Using single cell gene expression signatures, we show that each gene may contain multiple types of pathogenic variants that affect distinct downstream pathways. We demonstrate that single cell expression signatures outperform existing bulk experimental assays and computational models for predicting pathogenicity, and summarize both the utility and potential limitations of single cell sequencing as a general variant interpretation assay.

genomics↗

Large-scale phylogenomics uncovers a complex evolutionary history and extensive ancestral gene flow in an African primate radiation

Understanding the drivers of speciation is fundamental in evolutionary biology, and recent studies highlight hybridization as a potential facilitator of adaptive radiations. Using whole-genome sequencing data from 22 species of guenons (tribe Cercopithecini), one of the worlds largest primate radiations, we show that rampant gene flow characterizes their evolutionary history, and identify ancient hybridization across deeply divergent lineages differing in ecology, morphology and karyotypes. Lineages experiencing gene flow tend to be more species-rich than non-admixed lineages. Mitochondrial transfers between distant lineages were likely facilitated by co-introgression of co-adapted nuclear variants. Although the genomic landscapes of introgression were largely lineage specific, we found that genes with immune functions were overrepresented in introgressing regions, in line with adaptive introgression, whereas genes involved in pigmentation and morphology may contribute to reproductive isolation. This study provides important insights into the prevalence, role and outcomes of ancestral hybridization in a large mammalian radiation.

evolutionary biology↗

Integrative single-cell analysis of cardiogenesis identifies developmental trajectories and non-coding mutations in congenital heart disease

Congenital heart defects, the most common birth disorders, are the clinical manifestation of anomalies in fetal heart development - a complex process involving dynamic spatiotemporal coordination among various precursor cell lineages. This complexity underlies the incomplete understanding of the genetic architecture of congenital heart diseases (CHDs). To define the multi-cellular epigenomic and transcriptional landscape of cardiac cellular development, we generated single-cell chromatin accessibility maps of human fetal heart tissues. We identified eight major differentiation trajectories involving primary cardiac cell types, each associated with dynamic transcription factor (TF) activity signatures. We identified similarities and differences of regulatory landscapes of iPSC-derived cardiac cell types and their in vivo counterparts. We interpreted deep learning models that predict cell-type resolved, base-resolution chromatin accessibility profiles from DNA sequence to decipher underlying TF motif lexicons and infer the regulatory impact of non-coding variants. De novo mutations predicted to affect chromatin accessibility in arterial endothelium were enriched in CHD cases versus controls. We used CRISPR-based perturbations to validate an enhancer harboring a nominated regulatory CHD mutation, linking it to effects on the expression of a known CHD gene JARID2. Together, this work defines the cell-type resolved cis-regulatory sequence determinants of heart development and identifies disruption of cell type-specific regulatory elements as a component of the genetic etiology of CHD.

genomics↗

Chromatin and gene-regulatory dynamics of the developing human cerebral cortex at single-cell resolution

Genetic perturbations of cerebral cortical development can lead to neurodevelopmental disease, including autism spectrum disorder (ASD). To identify genomic regions crucial to corticogenesis, we mapped the activity of gene-regulatory elements generating a single-cell atlas of gene expression and chromatin accessibility both independently and jointly. This revealed waves of gene regulation by key transcription factors (TFs) across a nearly continuous differentiation trajectory into glutamatergic neurons, distinguished the expression programs of glial lineages, and identified lineage-determining TFs that exhibited strong correlation between linked gene-regulatory elements and expression levels. These highly connected genes adopted an active chromatin state in early differentiating cells, consistent with lineage commitment. Basepair-resolution neural network models identified strong cell-type specific enrichment of noncoding mutations predicted to be disruptive in a cohort of ASD subjects and identified frequently disrupted TF binding sites. This approach illustrates how cell-type specific mapping can provide insights into the programs governing human development and disease.

neuroscience↗