bioRxiv Science⌕ Search

Biology subjects

Brock, J. R.

Publications and source records attributed to Brock, J. R..

2 recordsLinked to original sources

Allopolyploidy expanded gene content but not pangenomic variation in the hexaploid oilseed Camelina sativa

Ancient whole-genome duplications (WGDs) are believed to facilitate novelty and adaptation by providing the raw fuel for new genes. However, it is unclear how recent WGDs may contribute to evolvability within recent polyploids. Hybridization accompanying some WGDs may combine divergent gene content among diploid species. Some theory and evidence suggest that polyploids have a greater accumulation and tolerance of gene presence-absence and genomic structural variation, but it is unclear to what extent either is true. To test how recent polyploidy may influence pangenomic variation, we sequenced, assembled, and annotated twelve complete, chromosome-scale genomes of Camelina sativa, an allohexaploid biofuel crop with three distinct subgenomes. Using pangenomic comparative analyses, we characterized gene presence-absence and genomic structural variation both within and between the subgenomes. We found over 75% of ortholog gene clusters are core in Camelina sativa and <10% of sequence space was affected by genomic structural rearrangements. In contrast, 19% of gene clusters were unique to one subgenome, and the majority of these were Camelina-specific (no ortholog in Arabidopsis). We identified an inversion that may contribute to vernalization requirements in winter-type Camelina, and an enrichment of Camelina-specific genes with enzymatic processes related to seed oil quality and Camelinas unique glucosinolate profile. Genes related to these traits exhibited little presence-absence variation. Our results reveal minimal pangenomic variation in this species, and instead show how hybridization accompanied by WGD may benefit polyploids by merging diverged gene content of different species.

genomics↗

Identification and Functional Annotation of Long Intergenic Non-coding RNAs in the Brassicaceae

Long intergenic noncoding RNAs (lincRNAs) are a large yet enigmatic class of eukaryotic transcripts with critical biological functions. Despite the wealth of RNA-seq data available, lincRNA identification lags in the plant lineage. In addition, there is a need for a harmonized identification and annotation effort to enable cross-species functional and genomic comparisons. In this study we processed >24 Tbp of RNA-seq data from >16,000 experiments to identify ~130,000 lincRNAs in four Brassicaceae: Arabidopsis thaliana, Camelina sativa, Brassica rapa, and Eutrema salsugineum. We used Nanopore RNA-seq, transcriptome-wide structural information, peptide data, and epigenomic data to characterize these lincRNAs and identify functional motifs. We then used comparative genomic and transcriptomic approaches to highlight lincRNAs in our dataset with sequence or transcriptional evolutionary conservation, including lincRNAs transcribed adjacent to orthologous genes that display little sequence similarity and likely function as transcriptional regulators. Finally, we used guilt-by-association techniques to further classify these lincRNAs according to putative function. LincRNAs with Brassicaceae-conserved putative miRNA binding motifs, short ORFs, and whose expression is modulated by abiotic stress are a few of the annotations that will prioritize and guide future functional analyses.

genomics↗