bioRxiv Science⌕ Search

Biology subjects

Honaas, L.

Publications and source records attributed to Honaas, L..

8 recordsLinked to original sources

Development of Rosaceae Crop-Specific Nanopore Models for Community Use Through the Genome Database for Rosaceae

Oxford Nanopore Technologies (ONT) sequencing platforms have enabled scientists to generate sequence reads of 20 kilobase or more, which has facilitated the assembly of complete genomes. Nanopore sequencing involves recording the change in electrical signals as the DNA molecule transverses the pore. Converting this signal information into bases - a process called basecalling - is challenging and require the implementation of machine learning. The models developed for ONT basecallers have continually improved. However, the mean quality of the reads generated fall below a Phred score of 20 (99% accuracy), the target quality for high-quality genome assembly, when used on challenging plant samples. These low-quality base calls limit the ability to conduct de novo genome assembly, accurate phasing of haplotypes, and long-read genotyping. To overcome these shortcomings, we fine-tuned ONTs high accuracy basecalling model using Bonito, an AI-based deep learning basecaller, to develop crop-specific models for five important species in the Rosaceae family. These species include highly valuable tree crops such as apple, peach, pear, plum, and sweet cherry. Through the development of an automated training pipeline, we were able to achieve [~]14% higher mean and median basecall quality, and a [~]20% increase in total reads with Phred scores >15. These results were achieved without significantly affecting the read length N50s and total reads called. As a result, these new crop-specific models will enable members of the Rosaceae genomics community to produce higher quality ONT sequencing data. Furthermore, we are releasing our automated pipeline to facilitate others to train their own organism or crop-specific models.

genomics↗

Predicting plant traits using large conglomerate RNA-seq datasets

RNA-seq datasets offer potential for predicting phenotypic traits, but the optimal dataset size for reliable predictions remains unclear. To explore this question, we compiled large-scale transcriptomic datasets across 12 plant species and used them to predict the phenotypic variables of tissue type and age. Predictions were made using random forest models. We created these models using an increasing number of samples to create performance curves. These curves show that only a few hundred samples are required to achieve maximum accuracy for tissue classification, while predicting age demands a few thousand samples. Acceptable prediction accuracy can be achieved at even lower numbers of samples. Our findings provide a benchmark for designing transcriptomic studies aimed at phenotype prediction and highlight the differing complexities involved in predicting simple versus more complex traits. CORE IDEAS- Plant transcriptomes are highly dynamic in response to internal and external conditions - There is interest in using transcriptomes for predicting phenotypic traits and disorders - It is unknown how many samples are required for accurate phenotype prediction - We created 12 massive RNA-seq datasets and corresponding phenotypic data and used these to create prediction models - Performance curves from our models provide insight into the required number of samples for accurate prediction

systems biology↗

The evolution of pectate lyase-like genes across land plants, and their roles in haustorium formation in parasitic plants

Parasitic plants in Orobanchaceae are noxious agricultural pests that severely impact crops worldwide. These plants acquire water and nutrients from their hosts through a specialized organ called the haustorium. A key step in haustorium development involves cell wall modification. In this study, we identified and analyzed the evolutionary relationships of pectate lyase-like (PLL) genes across parasitic plants and other non-parasitic land plant lineages. To support detailed examination of gene models and paralogous gene family members, we used published parasitic plant genomes, as well as a recently generated draft genome assembly and annotation of Triphysaria versicolor. One particular PLL gene, denoted as PLL1 in parasitic Orobanchaceae, emerged as an important candidate gene for parasitism. Our previous comparative transcriptomic analyses showed that PLL1 underwent neofunctionalization via an expression shift from floral tissues in non-parasitic relatives to haustoria in parasitic species. It belongs to the largest sub-clade of the PLL gene family, is highly upregulated in haustoria, and shows signatures of relaxed purifying selection and 15 individual sites with signatures of adaptive evolution. To explore its function in haustorium development, we manipulated PLL1 expression in T. versicolor, a model parasitic species from Orobanchaceae, using direct transformation with the parasite and host-induced-gene-silencing (HIGS). For HIGS, we generated transgenic Medicago hosts expressing hairpin RNAs targeting the PLL1 gene in T. versicolor. An average 60% reduction of PLL1 transcript level was observed in both direct transformation and HIGS treatments, leading to an increased frequency of poorly adhered parasites with fewer xylem connections and a smaller proportion of mature haustoria. These findings demonstrate that PLL1 plays a crucial role in haustorium development and suggest it as a promising target for managing parasitic weeds. Notably, the success of HIGS even before the establishment of a functional haustorium highlights the possibility of early intervention against parasitism.

plant biology↗

Rating Pome Fruit Quality Traits Using Deep Learning and Image Processing

Quality assessment of pome fruits (i.e. apples and pears) is used not only crucial for determining the optimal harvest time, but also the progression of fruit-quality attributes during storage. Therefore, it is typical to repeatedly evaluate fruits during the course of a postharvest experiment. This evaluation often includes careful visual assessments of fruit for apparent defects and physiological symptoms. A general best practice for quality assessment is to rate fruit using the same individual rater or group of individuals raters to reduce bias. However, such consistency across labs, facilities, and experiments is often not feasible or attainable. Moreover, while these visual assessments are critical empirical data, they are often coarse-grained and lack consistent objective criteria. Granny, is a tool designed for rating fruit using machine-learning and image-processing to address rater bias and improve resolution. Additionally, Granny supports backwards compatibility by providing ratings compatible with long-established standards and references, promoting research program continuity. Current Granny ratings include starch content assessment, rating levels of peel defects, and peel color analyses. Integrative analyses enhanced by Grannys improved resolution and reduced bias, such as linking fruit outcomes to global scale-omics data, environmental changes, and other quantitative fruit quality metrics like soluble solids content and flesh firmness, will further enrich our understanding of fruit quality dynamics. Lastly, Granny is open-source and freely available.

plant biology↗

A Phased, Chromosome-scale Genome for Malus domestica 'WA 38'

Genome sequencing for agriculturally important Rosaceous crops has made rapid progress both in completeness and annotation quality. Whole genome sequence and annotation gives breeders, researchers, and growers information about cultivar specific traits such as fruit quality, disease resistance, and informs strategies to enhance postharvest storage. Here we present a haplotype-phased, chromosomal level genome of Malus domestica, WA 38, a new apple cultivar released to market in 2017 as Cosmic Crisp (R). Using both short and long read sequencing data with a k-mer based approach, chromosomes originating from each parent were assembled and segregated. This is the first pome fruit genome fully phased into parental haplotypes in which chromosomes from each parent are identified and separated into their unique, respective haplomes. The two haplome assemblies, Honeycrisp originated HapA and Enterprise originated HapB, are about 650 Megabases each, and both have a BUSCO score of 98.7% complete. A total of 53,028 and 54,235 genes were annotated from HapA and HapB, respectively. Additionally, we provide genome-scale comparisons to Gala, Honeycrisp, and other relevant cultivars highlighting major differences in genome structure and gene family circumscription. This assembly and annotation was done in collaboration with the American Campus Tree Genomes project that includes WA 38 (Washington State University), dAnjou pear (Auburn University), and many more. To ensure transparency, reproducibility, and applicability for any genome project, our genome assembly and annotation workflow is recorded in detail and shared under a public GitLab repository. All software is containerized, offering a simple implementation of the workflow.

genomics↗

A chromosome-scale assembly for dAnjou pear

Cultivated pear consists of several Pyrus species with P. communis (European pear) representing a large fraction of worldwide production. As a relatively recently domesticated crop and perennial tree, pear can benefit from genome-assisted breeding. Additionally, comparative genomics within Rosaceae promises greater understanding of evolution within this economically important family. Here, we generate a fully-phased chromosome-scale genome assembly of P. communis cv. dAnjou. Using PacBio HiFi and Dovetail Omni-C reads, the genome is resolved into the expected 17 chromosomes, with each haplotype totalling nearly 540 Megabases and a contig N50 of nearly 14 Mb. Both haplotypes are highly syntenic to each other, and to the Malus domestica Honeycrisp apple genome. Nearly 45,000 genes were annotated in each haplotype, over 90% of which have direct RNA-seq expression evidence. We detect signatures of the known whole-genome duplication shared between apple and pear, and we estimate 57% of dAnjou genes are retained in duplicate derived from this event. This genome highlights the value of generating phased diploid assemblies for recovering the full allelic complement in highly heterozygous crop species.

genomics↗

PlantTribes2: tools for comparative gene family analysis in plant genomics

Plant genome-scale resources are being generated at an increasing rate as sequencing technologies continue to improve and raw data costs continue to fall; however, the cost of downstream analyses remains large. This has resulted in a considerable range of genome assembly and annotation qualities across plant genomes due to their varying sizes, complexity, and the technology used for the assembly and annotation. To effectively work across genomes, researchers increasingly rely on comparative genomic approaches that integrate across plant community resources and data types. Such efforts have aided the genome annotation process and yielded novel insights into the evolutionary history of genomes and gene families, including complex non-model organisms. The essential tools to achieve these insights rely on gene family analysis at a genome-scale, but they are not well integrated for rapid analysis of new data, and the learning curve can be steep. Here we present PlantTribes2, a scalable, easily accessible, highly customizable, and broadly applicable gene family analysis framework with multiple entry points including user provided data. It uses objective classifications of annotated protein sequences from existing, high-quality plant genomes for comparative and evolutionary studies. PlantTribes2 can improve transcript models and then sort them, either genome-scale annotations or individual gene coding sequences, into pre-computed orthologous gene family clusters with rich functional annotation information. Then, for gene families of interest, PlantTribes2 performs downstream analyses and customizable visualizations including, (1) multiple sequence alignment, (2) gene family phylogeny, (3) estimation of synonymous and non-synonymous substitution rates among homologous sequences, and (4) inference of large-scale duplication events. We give examples of PlantTribes2 applications in functional genomic studies of economically important plant families, namely transcriptomics in the weedy Orobanchaceae and a core orthogroup analysis (CROG) in Rosaceae. PlantTribes2 is freely available for use within the main public Galaxy instance and can be downloaded from GitHub or Bioconda. Importantly, PlantTribes2 can be readily adapted for use with genomic and transcriptomic data from any kind of organism.

bioinformatics↗

A phased, chromosome-scale genome of 'Honeycrisp' apple (Malus domestica)

Honeycrisp is one of the most valuable apple cultivars grown in the United States and a popular breeding parent due to its superior fruit quality traits, high levels of cold hardiness, and disease resistance. However, it suffers from a number of physiological disorders and is susceptible to production and postharvest issues. Although several apple genomes have been sequenced in the last decade, there is still a substantial knowledge gap in understanding the genetic mechanisms underlying cultivar-specific traits. Here we present a fully phased, chromosome-level genome of Honeycrisp apples, using PacBio HiFi, Omni-C, and Illumina sequencing platforms. Our genome assembly is by far the most contiguous among all the apple genomes. The sizes of the two assembled haplomes are 674 Mb and 660 Mb, with contig N50s of 32.8 Mb and 31.6 Mb, respectively. In total, 47,563 and 48,655 protein coding genes were annotated from each haplome, capturing 96.8-97.4% complete BUSCOs in the eudicot database, the most complete among all Malus annotations. A gene family analysis using seven Malus genomes shows that a vast majority of Honeycrisp genes are assigned into orthogroups shared with other genomes, but it also reveals 121 Honeycrisp-specific orthogroups. We provide a valuable resource for understanding the genetic basis of horticulturally important traits in apples and other related tree fruit species, including at-harvest and postharvest fruit quality, abiotic stress tolerance, and disease resistance, all of which can enhance breeding efforts in Rosaceae.

genomics↗