bioRxiv Science⌕ Search

Biology subjects

Baghbanzadeh, M.

Publications and source records attributed to Baghbanzadeh, M..

2 recordsLinked to original sources

seqLens: optimizing language models for genomic predictions

Understanding evolutionary variation in genomic sequences through the lens of language modeling has the potential to revolutionize biological research. Yet to maximize the utility of language modeling in genomics, we must overcome computational challenges in tokenization and model architecture adapted to diverse genomic features across evolutionary timescales. In this study, we investigated key elements in genomic language modeling (gLM), including tokenization, pretraining datasets, fine-tuning approaches, pooling methods, and domain adaptation, and applied the language models to diverse genomic data. We gathered two evolutionarily distinct pretraining datasets: one consisting of 19,551 reference genomes, including over 18,000 prokaryotic genomes (115B nucleotides) and the remainder eukaryotic genomes, and another more balanced dataset with 1,354 genomes, including 1,166 prokaryotic and 188 eukaryotic reference genomes (180B nucleotides). We trained five byte-pair encoding tokenizers and pretrained 52 gLMs, systematically comparing different architectures, hyperparameters, and classification heads. We introduce seqLens, a family of models based on disentangled attention with relative positional encoding, which outperforms relatively similar-sized models in 13 of 19 benchmarking phenotypic predictions. We further explore continual pretraining, domain adaptation, and parameter-efficient fine-tuning methods to assess trade-offs between computational efficiency and accuracy. Our findings demonstrate that relevant pretraining data significantly boost performance, alternative pooling techniques can enhance classification, tokenizers with larger vocabulary sizes negatively impact generalization, and gLMs are capable of understanding evolutionary relationships. These insights provide a foundation for optimizing genomic language models for identifying diverse evolutionary genomic features and improving genome annotations.

bioinformatics↗

Discovering genotype-phenotype relationships with machine learning and the Visual Physiology Opsin Database (VPOD)

BackgroundPredicting phenotypes from genetic variation is foundational for fields as diverse as bioengineering and global change biology, highlighting the importance of efficient methods to predict gene functions. Linking genetic changes to phenotypic changes has been a goal of decades of experimental work, especially for some model gene families including light-sensitive opsin proteins. Opsins can be expressed in vitro to measure light absorption parameters, including {lambda}max - the wavelength of maximum absorbance - which strongly affects organismal phenotypes like color vision. Despite extensive research on opsins, the data remain dispersed, uncompiled, and often challenging to access, thereby precluding systematic and comprehensive analyses of the intricate relationships between genotype and phenotype. ResultsHere, we report a newly compiled database of all heterologously expressed opsin genes with {lambda}max phenotypes called the Visual Physiology Opsin Database (VPOD). VPOD_1.0 contains 864 unique opsin genotypes and corresponding {lambda}max phenotypes collected across all animals from 73 separate publications. We use VPOD data and deepBreaks to show regression-based machine learning (ML) models often reliably predict {lambda}max, account for non-additive effects of mutations on function, and identify functionally critical amino acid sites. ConclusionThe ability to reliably predict functions from gene sequences alone using ML will allow robust exploration of molecular-evolutionary patterns governing phenotype, will inform functional and evolutionary connections to an organisms ecological niche, and may be used more broadly for de-novo protein design. Together, our database, phenotype predictions, and model comparisons lay the groundwork for future research applicable to families of genes with quantifiable and comparable phenotypes. Key PointsO_LIWe introduce the Visual Physiology Opsin Database (VPOD_1.0), which includes 864 unique animal opsin genotypes and corresponding {lambda}max phenotypes from 73 separate publications. C_LIO_LIWe demonstrate that regression-based ML models can reliably predict {lambda}max from gene sequence alone, predict non-additive effects of mutations on function, and identify functionally critical amino acid sites. C_LIO_LIWe provide an approach that lays the groundwork for future robust exploration of molecular-evolutionary patterns governing phenotype, with potential broader applications to any family of genes with quantifiable and comparable phenotypes. C_LI

bioinformatics↗