bioRxiv Science⌕ Search

Biology subjects

Goldman, S. L.

Publications and source records attributed to Goldman, S. L..

2 recordsLinked to original sources

Hackflex library preparation enables low-cost metagenomic profiling

Shotgun metagenomic sequencing provides valuable insights into microbial communities, but the high cost of library preparation with standard kits and protocols is a barrier for many. New methods such as Hackflex use diluted commercially available reagents to greatly reduce library preparation costs. However, these methods have not been systematically validated for metagenomic sequencing. Here, we evaluate Hackflex performance by sequencing metagenomic libraries from known mock communities as well as mouse fecal samples prepared by Hackflex, Illumina DNA Prep, and Illumina TruSeq methods. Hackflex successfully recovered all members of the Zymo mock community, performing best for samples with DNA concentrations <1 ng/uL. Furthermore, Hackflex was able to delineate microbiota of individual inbred mice from the same breeding stock at the same mouse facility, and statistical modeling indicated that mouse ID explained a greater fraction of the variance in metagenomic composition than did library preparation method. These results show that Hackflex is suitable for generating inventories of bacterial communities through metagenomic sequencing.

microbiology↗

FLIP: Benchmark tasks in fitness landscape inference for proteins

Machine learning could enable an unprecedented level of control in protein engineering for therapeutic and industrial applications. Critical to its use in designing proteins with desired properties, machine learning models must capture the protein sequence-function relationship, often termed fitness landscape. Existing bench-marks like CASP or CAFA assess structure and function predictions of proteins, respectively, yet they do not target metrics relevant for protein engineering. In this work, we introduce Fitness Landscape Inference for Proteins (FLIP), a benchmark for function prediction to encourage rapid scoring of representation learning for protein engineering. Our curated tasks, baselines, and metrics probe model generalization in settings relevant for protein engineering, e.g. low-resource and extrapolative. Currently, FLIP encompasses experimental data across adeno-associated virus stability for gene therapy, protein domain B1 stability and immunoglobulin binding, and thermostability from multiple protein families. In order to enable ease of use and future expansion to new tasks, all data are presented in a standard format. FLIP scripts and data are freely accessible at https://benchmark.protein.properties.

bioengineering↗