bioRxiv Science⌕ Search

Biology subjects

Resende, M. F. R.

Publications and source records attributed to Resende, M. F. R..

3 recordsLinked to original sources

Single-kernel near-infrared spectroscopy enables haploid kernel sorting in field and sweet corn using high-oil haploid inducers across diverse donor-inducer combinations

Key messageA single-kernel near-infrared reflectance spectroscopy-based sorter can effectively identify haploid kernels for doubled haploid production in field and sweet corn backgrounds. Doubled haploid (DH) technology significantly shortens the breeding cycle for developing homozygous inbred lines in maize (Zea mays). Manual sorting of haploids from a larger bulk of hybrid kernels in an induction cross is a major bottleneck in DH development. Automated systems based on near-infrared (NIR) reflectance spectroscopy can be valuable tools for rapid haploid sorting, provided that sorting accuracy is sufficient for incorporation into the DH process. In this study, we evaluated the accuracy of a custom-built single-kernel NIR (skNIR) sorter for classifying haploid kernels from 12 high-oil haploid induction populations generated from two sweet corn and two field corn donors and four high-oil haploid inducers (HOHIs). We evaluated several general classification models that can be applied without population-specific recalibration or prior genotyping, including models that classified haploids based solely on predicted oil content, as well as multivariate methods that used all wavelengths of the NIR spectra. The highest classification accuracy was obtained using a general multivariate support vector machine (SVM) model. When combined with the two best-performing HOHIs, the general SVM model accurately sorted induction populations from two of the three donor backgrounds crossed with these inducers. Two oil-based methods showed less accurate classification than the multivariate SVM model, due to overlapping oil content distributions across the two kernel classes. Overall, this study demonstrates effective skNIR-based sorting of haploid kernels from diverse induction populations using a single general model. The practical deployment of this instrument in maize breeding programs is discussed.

plant biology↗

InteracTor: Feature Engineering and Explainable AI for Profiling Protein Structure-Interaction-Function Relationships

Characterizing protein families structural and functional diversity is essential for understanding their biological roles. Traditional analyses often focus on primary and secondary structures, which may not fully capture complex protein interactions. Here we introduce InteracTor, a novel toolkit that extracts multimodal features from protein three-dimensional (3D) structures, including interatomic interactions like hydrogen bonds, van der Waals forces, and hydrophobic contacts. By integrating Explainable AI (XAI) techniques, we quantified the importance of the extracted features in the classification of protein structural and functional families. InteracTors interpretable features enable mechanistic insights into the determinants of protein structure, function, and dynamics, offering a transparent means to assess their predictive power within machine learning models. Interatomic interaction features extracted by InteracTor demonstrated superior predictive power for protein family classification compared to features based solely on primary or secondary structure, revealing the importance of considering specific tertiary contacts in computational protein analysis. This work provides a robust framework for future studies aiming to enhance the capabilities of models for protein function prediction and drug discovery. AUTHOR SUMMARYInteracTor is a computational toolkit designed to enhance our understanding of protein structure and function by focusing on three-dimensional (3D) structural interactions. Unlike traditional approaches that primarily rely on sequence or secondary structure data, InteracTor extracts biologically meaningful features such as hydrogen bonds, van der Waals forces, and hydrophobic contacts, which are critical for protein stability and dynamics. By integrating these features into machine learning models alongside explainable AI methods, InteracTor provides interpretable insights into how specific structural interactions influence protein behavior. Our results demonstrate that tertiary structure features significantly improve the accuracy of protein family classification compared to sequence-based methods alone, underscoring the importance of considering 3D interactions in computational protein analyses. The toolkits modular design makes it adaptable for diverse applications, including drug discovery and protein engineering. In a broader context, InteracTor bridges the gap between computational biology and practical applications in medicine and biotechnology by offering a transparent and robust framework for analyzing proteins at a molecular level. This work represents a step forward in leveraging structural data to advance predictive modeling and biological discovery.

bioinformatics↗

InteracTor: A new integrative feature extraction toolkit for improved characterization of protein structural properties

Understanding the structural and functional diversity of protein families is crucial for elucidating their biological roles. Traditional analyses often focus on primary and secondary structures, which include amino acid sequences and local folding patterns like alpha helices and beta sheets. However, primary and secondary structures alone may not fully represent the complex interactions within proteins. To address this limitation, we developed a new algorithm (InteracTor) to analyze proteins by extracting features from their three-dimensional (3D) structures. The toolkit extracts interatomic interaction features such as hydrogen bonds, van der Waals interactions, and hydrophobic contacts, which are crucial for understanding protein dynamics, structure, and function. Incorporating 3D structural data and interatomic interaction features provides a more comprehensive understanding of protein structure and function, potentially enhancing downstream predictive modeling capabilities. By using the extracted features in Mutual Information scoring (MI), Principal Component Analysis (PCA), t-distributed Stochastic Neighbor Embedding (t-SNE), Uniform Manifold Approximation and Projection (UMAP), and hierarchical clustering analysis as use cases, we identified clear separations among protein structural families, highlighting distinct functional aspects. Our analysis revealed that interatomic interaction features were more informative than protein secondary structure features, providing insights into potential structural and functional properties. These findings underscore the significance of considering tertiary structure in protein analysis, offering a robust framework for future studies aiming at enhancing the capabilities of models for protein function prediction and drug discovery.

bioinformatics↗