bioRxiv · 10.1101/2024.10.25.620288
Pre-processing annotated homologous regions in protein sequences concerning machine-learning applications
Abstract
Accurate preprocessing of annotated protein sequences with regard to homologies is essential for maintaining the integrity of machine-learning applications. This study presents two new tools--HAM (Homology-based Annotation Masking) and HAC (Homology Annotation Conflict)-- designed to address these challenges. HAM detects and masks homologous regions between datasets to prevent leakage, while HAC identifies and resolves annotation inconsistencies within datasets. Applying these tools to three benchmark datasets revealed substantial overlooked homology and annotation conflicts, even in datasets that had been previously clustered by sequence identity. These findings underscore the importance of homology-aware preprocessing to ensure the integrity of model training and evaluation. By integrating HAM and HAC into machine learning workflows, researchers can improve the consistency and trustworthiness of protein sequence-based predictions. Availabilitygithub.com/NawarMalhis/HAM.git
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Malhis, N.. 2024-10-25. Pre-processing annotated homologous regions in protein sequences concerning machine-learning applications. https://doi.org/10.1101/2024.10.25.620288
Cite the original work for its findings. Save a collection to share your selection of sources.