Tri-Modality Representation Learning for Molecular Property Prediction
Accurate molecular property prediction requires effective molecular representations that can describe a molecule from multiple complementary perspectives. Existing deep learning approaches typically use SMILES strings, two-dimensional molecular graphs, or three-dimensional conformations as inputs. These representations capture different aspects of molecular information: SMILES encodes a sequential description, molecular graphs describe atom-bond connectivity, and 3D conformations provide spatial information from atomic coordinates. However, these modalities are often learned separately or merged with low effective fusion operations, which may not sufficiently capture interactions among sequential, topological, and geometric features. In this work, we propose Tri-Modality Cross-Attention (TMCA), a multimodal framework that integrates 1D SMILES, 2D molecular graphs, and 3D molecular conformations for molecular property prediction. TMCA uses pretrained encoders for the 1D and 3D branches, with a SMILES-based Transformer for the 1D branch and a recent conformation-aware pretrained model for the 3D branch. At the same time, a trainable 2D graph encoder is designed to support modality fusion. All the parameters from the pretrained models are frozen to make overall training more efficient. We evaluate TMCA on four datasets (i.e., BBBP, BACE, ClinTox, and HIV) from the MoleculeNet database for classification tasks and compare its performance with state-of-the-art competing approaches. Experimental results show that TMCA achieves the best overall performance, demonstrating the value of multi-modality integration for drug property prediction and the effectiveness of our cross-attention-based fusion strategy. The source code of TMCA is available at: https://github.com/Ay-Zhao/TMCA