Document
1. Datasets
To develop and evaluate the cat-Fusion model, we adopted a benchmark dataset established in a previous study [1]. The positive set was assembled from several curated LLPS-related databases and published protein collections, including PRALINE [2], CD-CODE [3], DrLLPS [4], PhaSepDB [5], PhaSePro [6] and LLPSDB [7], together with LLPS-associated proteins reported in earlier prediction studies [8-10]. As these resources contain substantial overlap, duplicate entries were first merged, resulting in 5,656 unique LLPS-associated proteins. Sequence redundancy was then reduced using CD-HIT at a 50% sequence identity threshold, yielding 4,807 non-redundant positive proteins.
The negative set was derived from the human proteome. To reduce the likelihood of including unannotated or potentially LLPS-associated proteins as negative samples, all positive proteins were removed, together with their first-order interaction partners recorded in BioGRID (v3.5.175)[11]. The remaining proteins were subjected to the same CD-HIT filtering procedure, after which negative proteins were randomly sampled to maintain a class distribution comparable to that of the positive set.
The resulting dataset was divided into a training set and an independent test set in the original study. The training set contained 3,333 LLPS-associated proteins and 3,252 non-LLPS proteins, whereas the independent test set comprised 1,422 positive and 1,376 negative proteins. Detailed dataset statistics are provided in Table 1. We retained the original data partitions to ensure consistency with previous studies and facilitate fair comparisons with existing LLPS predictors. The training set was used for model development and hyperparameter optimization, while the independent test set was reserved exclusively for evaluating the predictive performance of cat-Fusion.
Training Dataset: 6585_training.fasta
Independent Test Datasets: 2798_test.fasta
2. Feature Extraction Tools
In cat-Fusion, we employed two types of residue feature extraction methods:
(1)physicochemical and structural features extraction: including secondary structure composition (α-helix, β-sheet, turn, coil, bend), intrinsic disorder, aggregation propensity, hydrophobicity, net charge, radius of gyration, residue contact number, accessible surface area, and AlphaFold predicted confidence (pLDDT), totaling 128 dimensions.
(2)semantic features extracted from PPLMs: ESM2[12].
3. References
[1] Monti, M., et al. catGRANULE 2.0: accurate predictions of liquid-liquid phase separating proteins at single amino acid resolution. Genome biology 2025;26(1):33. [2] Vandelli, A., et al. The PRALINE database: protein and Rna humAn singLe nucleotIde variaNts in condEnsates. Bioinformatics 2023;39(1). [3] Rostam, N., et al. CD-CODE: crowdsourcing condensate database and encyclopedia. Nature Methods 2023;20(5):673–676. [4] Ning, W., et al. DrLLPS: a data resource of liquid-liquid phase separation in eukaryotes. Nucleic Acids Res 2020;48(D1):D288–d295. [5] You, K., et al. PhaSepDB: a database of liquid–liquid phase separation related proteins. Nucleic acids research 2020;48(D1):D354–D359. [6] Mészáros, B., et al. PhaSePro: the database of proteins driving liquid–liquid phase separation. Nucleic Acids Research 2020;48(D1):D360–D367. [7] Wang, X., et al. LLPSDB v2.0: an updated database of proteins undergoing liquid–liquid phase separation in vitro. Bioinformatics 2022;38(7):2010–2014. [8] Kuechler, E.R., et al. Distinct features of stress granule proteins predict localization in membraneless organelles. Journal of molecular biology 2020;432(7):2349–2368. [9] Kuechler, E.R., et al. Comparison of Biomolecular Condensate Localization and Protein Phase Separation Predictors. Biomolecules 2023;13(3):527. [10] Youn, J.-Y., et al. Properties of stress granule and P-body proteomes. Molecular cell 2019;76(2):286–294. [11] Oughtred, R., et al. The BioGRID database: A comprehensive biomedical resource of curated protein, genetic, and chemical interactions. Protein Science 2021;30(1):187–200. [12] Lin, Z., Akin, H., Rao, R., Hie, B., Zhu, Z., Lu, W., ... & Rives, A. Evolutionary-scale prediction of atomic-level protein structure with a language model.Science 2023;379(6637),1123-1130.