CAMELEON-AI X v3.0: Enhanced Binding Affinity Prediction Using Stacking Ensemble with Ridge Meta-Learner
Abstract
Binding affinity estimation is an important computational task in modern drug discovery because it helps prioritize compounds according to their expected interaction strength with protein targets. This study presents CAMELEON-AI X v3.0, a machine-learning framework designed to improve binding affinity prediction on heterogeneous multi-target data. The proposed pipeline combines target-stratified Z-score normalization, target-level statistical encoding, complementary molecular fingerprints, two-dimensional molecular descriptors, and a stacking ensemble whose second-level learner is Ridge regression. Four diverse first-level models, namely XGBoost, LightGBM, RandomForest, and ExtraTrees, generate out-of-fold predictions that are subsequently used by the meta-learner. The feature representation contains 2,760 dimensions, including ECFP4, ECFP6, MACCS keys, an RDKit fingerprint, 30 molecular descriptors, and three target-encoding variables. Evaluation was performed on 29,630 BindingDB measurements, using an 85% training split and a 15% test split. The stacking model obtained R² = 0.7942, RMSE = 0.9712 kcal/mol, MAE = 0.7058 kcal/mol, PCC = 0.8913, and SCC = 0.8825. Relative to the RandomForest baseline with R² = 0.4900, the resulting R² gain was 0.3042, corresponding to a 62.1% improvement. The model also exceeded the simple-average ensemble, which obtained R² = 0.657. Five-fold cross-validation produced a mean R² of 0.7750 with a standard deviation of approximately 0.007, indicating stable performance across folds. The findings show that learned ensemble weighting and target-aware preprocessing can provide a substantial benefit for binding affinity prediction while retaining a comparatively lightweight machine-learning pipeline.