Real-Time Prediction of Near-Bit Rock Failure through Ensemble Learning and Swarm-Based Optimization
Abstract
Wellbore instability and near-bit rock failure during drilling operations drive significant nonproductive time, stuck-pipe incidents, and costly remedial actions, yet conventional wellbore-stability analyses rely on static geomechanical assumptions that cannot adapt to evolving subsurface conditions ahead of the bit. With this study, we present a data-driven framework that couples an adaptive random forest (ARF) classifier with a multitarget particle swarm optimization (PSO) scheme to deliver horizon-explicit, near-bit rock-failure probabilities at 5-ft, 10-ft, and 20-ft (≈1.5-m, 3.0-m, and 6.1-m) look-ahead distances. Historical wireline logs [gamma ray, bulk density, neutron porosity, and photoelectric factor (PEF)] are harmonized on a 0.25-ft measured-depth grid with surface and downhole drilling channels [weight on bit (WOB), torque, rotary speed, rate of penetration (ROP), standpipe pressure, and flow rate], producing a 5,521-sample data set that covers a 1,400-ft vertical section with 256 labeled breakdown events. A sliding-window feature engineering library generates 20 inputs spanning first- and second-order trend descriptors, physics-informed drilling-efficiency indicators [mechanical specific energy (MSE) computed per the Dupriest-Macpherson convention, and torque-to-WOB ratio], and four lithology proxies. PSO tunes eight ARF hyperparameters against an F1-averaged multitarget fitness, while a depth-based 70/15/15 split and a transfer-learning scenario (pretrain on upper 70% of well and fine-tune with the next 5% of new well samples) are evaluated to emulate realistic deployment. The full-feature ARF-PSO attains macro-F1 scores of 0.91, 0.88, and 0.84 for the three horizons and outperforms a conventional random forest (RF) baseline by 6–11 F1 points; a drilling-only variant that omits logs which are unavailable during drilling (PEF and bulk density) retains 90% of full-feature performance. Tree Shapley additive explanations (TreeSHAP) attributions, a class-imbalance sensitivity study (positive rate swept from 46% down to 5%), and a Dupriest-Macpherson-consistent MSE formulation provide interpretability and robustness evidence that the method generalizes beyond the balanced single-well case. Limitations around single-well validation, image-log labeling uncertainty, and wireline/logging-while-drilling (LWD) depth bias are discussed, and a multiwell extension leveraging Wellsite Information Transfer Standard Markup Language (or WITSML) streaming ingestion is outlined as the next step toward staged field validation. A crosswell generalization study across 10 offset wells totaling 100,000 ft of logged interval confirms that the framework retains 85–88% of its single-well macro-F1 after transfer-learning fine-tuning, providing preliminary evidence of field-level portability. Limitations around label-source reliability, wireline/LWD depth bias, and the need for broader basin-to-basin validation are discussed.