Algorithmic Diffusion on YouTube: A Machine Learning Analysis of Channel-Level Information Spread and Its Cross-Platform Generalisability
Abstract
(1) Background: Information diffusion models developed for graph-based platforms such as Reddit and broadcast architectures such as Telegram identify temporal features—particularly the timing of peak spread—as dominant predictors of coverage. Whether these predictors generalise to platforms where content is distributed through algorithmic recommendation rather than social-graph contagion remains an open question. (2) Methods: We analyse the YouNiverse dataset, comprising 133,364 English-language YouTube channels observed weekly from January 2015 to September 2019 (18.9 million observations). We derive channel-level diffusion features—including time-to-peak, post-peak decay rate, diffusion volatility, and upload frequency—and train three machine learning models (Linear Regression, Random Forest, and LightGBM) on two tasks: predicting peak weekly view growth (regression) and identifying viral channels (classification). A single-feature naive baseline (subscriber count alone) establishes the marginal contribution of the broader feature set beyond subscriber count alone, and a temporal split experiment (training on channels peaking before 2018, testing on 2018–2019) assesses cross-temporal stability. Because subscriber count and subscriber rank are measured at the October 2019 crawl, this is a retrospective characterisation rather than a strict real-time forecasting design. (3) Results: LightGBM achieves R2=0.776 (5-fold CV: 0.778±0.003) compared with R2=0.548 for the naive baseline, a net gain of +0.228R2. Because subscriber rank and subscriber count are near-perfectly collinear, we interpret them jointly as a channel-size dimension (42.2% of total mean absolute SHAP attribution), rather than as independent effects. Time-to-peak ranks fourteenth (1.1%), in contrast to its dominant role on Reddit (r=0.995, rank #1). For virality classification, LightGBM achieves ROC-AUC =0.967. Under the temporal split, Random Forest (R2=0.703) outperforms LightGBM (R2=0.683), showing greater cross-temporal stability within this retrospective split. (4) Conclusions: Within the 2015–2019 data, the results are consistent with algorithmic recommendation weakening the relationship between temporal diffusion dynamics and coverage magnitude at the channel level. Time-to-peak is weakly informative in this setting, while generalisation to the current recommendation system requires validation on newer data.