Skip to content

Beyond Model Complexity: A Reproducible Comparison of Classical Machine Learning, Matrix Factorization, Graph Embeddings, and LightGCN for Recommendation

Sep 2026 · Algorithms · 0 citations · 17 references
Recommender Systems and Techniques

TL;DR

Findings show that under the evaluated setting, greater model complexity did not consistently translate into higher recommendation effectiveness, and they thus highlight the importance of strong baselines, model tuning, standardized evaluation, and reproducible experimental protocols.

Abstract

Recommender systems increasingly incorporate graph embeddings and graph neural networks to capture high-order relationships between users and items. However, the additional complexity of these approaches does not necessarily guarantee better recommendation quality than strong classical and latent-factor baselines. This study presents a reproducible comparison of six recommendation models representing four methodological families: Logistic Regression and Random Forest; Matrix Factorization with Bayesian Personalized Ranking; DeepWalk and node2vec; and LightGCN. The experiments were conducted on the MovieLens 1M dataset using a per-user temporal split. For each user, the most recent positive interaction was assigned to testing, the preceding interaction to validation, and all earlier positive interactions to training. The primary evaluation used identical candidate sets containing one held-out positive movie and 99 sampled unobserved movies. Performance was measured using Recall, Precision, Hit Rate, and NDCG at multiple cutoffs, complemented by bootstrap confidence intervals, paired statistical tests, computational-efficiency measurements, and analyses by user activity and movie popularity. Matrix Factorization achieved the best overall performance, reaching a Recall@10 of 0.7458 and an NDCG@10 of 0.4558, representing an approximately 56% improvement in NDCG@10 over Random Forest, the strongest classical baseline. Validation-based tuning improved LightGCN to an NDCG@10 of 0.2875; it significantly outperformed Logistic Regression but remained statistically indistinguishable from Random Forest after Holm correction. Tuned node2vec also significantly outperformed DeepWalk, reaching an NDCG@10 of 0.1593, although both random-walk embedding methods’ results remained substantially below than the strongest baselines. Popularity-based analysis further revealed that classical models and LightGCN achieved substantially higher ranking effectiveness for popular movies, whereas Matrix Factorization maintained comparatively stronger performance for less-popular items. These findings show that under the evaluated setting, greater model complexity did not consistently translate into higher recommendation effectiveness, and they thus highlight the importance of strong baselines, model tuning, standardized evaluation, and reproducible experimental protocols.

Read PDF

Similar papers

Co-occurrence graph neural network for recommender systems

A novel neural network called the Co-occurrence Graph Neural Network (CoGNN), which utilizes two co-occurrence graphs to establish user and item relationships and outperforms various baseline models in terms of recommendation accuracy and algorithm convergence.

Chao Lin, Y. Lin, You-Yu Wang et al. · 0 citations
Conference Sep 2026

A graph contrastive learning recommendation algorithm based on variational inference

The results confirm that SGICL effectively enhances recommendation accuracy, robustness, and the capacity to learn rich latent graph representations, providing a practical and scalable framework for graph-based recommender systems.

Chi Ma, Haohao Cao, Hui Hu et al. · 0 citations
Review Open access 2026

DISR: Distributional Item Representations for Sequential Recommendation

Sequential recommender systems predict the next item from a chronological interaction history, typically representing users and items as point embeddings scored by dot product. DISR (Distributional Item Representations for Sequential Recommendation) instead represents items and sequence queries as diagonal Gaussian dis...

Abdelilah Bajjou, E. Nfaoui · 0 citations
Preprint Aug 2026

POI Recommendation with LLM-Augmented Multi-Graph Learning and Contrastive Alignment

The proposed LLM-augmented Multi-Graph Contrastive Learning (LLM-MGCL) is a multi-graph neural network that uses semantic and spatial information about items to extend the LightGCN backbone with two auxiliary item-item graphs that outperforms classical collaborative filtering, matrix factorization, and interaction-only...

Burak Tamer, W. Höpken, Zehui Wang · 0 citations

Related blog posts

Microsoft Research Blog Jul 13, 2026

Verifying Rust cryptography in SymCrypt, from standards to code

Cryptographic code supports vital protections in modern computing systems. Learn how a new method helps verify code as developers write it while preserving speed and adaptability as it gets implemented and evolves. The post Verifying Rust cryptography in SymCrypt, from standards to code appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.