Skip to content
Open access

Is There a Best Hypergraph Neural Network? A Significance-Aware Recomputation and Statistical Audit of DHG-Bench

Jul 2026 · Machine Learning and Knowledge Extraction · 0 citations · 34 references

Abstract

Deep hypergraph learning is evaluated almost entirely through leaderboards that rank methods by mean accuracy over a few random seeds, usually without significance testing. Is there a best hypergraph neural network, or does the apparent ordering reflect seed noise? We independently recomputed the node-classification track of DHG-Bench on a single GPU with twenty random seeds (against five upstream) and a different software stack, and applied a four-layer statistical audit to the per-seed accuracies: a reproducibility check, per-dataset paired Wilcoxon tests with Holm correction, an across-datasets Friedman/Iman–Davenport omnibus with Nemenyi and Holm-corrected pairwise tests, and a variance decomposition. Within a single dataset, twenty seeds distinguish most method pairs (74–98%), so the protocol is not underpowered. Across the nine datasets where all 17 methods complete, the omnibus rejects global equality (Kendall’s W=0.45), yet no pair survives Holm correction, and the top methods fall within one critical-difference band. One dataset carries more seed noise than between-method signal and cannot rank methods. The recompute also documents a non-reproducible method, a label-range data fault, and missing per-dataset configurations in the public release. No single method is statistically best across these datasets, so single-leader claims are not supported; we release a reusable significance-aware evaluation protocol.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.