Metadata Compressibility and Evaluation Bias in Malicious Package Detection for NPM and PyPI
Abstract
Malicious packages in open source-software supply chains are a growing security concern, yet machine learning detectors built on registry metadata are difficult to interpret and are typically evaluated under protocols susceptible to data leakage. We construct a dataset of 3330 package versions from NPM and PyPI in which every attribute is reconstructed at its exact publication timestamp, and we partition the data so that all versions of a package, all packages of a maintainer, and all members of a name-similarity family reside within a single fold. The primary analysis is restricted to NPM, where 18 of 456 point-in-time attributes attain an AU-PR of 0.9972 against a positive-class baseline of 0.8755 and an AUC-ROC of 0.9828 on 763 held-out versions. A secondary pooled analysis over both ecosystems is reported: the class priors differ by nearly a factor of nine, a classifier using only the ecosystem of origin attains an AUC-ROC of 0.8036, and the pooled figures should be read with that composition in mind. Confidence intervals are obtained by cluster-bootstrap resampling at the composite-group level and are 3.7-times wider than version-level intervals in AU-PR (0.0201 against 0.0055); at that scale, none of the differences between feature subsets, selection-stage orderings, or metadata families reported here is distinguishable from sampling variation. At an operational prevalence of 0.1%, the positive predictive value is 4.3% and recall at the selected threshold is 0.537, indicating that metadata-based detection functions as a triage filter rather than a definitive verdict.