Skip to content
Preprint

Where the Numbers Come From: Auditing Evaluation in Provenance-Based Intrusion Detection

Sep 2026 · 0 citations · 39 references
Computer Science

Abstract

Reproducing a provenance-based intrusion detector's score does not establish what that score says about its emitted alarms or the information its encoder uses. We audit nine released implementations, execute four detectors using their own code, and isolate three measurement effects. First, a fixed-alert comparison separates label choice from neighbourhood credit: ThreaTrace reports precision 0.938 with neighbourhood credit, although only seven of its 994 alarms carry its own attack label. Second, removing test-label checkpoint selection lowers attack detection precision (ADP) by 0.16 to 0.33 across four forty-member word2vec configurations without changing detector order. Third, a buffer-reuse defect gives a linear encoder unintended degree-dependent inputs. Correcting it lowers type-only ADP in every identical-input initialization pair on two hosts, while historical word2vec effects depend on the host. These findings qualify claims from the inspected implementations about alarm precision, performance magnitude and static-attribute sufficiency. STRICT connects them to six checkable reporting requirements. Because the comparisons condition on benchmark targets, they neither validate those labels nor establish a universal detector ranking.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.