Skip to content
Open access

Dual-Threshold Conformal Deferral for Trustworthy Security Alert Triage

Sep 2026 · Electronics · 0 citations · 23 references

Abstract

Automated alert triage can reduce Security Operations Center (SOC) workload, yet the validation-tuned thresholds deployed systems rely on carry no finite-sample control of their operational error rates and degrade unpredictably under distribution shift. We present a model-agnostic dual-threshold conformal deferral architecture: high-score alerts are auto-escalated under finite-sample marginal class-conditional control of the benign-escalation probability (budget α), low-score alerts are auto-closed under matching control of the threat-miss probability (budget β), and the rest are deferred to an analyst. It needs no retraining and closes an automatic zone rather than certifying what the calibration data cannot support. We evaluate it on a reinforcement-learning investigation agent in a simulated SOC and on four classifiers trained on CIC-IDS2017 and tested on CSE-CIC-IDS2018, using stratified 25,000-flow calibration and evaluation samples, with attack-type recall computed over the full 16.2-million-flow corpus. Pooling episodes from ten trained policies across two evaluation datasets, the architecture automated 73.7% of decisions at α = β = 0.01—a figure for that predefined pooled mixture rather than a per-policy or per-dataset guarantee—realizing benign auto-escalation and threat auto-close rates of 0.0099 and 0.0101 and deferring the hardest ~26% of alerts. After recalibration on labeled target-domain data, severe cross-dataset degradation appears not as a silent error but as sharply reduced certifiable automation, with deferral rising to 79–99% for the most affected classifiers. This visibility is a property of the recalibrated layer: thresholds left un-recalibrated after a shift continue to certify nothing while still deciding, so the architecture requires periodic recalibration on labelled target-domain alerts to deliver it. Substituting open-weight language models for the analyst inside the band failed a pre-specified criterion at every scale tested from 7B to 32B across two model families, with the discriminative signal flat in model size and far below the first-stage policy’s own.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.