Skip to content
Preprint

X-WAD: eXplainable Web Anomaly Detection

Aug 2026 · 0 citations · 25 references
Computer Science

TL;DR

This paper investigates the effectiveness of Transformer-based Language Models in the detection of anomalies in HTTP requests, focussing on providing detailed explanations for the detected anomalies, and employs token-level logit-based surprisal mapping to provide both an anomaly score and a direct, detailed explanation via heatmap-like highlighting.

Abstract

The rapid growth of web-based services, particularly API-driven architectures, reflects an increasing reliance on distributed systems, exposing sensitive data to security risks and making the adoption of automated defensive mechanisms essential. In this context, where benign traffic predominates in real-world settings, modern defenses increasingly model normal behavior, relying on semi-supervised approaches trained on only normal data. However, ensuring the complete absence of anomalous instances in such training data is inherently difficult in practice, and mislabeled or contaminated attack samples can introduce backdoors into the learned defense, causing the model to silently misclassify certain attack patterns as normal behavior. This paper investigates the effectiveness of Transformer-based Language Models (TLMs) in the detection of anomalies in HTTP requests, focussing on providing detailed explanations for the detected anomalies. The study employs token-level logit-based surprisal mapping to provide both an anomaly score and a direct, detailed explanation via heatmap-like highlighting. The effectiveness of the proposed explainability approach is demonstrated by the discovery of labelling inconsistencies in a popular public dataset, revealing how anomalous contamination in the training data had induced backdoor-like failures in the detection models.

View source

Similar papers

Open access Aug 2026

Profiling the Invisible Insider: A UEBA-Based Machine Learning Framework for Low-and-Slow Data Exfiltration Detection

A UEBA-based machine learning framework that constructs per-user behavioral profiles from enterprise proxy and access log data, scoring sessions against a 30-feature behavioral representation spanning temporal patterns, data-transfer anomalies, domain interactions, HTTP characteristics, and session-device signals is pr...

L. Lanuwabang, S. Suprakash · 0 citations
#artificial intelligence Preprint Sep 2026

Comparative Evaluation of Static Embedding Models for HTTP Request Anomaly Detection

Web applications are increasingly targeted by cyberattacks that exploit HTTP requests to evade security mechanisms. Traditional web application firewalls (WAFs) rely on rule-based approaches that often exhibit high false positive rates and limited adaptability. Recent studies have explored machine learning techniques a...

A. Riverol, Gustavo Betarte, R. Martínez et al. · 0 citations
Conference Aug 2026

Explainable Malware Detection from Noisy API Sequences with RAG-Based MITRE ATT&CK Mapping

As sophisticated evasion techniques like polymorphism and staged execution increasingly neutralize conventional signature-based defenses, dynamic API sequence analysis has emerged as an effective approach for malware detection. However, extracting actionable intelligence from noisy execution logs while maintaining mode...

Dat Quoc Phan, Tien Duc Anh Hao, Nghi Hoang Khoa et al. · 0 citations
Sep 2026

Scalable and Adaptive Log-based Anomaly Detection: A Synergistic Approach

System logs are critical for software reliability. While many automated log-based anomaly detection methods exist, they often falter in large-scale cloud systems due to high resource consumption and poor adaptability to evolving logs. In this paper, we present SeaLog, an accurate, lightweight, and adaptive log-based an...

Jinyang Liu, Junjie Huang, Zhihan Jiang et al. · 0 citations

LADE: LLM-Assisted Advanced Persistent Threat Detection and Explanation

Experimental results show that LLMs, when guided by rubric-based prompts and supplemented with ATT&CK domain knowledge, achieve robust performance across detection, localization, and TTP mapping tasks.

Joon-Young Gwak, Aubrey Strier, Zhaohan Xi et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.