Skip to content
Conference

A Dataset-Aware Drift-Robust Log Anomaly Detection Pipeline for Cloud Service Reliability

Aug 2026 · Moratuwa Engineering Research Conference · pp. 181-186 · 0 citations · 18 references

Abstract

Cloud services require continuous log monitoring to detect failures before outages occur. Most log anomaly detectors rely on fixed event identifiers or templates, which degrade as services evolve, parsers change, and new templates appear, leading to a mismatch between offline accuracy and online reliability. We propose a drift-resistant log anomaly detection pipeline that is simple, data-sensitive, and resilient to template evolution. The method builds session-level TF-IDF-based representations combining semantic, event-identity, and structural features. A logistic regression scorer processes chronological log sessions, while a streaming controller increases the decision threshold when the out-of-vocabulary ratio indicates drift. Evaluation on the public Hadoop Distributed File System dataset uses a reproducible chronological split and a semantic-preserving drift protocol. The method achieves an F1-score of 0.567 with zero false positives in the in-distribution setting and retains an F1-score of 0.558 under drift, while improving AUC from 0.696 to 0.748 over the best semantic baseline and avoiding the collapse seen in event-count methods. These results show that lightweight semantic modeling combined with explicit drift handling improves log reliability without online retraining.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.