Sep 2026· Proceedings of the 19th ACM International Systems and Storage Conference· 0 citations· 14 references
TL;DR
The semantic parser is treated as an existing component, and the contribution is evidence about when that component changes downstream detection behaviour enough to justify pilot testing, which supports cautious, workload-aware use of semantic parsing in observability pipelines.
Abstract
Parsing decisions in log anomaly detection pipelines directly shape downstream detection effectiveness. Despite growing work on semantic parsers and LLM-assisted parsing, the deployment conditions under which semantic enrichment is preferable to efficient syntactic parsing have not been empirically characterised. Semantic-enriched parsing can improve template distinguishability by using lightweight language-model representations. It also introduces extra representation complexity and embedding inference overhead. Systems operators therefore need evidence about when this additional complexity is useful, negligible, or unstable. This is a systems observability design study, not a parser-proposal paper: the semantic parser is treated as an existing component, and the contribution is evidence about when that component changes downstream detection behaviour enough to justify pilot testing. We empirically characterise these boundary conditions by evaluating a semantic-enriched parser against three syntactic baselines (Drain, IPLoM, and Spell) across five public LogHub datasets spanning HPC, distributed storage, and cloud logs. We use three anomaly detectors with different modelling assumptions: DeepLog, LogAnomaly, and CNN. We report mean and standard deviation over five random seeds, define semantic-enrichment benefit as ΔSE = F1(semantic) – F1(Drain), and also report ΔSE* relative to the best syntactic baseline. We introduce the Template Expansion Ratio (TER), a simple post-parsing diagnostic for how much additional template vocabulary the semantic parser creates relative to Drain. Material gains concentrate on BGL and Spirit, negligible gains occur on HDFS and OpenStack, and Thunderbird shows unstable behaviour under very low anomaly prevalence. The results support cautious, workload-aware use of semantic parsing in observability pipelines rather than treating semantic enrichment as a universally beneficial replacement for syntactic parsers.
This work proposes SLOPE, a fine-grained log parser combining syntax with semantics, which achieves an average parsing accuracy improvement of 61.2% and 5.9 times higher throughput over all baselines, exhibiting state-of-the-art robustness.
Shu-Ting Lai, Haiyu Huang, Peng-Fei Chen et al.· ACM Transactions on Software...· 0 citations
PASK (Parser-Aware Structural KV Persistence), which turns parser-derived structure into layer-group-specific KV persistence decisions by using task-error sensitivity to set minimum protection floors and attention-output distortion to allocate residual KV capacity.
FlanBC is presented, a log parsing framework that formulates template extraction as a BIO (Beginning, Inside, Outside) sequence-labeling task and integrates a Flan-T5 semantic encoder, Bidirectional Long Short-Term Memory (BiLSTM) layers for local sequential modeling, and a Conditional Random Field (CRF) decoder for st...
Jinhui Yuan, Bin Guan, Kun Wen et al.· Information· 0 citations
SETYPE is presented, a semantics-aware type system that can be derived directly from source code based solely on the meanings of symbols and expressions in natural language that achieves 87% detection precision and 88% detection accuracy on real-world applications.
Tabular Anomaly Detection (TAD) plays a fundamental role in securing real-world applications. Despite rapid advances in TAD, the prohibitive cost of human-centric label annotation remains a primary bottleneck for large-scale production systems. To alleviate this bottleneck, we propose a novel ''coarse-to-fine'' label a...
Haihong Zhao, Aochuan Chen, Miao Peng et al.· Proceedings of the 32nd ACM...· 0 citations
HGX is introduced, which natively embeds domain dependencies into an expressive parent-child hierarchy, and achieves up to a 92% reduction in ASP solver invocations compared to flat evaluation strategies, while strictly preserving extraction accuracy.
Unknown authors· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.