Skip to content
Open access

FlanBC: A Semantic-Structural Sequence Labeling Framework for Log Parsing

Aug 2026 · Information · 0 citations · 12 references

TL;DR

FlanBC is presented, a log parsing framework that formulates template extraction as a BIO (Beginning, Inside, Outside) sequence-labeling task and integrates a Flan-T5 semantic encoder, Bidirectional Long Short-Term Memory (BiLSTM) layers for local sequential modeling, and a Conditional Random Field (CRF) decoder for structured label prediction.

Abstract

Log parsing converts raw system logs into structured templates and is a key preprocessing step for Artificial Intelligence for IT Operations (AIOps). Existing parsers face a practical trade-off: rule-based methods offer high throughput but limited adaptability across heterogeneous log sources, whereas Large Language Model (LLM)-based parsers achieve broader semantic coverage at the cost of inference latency, privacy exposure, and cloud dependency. This paper presents FlanBC, a log parsing framework that formulates template extraction as a BIO (Beginning, Inside, Outside) sequence-labeling task and integrates a Flan-T5 semantic encoder, Bidirectional Long Short-Term Memory (BiLSTM) layers for local sequential modeling, and a Conditional Random Field (CRF) decoder for structured label prediction. Log-specific preprocessing and a subword-to-token alignment mechanism adapt the general-purpose encoder to semi-structured log data. A layer-freezing strategy reduces the number of parameters updated during training. The framework supports local inference without external API dependency. Experiments on three benchmark datasets from LogHub (HDFS, BGL, OpenStack) under a supervised random-split setup evaluate parsing accuracy, training efficiency, statistical stability across random seeds, and component contributions. FlanBC achieves a Group Accuracy of 99.32% on HDFS and 98.47% on BGL, with an inference throughput of 700+ logs/s on a consumer-grade GPU. On OpenStack, performance is lower (GA = 92.54%), reflecting the challenge that diverse natural-language-like logs pose for compact encoder-based models. Under a stricter template-disjoint split that prevents template overlap between training and test sets, FlanBC achieves an average Group Accuracy of 91.14%, indicating that the model generalizes to unseen templates beyond in-distribution recognition. Ablation results indicate that the semantic encoder, BiLSTM module, and CRF decoder each contribute to prediction accuracy. These findings suggest that domain-adapted semantic encoders combined with structured decoding offer a practical accuracy–efficiency balance for log parsing in settings where local, cloud-free inference is preferred.

Read PDF

Similar papers

Aug 2026

SLOPE: Fine-Grained Log Parser Combining Syntax with LLM-Distilled Semantic

This work proposes SLOPE, a fine-grained log parser combining syntax with semantics, which achieves an average parsing accuracy improvement of 61.2% and 5.9 times higher throughput over all baselines, exhibiting state-of-the-art robustness.

Shu-Ting Lai, Haiyu Huang, Peng-Fei Chen et al. · 0 citations
2026

LogPISA: An Improved Pre-Training and Tuning Pipeline for Log Understanding With Invariant and Semantic-Aware Objectives

With the rapid development of computer and network technology, network and software logs generated by a multitude of devices contain a wealth of knowledge and serve as a critical resource for intelligent fault diagnosis and efficient system operations. In recent years, various deep learning methods and the pre-training...

Lanlan Rui, Yuan-Rui Yang, Peng Yu et al. · 0 citations
Open access Aug 2026

Neural Turing Machines for efficient natural language summarization: architecture, optimization, and performance analysis

Introduction Abstractive text summarization remains a fundamental challenge in Natural Language Processing (NLP), particularly for long documents that require models to preserve long-range dependencies and maintain semantic coherence. Although Transformer-based architectures have achieved strong summarization performan...

K. Katti, K. Katti, Amanul Islam · 0 citations
Conference Open access Aug 2026

A Cost-Efficient Routing Pipeline for Multilingual Short-Text Classification Using Small Language Models

This work evaluates a fixed-list routing strategy that keeps stronger languages on a direct multilingual path and selectively sends weaker languages through translation into English before zero-shot classification, and reports routing through tier-level quality gains and tier-level latency rather than a single global e...

Wajdi Ben Saad, Safa Madiouni · 1 citation
#small language model Open access Sep 2026

Low-Resource SLM Models for Elastic Semantic Search

Semantic search over domain-specific corpora requires an effective embedding model and infrastructure. Elasticsearch’s native semantic_text field and ELSER sparse-vector inference require a commercial Enterprise subscription, inaccessible to most academic institutions. This paper documents the licensing barrier and pre...

S. Cheresharov, Georgi Gustinov, V. Tabakova-Komsalova et al. · 0 citations
Open access Aug 2026

Ensemble-Based Approach for Amazigh POS Tagging: Leveraging Multiple Models for Enhanced Performance in Low-Resource Language Processing

The results show that hybrid ensemble methods can improve token-level accuracy in low-resource POS tagging, while also revealing a trade-off between frequent-tag accuracy and rare-tag robustness.

Abdelouahed Moussaoui, Nor-Eddine Azalmad, Said Bahassine et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.