Skip to content
Review Open access

An Integrated Honeypot and LLM-Based Framework for near Real-Time Detection and Behavioral Analysis of Malicious Activities

Aug 2026 · Journal of Cybersecurity and Privacy · 0 citations · 45 references

TL;DR

Integration of deterministic preprocessing with LLM-based reasoning enables the transformation of raw honeypot logs into structured and actionable cybersecurity intelligence, reducing analyst workload while improving the explainability and reliability of intrusion analysis in near-real-time environments.

Abstract

Malicious activity detection in honeypot environments remains challenging due to the volume and heterogeneity of captured data, as well as the sequential nature of attacker behavior. This study proposes an integrated framework combining a Cowrie-based honeypot with a locally deployed Large Language Model (LLM) to enable automated detection and near-real-time behavioral analysis. Attacker interactions are captured and processed through a structured preprocessing stage that reconstructs session-level activity. These representations are analyzed using an LLM, allowing contextual interpretation of authentication patterns, command execution sequences, and post-compromise behavior. Structured analytical outputs are generated, including severity classification, reasoning, and recommended actions. Evaluation was conducted using isolated and concurrent attack scenarios. Results indicate effective identification of brute-force attacks, reconnaissance activity, persistence staging, and download-and-execute patterns, achieving a 94.23% Accuracy (95% CI: 84.05–98.79%), 100.0% Precision, 87.50% Recall, and an F1-score of 93.33% across an expanded evaluation of 52 independent observations, with zero false positives (FPs). These figures are derived from a single evaluation run and are reported as preliminary, proof-of-concept estimates rather than as stable, production-grade performance. Analytical output remained stable for sessions governed by deterministic severity overrides. However, a low-intensity multi-stage session produced inconsistent severity classifications under concurrent conditions, indicating that analyst review is still required for borderline cases. Integration of deterministic preprocessing with LLM-based reasoning enables the transformation of raw honeypot logs into structured and actionable cybersecurity intelligence, reducing analyst workload while improving the explainability and reliability of intrusion analysis in near-real-time environments.

Read PDF

Similar papers

Preprint Sep 2026

LLM-Based Penetration Testing in the Presence of Honeypots

This work presents a systematic study of honeypot-aware budget allocation for LLM attack agents and shows that with the proposed detector-guided policy, LLM agent attackers can effectively allocate budget to compromise hosts in a host pool, highlighting the importance of dynamically allocating budget in a controlled mi...

Xin-Hong Xie, Piyush Nagasubramaniam, Neeraj Karamchandani et al. · 0 citations
Open access Aug 2026

Profiling the Invisible Insider: A UEBA-Based Machine Learning Framework for Low-and-Slow Data Exfiltration Detection

A UEBA-based machine learning framework that constructs per-user behavioral profiles from enterprise proxy and access log data, scoring sessions against a 30-feature behavioral representation spanning temporal patterns, data-transfer anomalies, domain interactions, HTTP characteristics, and session-device signals is pr...

L. Lanuwabang, S. Suprakash · 0 citations
#generative ai Open access Sep 2026

Hive-AI: a defended multi-service honeypot framework for generative AI APIs

HIVE-AI, a 47,578-LoC honeypot framework deployed continuously on a single 4-vCPU/4-GB Virtual Private Server since 6 April 2026, is presented, a promising low-cost alternative rather than a full substitute for open-source honeypot frameworks.

Sebastián Vargas Yáñez, Sergio Tobón · 1 citation
Review Open access Aug 2026

Adaptive Defense: Enhancing NIDS with Smart Honeypots and Attack Profiling

The research concludes that combining honeypot intelligence with machine learning improves real-time identification, proactive defence, as well as reducing the false alarms.

Unknown authors · 0 citations
Open access Sep 2026

AI-Augmented Network-Forensics: Leveraging LLMs for Real-Time Threat Detection and Automated Response in Enterprise Environments

In modern enterprise networks, complicated rule-based signatures, fragmented alerts, encrypted traffic, and analyst workloads are delaying the ability to recognize and contain incidents, as the need grows for faster correlation of heterogeneous telemetry. This study evaluates an LLM-augmented network-forensics architec...

Mohammed Imran Choudhary · 0 citations
Preprint Sep 2026

OllamaDrama: Designing and Deploying a Honeypot to Measure Attacks on Exposed LLM Infrastructure

Publicly exposed large language model (LLM) infrastructure creates a growing attack surface, yet real-world targeting remains poorly understood. We present Ollure, a low- and medium-interaction honeypot that emulates the Ollama API without a backend LLM. Spanning four deployments across cloud and university networks, O...

Karina Elzer, Niklas Netterstrøm Johansen, Emmanouil Vasilomanolakis · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.