Skip to content
Open access

Real-time NLP moderation on XMPP with GPU-accelerated inference and adaptive batching

Aug 2026 · Scientific and Technical Journal of Information Technologies, Mechanics and Optics · 0 citations · 7 references

Abstract

Real-time communication platforms generate continuous streams of short, latency-sensitive messages, creating a demanding environment for automated toxicity detection. Transformer-based language models offer strong contextual accuracy, but their inference cost makes them difficult to deploy efficiently in decentralized messaging systems such as the Extensible Messaging and Presence Protocol, where requests arrive asynchronously from multiple servers. Traditional batching strategies struggle in this setting because traffic patterns fluctuate, message lengths vary, and strict latency budgets prevent the accumulation of large batches. This work introduces a full-stack moderation system that combines an OpenFire server plugin, a Graphics Processing Unit (GPU) accelerated microservice for toxicity classification, and an adaptive batch processing method suitable for use in multi-server systems. The plugin intercepts messages at the packet-processing layer and forwards them to a lightweight external inference service that can run on either a Central Processing Unit or a GPU. We introduce Adaptive Cross-Domain Batching (ACDB) as a method to dynamically adjust batch sizes based on the state of the queues, characteristics of the messages being processed, and real-time feedback from GPU usage. Experiments demonstrate that GPU inference provides significant improvements in both throughput and latency, with compute accelerations of up to 28× and total end-to-end accelerations of up to 9×. Comparative evaluation against static batching baselines across multiple traffic scenarios shows that ACDB dynamically adapts batch sizes from 3 to 25 depending on load conditions, reducing latency by 52–59 % compared to large-batch static configurations while maintaining 85–94 % of their throughput. Overall, this system enables scalable, real-time toxic content moderation capabilities for federated chat systems using adaptive batching for Transformer-based inference on dynamic workloads, while preserving strict message delivery constraints under realistic operating conditions.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.