Structured PREreview of "Mitigating LLM Prompt Injections via Multi-Vector Representation and Voting-Based Routing"
Abstract
This Zenodo record is a permanently preserved version of a Structured PREreview. You can view the complete PREreview at https://prereview.org/reviews/22919717. Does the introduction explain the objective of the research presented in the preprint? Yes Justification: The introduction and abstract clearly articulate the security problem, architectural motivation, and core research objectives: Problem Context: Existing prompt injection defenses often suffer from high false-positive rates (the base-rate fallacy), disrupt user experience, or rely on proprietary black-box APIs or intrusive weight fine-tuning . Furthermore, single-embedding classifiers exhibit latent-space blind spots and tokenizer dropouts when handling non-standard Unicode or dense code blocks. Core Research Objectives: The introduction explicitly formulates three guiding research questions: RQ1: How does integrating a multi-vector semantic diversity layer with a voting-based consensus routing mechanism influence detection efficiency and false-positive rates? RQ2: Can a bidirectional verification framework (filtering both input prompts and generated responses) eliminate "silent failures" when input guardrails are bypassed? RQ3: Is it possible to deploy an enterprise-grade prompt injection defense using exclusively sub-billion parameter, open-source models on local edge hardware? Scope & Pipeline Architecture: The authors present a multi-layered defense pipeline combining a Triple Modular Redundancy (TMR) embedding layer (Snowflake Arctic-Embed, IBM Granite, MiniLMv2), a 6-classifier ensemble (XGBoost and LightGBM), an adaptive consensus router, and Llama Guard 3 as an expert fallback and output interceptor . The framework is empirically evaluated on a stratified dataset of 93,398 test prompts targeting a local Small Language Model (Gemma 3:1B). Are the methods well-suited for this research? Highly appropriate Justification: The multi-layered defense methodology is well-engineered, sound, and directly addresses single-point-of-failure vulnerabilities inherent in single-embedding classifiers. Methodological Strengths: Triple Modular Redundancy (TMR) Embedding Layer: User inputs are simultaneously mapped into three distinct vector spaces using small open-source encoders under one billion parameters: Snowflake Arctic-Embed (33 million parameters), IBM Granite-Embedding (30 million parameters), and Sentence-Transformers MiniLMv2 (33 million parameters). This structural diversification prevents individual tokenizer dropouts and mitigates latent-space blind spots. Dual Gradient Boosting Ensemble: The framework trains two complementary gradient boosting decision tree models—XGBoost (optimized for high specificity) and LightGBM (optimized for high sensitivity)—across all three vector representations. This produces a robust six-classifier ensemble. Confidence-Based Consensus Routing: Predictions are evaluated using an adaptive vote margin defined as the absolute difference between safe and malicious votes across the six models. Queries with strong consensus take the Fast Route for immediate forwarding or blocking. Uncertain queries with a vote margin of two or fewer trigger the Slow Route, delegating the decision to an expert fallback model, Llama Guard 3. Bidirectional Output Interception: A secondary Llama Guard 3 filter inspects generated LLM output responses to catch late-execution payloads, system prompt leaks, and data exfiltration attempts that managed to bypass input guardrails. Edge Hardware Validation: The entire system was implemented and evaluated on local consumer hardware (an Intel i7 processor with an NVIDIA GeForce RTX 4070 GPU running Ollama) using a deduplicated, stratified dataset of 93,398 test prompts. Constructive Feedback: Quantifying Edge Latency Bottlenecks: While the Fast Route achieves fast decision times, processing ambiguous queries through the Slow Route introduces an average end-to-end latency peak of over twelve seconds on edge hardware. Adding a detailed breakdown separating local embedding calculation times from Ollama server overhead would provide clearer insights for real-time production deployment. Are the conclusions supported by the data? Somewhat supported Justification: The empirical findings across the 93,398 test prompts provide strong statistical support for the core architecture claims, though certain boundary conditions and failure modes are explicitly noted. Key Empirical Strengths: End-to-End Pipeline Metrics: Enabling the full end-to-end pipeline (confidence routing plus output filtering) increased overall classification accuracy from an ensemble baseline of 89.23% to 90.69%. Overall recall (sensitivity) improved from 72.11% to 78.68%, while maintaining a low False Positive Rate of 5.89% and a high specificity. Bidirectional Output Interception: The secondary Llama Guard 3 output filter successfully intercepted 1,359 malicious responses that had bypassed initial input classification, reducing False Negatives by 23.5% without increasing false alarms . Triple Modular Redundancy (TMR) Fault Tolerance: While individual embedding tokenizers experienced dropouts on complex Unicode or dense code blocks (Snowflake dropped 8,427 prompts, MiniLM dropped 5,525, and Granite dropped 2,072), zero concurrent tri-model tokenization dropouts occurred across all 93,398 test cases. Constructive Bounds & Limitations: Confidently Wrong High-Consensus Errors: The authors transparently identify a key vulnerability: 1,987 false negatives (representing 19.7% of all errors) occurred when all ensemble classifiers were confidently incorrect with high probability (greater than or equal to 0.80). Because these misclassifications exhibited low entropy, they bypassed the Llama Guard 3 fallback. Model Scale Scope: Testing was conducted using Gemma 3:1B as a representative edge Small Language Model. While Gemma 3:1B serves as a useful worst-case baseline for edge vulnerability, evaluating whether these exact interception rates generalize to multi-billion parameter enterprise LLMs or agentic tool-use environments remains an important area for future study. Are the data presentations, including visualizations, well-suited to represent the data? Somewhat appropriate and clear Justification: The visualizations and data presentations in the manuscript are informative, logically structured, and aligned with the defense pipeline, though minor visual formatting adjustments and an additional latency chart would further enhance readability. Visual Strengths: Architectural Flowchart (Figure 1): Figure 1 provides a clean, self-contained overview mapping the five-stage defense pipeline (from multi-vector mapping to consensus voting, Small Language Model execution, and output verification). Performance Benchmark Charts (Figures 3 and 4): Figures 3 and 4 effectively contrast individual sub-model performance, clearly highlighting the operational trade-off between XGBoost (higher specificity and lower false positive rate) and LightGBM (higher sensitivity and lower false negative rate) across all three embedding models. Threat Interception Breakdown (Figure 6): Figure 6 presents a clear bar chart breaking down the top threat categories caught at the input guardrail versus the output guardrail, demonstrating where the secondary filter adds value. Side-by-Side Confusion Matrices (Figure 7): Figure 7 displays side-by-side confusion matrices (evaluated on the 93,398 test set) comparing the baseline ensemble against the full end-to-end pipeline, making the reduction in false negatives easy to interpret. Constructive Feedback & Areas for Improvement: Visual Formatting Inconsistencies (Figure 2): In Figure 2 (showing missing predictions per embedding model due to tokenization dropouts), the bar styling and axis formatting could be harmonized with the rest of the manuscript's visual theme. Text Label Sizing in Summary Diagrams (Figure 5): In Figure 5 (ensemble voting accuracy breakdown), the text annotations and percentage labels are relatively small and can be difficult to read in lower-resolution print views. Latency Profile Plot: While the manuscript discusses latency metrics (such as Fast Route execution times and 95th-percentile tail latency peaks), adding a dedicated line plot or bar chart illustrating inference time distributions across Fast Route versus Slow Route calls under load would make the latency trade-off much clearer. How clearly do the authors discuss, explain, and interpret their findings and potential next steps for the research? Somewhat clearly Justification: The authors provide clear analytical explanations for their architectural choices, transparently analyze system failure modes, and outline actionable future research directions. Strengths in Discussion and Interpretation: Architectural Rationale: The discussion clearly justifies why an even number of six classifiers (3 embedding spaces multiplied by 2 decision tree algorithms) was intentionally chosen: to induce statistical ambiguity (a vote margin of 2 or fewer) during obscure attacks, effectively triggering the Slow Route fallback to Llama Guard 3. Analysis of Failure Modes: The authors transparently examine the primary vulnerability of their system: "confidently wrong" low-entropy false negatives. They analyze how 1,987 malicious prompts bypassed the fallback filter because all six ensemble models misclassified them with high confidence (probability greater than or equal to 0.80). Va