Skip to content

An Embarrassingly Simple Detector for Model Extraction Attacks in Large Language Model API Traffic

Jun 2026 · arXiv.org · Vol abs/2606.05725 · 1 citation · 50 references
Computer Science

TL;DR

This work forms model extraction monitoring as benign-calibrated traffic-window distribution testing: embed incoming queries into a semantic space and test whether their aggregate distribution deviates from historical benign traffic.

Abstract

Large language models (LLMs) are increasingly deployed through hosted APIs, making model extraction a practical threat to model ownership and service security. Individual extraction queries often resemble benign requests, while existing evaluations often focus on single-query anomaly scoring or pure benign-versus-attacker user settings. We formulate model extraction monitoring as benign-calibrated traffic-window distribution testing: embed incoming queries into a semantic space and test whether their aggregate distribution deviates from historical benign traffic. We instantiate this formulation with maximum mean discrepancy (MMD), using only benign-vs-benign comparisons to set the decision threshold. We evaluate on fourteen attacker-normal query pairs from four extraction scenarios and compare with adapted PRADA, SEAT, CAP, DATE, marginal Mahalanobis, and pseudo-class energy baselines. Across three random seeds, MMD achieves 0.3% benign FPR, 100.0% pure-attacker TPR, 90.5% average TPR over attacker fractions, and 95.1% balanced accuracy. These results show that benign-calibrated distribution testing is a strong empirical baseline for model extraction detection in both user-level and mixed multi-user LLM API traffic.

View source

Similar papers

2026

LLMKernelBench: Benchmarking Large Language Models on Software Vulnerability Detection in Linux Kernel

Large language models (LLMs) demonstrate strong capabilities in code-related tasks, however their effectiveness in software vulnerability detection (SVD) remains poorly understood due to inadequate evaluation frameworks. Existing benchmarks suffer from training data contamination, isolated function evaluation without c...

Arastoo Zibaeirad, Rodrigo Pato Nogueira, Marco Vieira · 0 citations
Open access Aug 2026

Real-Time Detection and Mitigation of Prompt Injection Attacks in LLM-Integrated Enterprise Systems

Large language models (LLMs) embedded in enterprise workflows cannot structurally distinguish legitimate instructions from adversarial ones in the same token stream, making prompt injection OWASP's top LLM risk for two consecutive editions a persistent threat across direct and indirect vectors. This paper presents Prom...

Fatimah Alhamzawi · 0 citations
#artificial intelligence Preprint Sep 2026

SCRIPTIOC-BENCH: A Benchmark for Recognizing Actionable Threat Intelligence from Script-Based Malware using LLMs

This work presents SCRIPTIOC-BENCH, a benchmark for measuring static IOC extraction capability on real-world malicious scripts, and evaluates a broad range of proprietary and open-weight LLMs, showing that IOC recovery without execution remains challenging across model scales.

Hanna Kim, Jian Cui, Minkyoo Song et al. · 0 citations
Preprint Sep 2026

Do Defenses Against LLM Extraction Work Across Attacks? A Lifecycle Benchmark of Black-Box Model Extraction

Large language models (LLMs) deployed through text-only APIs face model extraction risks, as adversaries can collect their responses to train surrogates that reproduce their capabilities. While prior work has developed diverse attacks and defenses, evaluations remain fragmented across access assumptions, model configur...

Shu-Ze Liu, Kai-Xiang Zhao, Run-Yang Xu et al. · 0 citations
Conference Aug 2026

Explainable Malware Detection from Noisy API Sequences with RAG-Based MITRE ATT&CK Mapping

As sophisticated evasion techniques like polymorphism and staged execution increasingly neutralize conventional signature-based defenses, dynamic API sequence analysis has emerged as an effective approach for malware detection. However, extracting actionable intelligence from noisy execution logs while maintaining mode...

Dat Quoc Phan, Tien Duc Anh Hao, Nghi Hoang Khoa et al. · 0 citations
Open access Dec 2025

AI security beyond core domains: resume screening as a case study of adversarial vulnerabilities in specialized LLM applications

A benchmark for this vulnerability in LLM-based resume screening is introduced: 463 job-candidate pairs drawn from a 14-domain corpus, with the evaluated sample covering 13 domains, attacked through a taxonomy of four attack types and four injection positions.

Hong-Lin Mu, Jinghao Liu, Kai-Yang Wan et al. · 4 citations · ⚡1

Related blog posts

MIT News · Artificial Intelligence Sep 24, 2026

Estimating suicide risk from text

A new language-processing tool could help identify the highest-risk individuals from natural language, enabling swifter interventions.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.