As large language model (LLM) inference becomes increasingly expensive, resource-consumption attacks pose a growing threat to model providers. Existing attacks typically amplify cost by inducing abnormally long or repetitive outputs on attacker-controlled or triggered requests, making them easier to detect and limiting...
Zi-Han Wang, Rui Zhang, Xin-Yuan Qian et al.· 0 citations
The transport layer is a critical component of the Internet protocol stack, providing reliable data delivery and securing the communication channel between clients and servers. TLS 1.3, the de facto standard protocol for this layer, is deployed on billions of devices and provides the confidentiality and integrity of mo...
Xin-Yue He, Yuan Zhang, Guo-Min Yang et al.· IEEE Transactions on Informa...· 0 citations
Expert parallelism (EP) is a common strategy for serving large Mixture-of-Experts (MoE) models across multiple GPUs by distributing experts among devices. Router decisions then determine both which experts process each token and which GPUs execute the resulting work. This procedure exposes a supply-chain attack surface...
Rui Zhang, Wenbo Jiang, Hong-Wei Li et al.· 1 citation
Diffusion-based text-to-image (T2I) models are increasingly used for visual content creation, making their generation capability a valuable intellectual property asset. However, this capability is vulnerable to black-box output-based distillation, where an adversary queries the service, collects prompt-image pairs, and...
Zi-Han Wang, Bo-Heng Li, Rui Zhang et al.· 0 citations
Retrieval-Augmented Generation (RAG) improves the factuality of large language models with external knowledge, yet conflicting evidence remains a fundamental challenge in dynamic and adversarial environments. Existing approaches often treat conflicts as static inconsistencies and select more reliable knowledge, overloo...
Xiao-Rui Nie, Hong-Wei Li, Shenghao Wu et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.