Skip to content

Fine-tuning as Jailbreaking: A data-centric red teaming framework via logic injection

Aug 2026 · International Conference on Automated Software Engineering · Vol 33 · 0 citations · 31 references

TL;DR

A red-teaming testing method for fine-tuning-stage defenses that audit the harmful content of poisoned fine-tuning data, exposing a vulnerability of the RFT data supply chain to logic injection and point to the need for fine-tuning-stage defenses that audit the harmful content of poisoned fine-tuning data.

View source

Similar papers

Conference Open access 2026

LogSanitizer: Defending LLM-Integrated SOCs against Backdoor Triggers Delivered through Firewall Logs

LogSanitizer is proposed, a family of input sanitization defenses operating at two levels: a pre-prompt log-transformation pipeline that disrupts trigger patterns in the structured log representation, and a post-tokenizer perturbation strategy that corrupts trigger-bearing token configurations before they reach the mod...

Leszek Wronski, Bogdan Ksiezopolski · 0 citations
Jul 2026

On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment

This work proposes Routing-based On-Policy Distillation (ROPD), a novel realignment framework that models the divergence between aligned and compromised output probability distributions rather than fitting specific prompt templates, establishing a new standard for robust LLM realignment.

Yong-Jian Guo, Wanlun Ma, Lingyu Shen et al. · 0 citations
#small language model Preprint Aug 2026

RTLGuard: A Lightweight Teacher-Student Defense for Poisoned RTL Code Generation Models

RTLGuard leverages a teacher-student framework designed to sanitize compromised RTL generation models by fine-tuning a small-scale,"clean"teacher model on a limited set of trusted RTL data, and incorporating feature alignment and knowledge distillation to suppress malicious behaviors.

Mahshid Rezakhani, K. Azar, H. Kamali · 0 citations
Book Open access Aug 2026

The 2nd SeT-LLM Workshop on Secure and Trustworthy Large Language Models

Large language models (LLMs) are increasingly embedded as core components of data-centric systems, supporting analytical decision making, and automated reasoning over large-scale, heterogeneous datasets. Yet their deployment in open-world environments raises fundamental challenges to security and trustworthiness: LLMs...

Lu Lin, Jinghui Chen, Ting Wang et al. · 0 citations
Jul 2026

Just Testing, Move Along: Evasion of LLM-based System Log Interpretation by Prompt Injection

This paper presents a framework for evaluating prompt injection attacks against LLM-based log interpretation using log traces generated during real cyber attacks, and creates adversarial examples through generic injection generation, refinement, and attack-specific optimization.

Max Landauer, F. Skopik, Markus Wurzenberger et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.