Retrieval-Constrained Policy Optimization for Attack Technique Extraction from Cyber Threat Intelligence
TTP-R1 is proposed, a two-stage framework that combines retrieval-augmented supervised fine-tuning (SFT) with reinforcement learning using verifiable rewards (RLVR) and applies Group Relative Policy Optimization with a decomposed reward that directly supervises the precision, recall, and output format of the predicted...