Preprint
Aug 2026
HiRoute: Hierarchical Routed Prompt Tuning for Safety Alignment of Large Language Models
Experiments across three instruction-tuned models show that HiRoute achieves high safety rates across multiple safety benchmarks while preserving safe-response helpfulness, reducing over-refusal, and maintaining competitive performance on general-purpose tasks.
Fangzhou Chen, Shiji Zhao, Mengyan Wang et al.
· 0 citations