Fine-tuning adapts aligned large language models (LLMs) to downstream tasks, but a few dozen harmful examples can remove their refusal of harmful requests. Prior work localizes safety-related behavior to specific layers, directions, and tokens, suggesting targets for protection. We test whether successful localization...
Jungseob Lee, D. Lee, Sugyeong Eo et al.· 1 citation
PriceCheck is introduced, which builds a compact family of decision rules from label-free checks such as re-solving a problem, and shows that choosing how checks are combined and stopped matters alongside how well a verifier ranks answers.
D. Lee, Jungseob Lee, Chanjun Park et al.· 2 citations
Redline is a finite-sample procedure that selects operating points, hand-picked or learned, from the correctness of their answers on calibration prompts, and keeps the reference-relative risk, the joint probability that the reference answers correctly and a candidate configuration does not, within a user-chosen budget...
Jungseob Lee, D. Lee, Chanjun Park et al.· 1 citation
GLANCE is presented, the first one-pass block drafter that is lossless on an unmodified VLM target, and it breaks the cycle at both ends of a self-defeating cycle.
Jungseob Lee, Seongtae Hong, D. Lee et al.· 0 citations
This work shows that training a strong instruction-tuned reasoning model on its own answer-conditioned chains sharply lowers its verifiable-reasoning accuracy, and generates answer-blind data, because no correctness filter can see this damage in the data.
Jungseob Lee, Seungyoon Lee, Suhyune Son et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.