Value-model-free RLVR methods such as GRPO assign uniform advantages to all tokens in a rollout, ignoring that tokens contribute unequally. Recent methods use token entropy as an importance proxy but compute it globally across the batch, conflating importance with prompt difficulty and positional trends. We argue that...
The results demonstrate the feasibility of harness distillation for robotics: intervention experience can be accumulated, refined through execution, and reused across agents to improve manipulation.
Seungyeon Kim, Junhoo Lee, Minkyu Kim et al.· 0 citations
Evaluation of a video authoring pipeline featuring two layers of structured refusal shows that both layers independently improve the same instructional dimensions, suggesting that thoughtful resistance and generative AI are not opposites but partners.
Yearim Kim, In-Chang Baek, Nojun Kwak· 0 citations
This work proposes LandingAgent, a three-phase agentic framework that profiles the target, constructs a reference-guided wireframe, and refines the page through critique-guided polishing, and evaluates it against direct prompting on faithfulness, conciseness, readability, aesthetics, and structural diversity.
In-Chang Baek, Hyeongseok Lee, Yearim Kim et al.· 0 citations
String is presented, an open-source runtime that gives this new class of software user an interface of its own and treats the job as an operating-systems problem and what three months of production use taught us.