DARS: Dual-Level Credit Assignment RL with Structured Reasoning for Instruction-Based Image Editing
Instruction-based image editing uses a planner-renderer pipeline: a vision-language model (VLM) first converts the instruction into an edit plan, and a diffusion model then executes that plan. Training such systems with only final-image rewards is inefficient because a poor edit does not reveal whether additional optim...