ThinkBLOX: 3D Indoor Scene Generation with Progressive Reasoning
This work proposes ThinkBLOX, a VLM-based progressive reasoning framework that iteratively designs and refines 3D scenes and introduces Tier-Decoupled GDPO, a reinforcement learning scheme that organizes heterogeneous rewards into distinct tiers, stabilizing policy optimization across physical validity, semantic plausibility, and reasoning-action consistency.