Multimodal Thinking with Renderable Programs
This work introduces SVGLM, a framework that uses scalable vector graphics (SVG) primitives to connect text and image in reasoning tasks, and exploits the duality of SVG as both image description and text instructions, yielding a more compact, interpretable solution to equip general VLMs with the capability of generati...