Large Vision-Language Models (LVLMs) have shown strong promise for multimodal reasoning, yet often struggle with tasks requiring concepts beyond what is directly observable in the input image. Existing methods generate intermediate images or latent visual tokens to guide reasoning, but these representations can introdu...
Wen-Han Yang, Nilay Naharas, Ali Payani et al.· 0 citations
Universal Activation Verbalizer (UAV), a framework that uses a shared decoder to explain activations from heterogeneous donor models, and provides the activation-grounded factual and semantic information needed for faithful explanations.
Hai-Yan Zhao, Zirui He, Guan-Chun Wang et al.· arXiv.org· 1 citation
D3-Gym is introduced, the first automatically constructed dataset with verifiable environments for scientific Data-Driven Discovery, and how D3-Gym environments can serve as a testbed for studying agentic optimization loops such as Autoresearch on real scientific workflows.
Hanane Nour Moussa, Yi-Fei Li, Zhuo-Yang Li et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.