In real-world applications, current multimodal large models are often overestimated in their ability to understand scientific charts. To assess their true capabilities and identify key performance bottlenecks, we conducted an in-depth study on scientific chart understanding. Charts in scientific literature often featur...
Ling-Dong Shen, Qigqi, Kun Ding et al.· IEEE Transactions on Image P...· 0 citations
Diagnostics reveal that RL on PLMs is governed by two reward properties: verifiability, whether the reward is a fixed environment or a learned surrogate vulnerable to distribution shift, and coverage, the fraction of sequence space giving an informative gradient.
Hanqun Cao, Hongrui Zhang, Junde Xu et al.· Proceedings of the 32nd ACM...· 0 citations
Reinforcement learning (RL) is increasingly applied to Protein Language Models (PLMs), yet its effectiveness varies across tasks, and standard metrics such as pass@k can rise even when the model's solvable problem set is shrinking. We introduce two capability-level diagnostics. The Expansion-Shrinkage Ratio (ESR) measu...
Hanqun Cao, Hongrui Zhang, Junde Xu et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.