With the rapid development of multimodal large language models (MLLMs) and the cost of large amounts of data and resources, researchers are more inclined to use public datasets and fine-tune open-source MLLMs to achieve excellent performance in audiovisual related tasks. While this trend accelerates progress, it also i...
Jin-Ming Wen, Xin-Yi Wu, Yuwen Li et al.· ACM Computing Surveys· 0 citations
The rapid progression of large language models is extending AI from passive content generation into the active workflows of engineering and scientific discovery. This shift raises a compelling question: can AI be both the object of development and an active participant in building next-generation AI systems? We explore...
Tingyu Qu, Wei-Gao Sun, Yuecheng Liu et al.· 0 citations
Long-horizon agentic tasks require an agent to modify an environment through a sequence of tool calls, with success determined by the final state. The standard recipe assigns a single outcome reward at the end and compares trajectories sampled for the same task. As a result, a group with no successful trajectory yields...
Ming Ma, Yi Zhu, Yi-Ran Zhong et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.