Best-of-$N$ is a widely used inference strategy for complex reasoning, whose effectiveness depends on whether sampled candidates can cover diverse and high-quality reasoning paths. However, post-trained reasoning models often suffer from \emph{exploration collapse}, where independent rollouts repeatedly follow similar...
Hengyuan Zhang, Chenming Shang, Zunhai Su et al.· 0 citations
Inspired by dual-coding theory, this work proposes a memory architecture that uses parallel visual and verbal codes, which it calls DualMem, and views this as a step towards memory systems that preserve a richer record of agents'observations.
The findings suggest world knowledge and task-directed ability can be learned in geometrically complementary forms, and that future post-training pipelines should consider how best to engineer the interface between them.
Rui-Ze Xu, Xiao Yu, Yu-Jin Tang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.