KV-streams is proposed, a plug-and-play strategy compatible with any compaction strategy that substantially increases throughput while showing no evidence of hindering performance, and is shown to be an efficient and lightweight plug-and-play addition to any post-training pipeline.
Emiliano Penaloza, Dane Malenfant, Dheeraj Vattikonda et al.· 0 citations
R-Qwen is a recursive reasoning framework built upon a pretrained Qwen backbone, which consistently outperforms prior recursive reasoning models and substantially larger LLMs while using a comparable number of trainable parameters, suggesting that recursive reasoning mechanisms and pretrained language model priors are...
Omid Nejati Manzari, Guillaume Lajoie, H. Rivaz· 0 citations
This work interprets chain-of-thought reasoning as a latent variable modeling problem and demonstrates that this distribution-matching paradigm of LLM fine-tuning can serve as an effective alternative to maximum-likelihood training and reward-maximizing policy optimization.
Edward J. Hu, Moksh Jain, Eric Elmoznino et al.· International Conference on...· 110 citations· ⚡19
A new algorithm for amortized inference in sparse probabilistic graphical models (PGMs) is presented that enables off-policy training but avoids the need to instantiate all the random variables for each parameter update, thus speeding up training considerably.
J. Falet, Haebeom Lee, Esmeralda S. Whitammer et al.· International Conference on...· 9 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.