It is argued that it is difficult to predict the alignment state of a model after training by inspecting the training data alone, which highlights the importance of comprehensive post-training alignment evaluations.
Yavuz Faruk Bakman, D. Yaldiz, Baris Askin et al.· 0 citations
Long-context LLM decoding reads the key-value (KV) cache at every step. Loading it takes longer than computing attention over it, so throughput is bandwidth-bound. Hence, reducing the cache size can raise both decoding speed and serving capacity. The challenge is to reduce cache size while preserving the attention prod...
Samuel Fernández-Menduiña, Amir Ziashahabi, Eduardo Pavez et al.· 0 citations
This work proposes ASPIRE, a non-synchronized batched self-speculative decoding framework built on three components, which achieves speedup in decoding throughput over autoregressive baselines and improves average speedup by approximately $27\% over the strongest prior self-speculative baselines.
Amir Ziashahabi, Hossein Entezari Zarch, Lei Gao et al.· 0 citations
Multi-agent systems (MAS) are enabling increasingly complex, collaborative applications in autonomous driving, smart logistics, robotic coordination, and distributed sensing. Their effectiveness depends on collective intelligence emerging from multiple distributed agents, each operating with partial information and oft...
Hao-Zhao Wang, Zhuangdi Zhu, Zheng Xu et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.