XBRIDGE is proposed, a decode-free communication protocol that outperforms text-based communication on all seven tasks for each model pair while achieving 11x lower latency, and in a same-architecture setting it also exceeds a KV-sharing baseline on six of seven tasks.
Wooseong Yang, Wei-Chieh Huang, Wei-Zhi Zhang et al.· 3 citations
An agent-based recommendation framework, memory-based Personalized Recommendation Tool learning via autonomous language Agents (PRTA), in which an LLM acts as a central planner interacting with multiple recommendation models as tools, demonstrating the superiority of PRTA over traditional recommendation and LLM-based b...
Ming-Dai Yang, Zhi-Wei Liu, Wei-Zhi Zhang et al.· Proceedings of the 20th ACM...· 1 citation
Random Attention keeps the prompt and evicts uniformly at random within each attention head, computing no score at all; across four models and six reasoning tasks, it matches the strongest baseline in task performance while delivering 32-43% higher throughput than that method when deployed with vLLM.
Heng-Yi Wang, Jie-Lin Qiu, Wenting Zhao et al.· 0 citations
BlockServe is presented, a continuous batching framework that integrates block-grained scheduling -- immediately evicting completed requests at block boundaries -- with mixed-state execution that extends dual cache and parallel decoding to heterogeneous batches via gather-scatter indexing.
Yuanjie Zhu, Liangwei Yang, Ke Xu et al.· arXiv.org· 0 citations
A controlled harness evaluation of memory substrates for memory-augmented agents, covering dense and sparse indices, text records, structural stores, hierarchical stores, refinement-based memories, parametric updates, and activation-compatible context mechanisms, shows that no single substrate consistently dominates.
Wei-Chieh Huang, Wei-Zhi Zhang, Yu-Chen Wu et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.