C2KV is proposed, a unified framework for non-prefix KV reuse that jointly optimizes KV cache compression and concatenation that significantly reduces KV cache storage and transfer costs.
Chuheng Du, Jun-Yi Chen, Hanlin Tang et al.· Proceedings of the 32nd ACM...· 2 citations
Evidence-First Reflection (EFR), a two-stage reflector that explicitly decouples action-induced visual differences extraction from outcome verification, makes reflection better grounded in screen transitions, while reducing both visual search complexity and reasoning burden.
Gleam is a novel and network-efficient framework for task-generic GPU sharing across local-area CUDA devices, with three key contributions that reduce bandwidth overhead in CUDA API remoting through automatic model weight caching, and mitigate accumulated latency from frequent API calls by asynchronous execution.
This work proposes a novel topology named the Balanced Sparse Tree (BST), which is a topology characterized by symmetric design and sparse connections, motivated by hypergraph theory and Steiner Systems, and demonstrates the superiority of BST over the state-of-the-art in network scale, latency, bandwidth, and cost.
Shaoteng Liu, Dejun Kong, Huitian Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.