The key-value (KV) cache stores information from past tokens and is a major memory bottleneck in long-context inference. Existing quantization methods address this bottleneck by representing the KV cache uniformly with lower-precision data types and designing quantization schemes to minimize reconstruction error in the...
Hannah Laus, C. M. Verdun, Hao Wang et al.· 1 citation
LOGIC transforms the exploration paradigm: instead of exploring at the action level, it employs an integrated critic to guide a generator in discovering optimal, high-value goals (RTGs), enabling robust exploration beyond the dataset's frontier while avoiding the fragility of complex multi-loss objectives.
Shi-Hao Shu, Rujie Zhong, Hao Wang et al.· Proceedings of the 32nd ACM...· 0 citations
Decision Transformers (DTs) have emerged as a powerful paradigm for sequential decision-making in offline learning, yet they face two intrinsic deficiencies: the challenge of specifying optimal Return-to-Go (RTG) targets and a fundamental lack of exploration beyond the offline dataset. While prior studies have attempte...
Shihao Shu, Rujie Zhong, Hao Wang et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.