Disaggregated LLM serving separates prefill and decode into distinct node pools, interposing a network fabric between the moment a key-value (KV) cache is computed and the moment it is consumed. This architectural shift invalidates a core assumption of classical cache policies: that the cost of a miss is simply recompu...
Dong Liu, Yan-Xuan Yu, Eric Jiang et al.· Proceedings of the 19th ACM...· 0 citations
Codebook Agent is the most accurate method on all six benchmarks the authors compare, and an MLP proxy that reads the flattened adjacency, regressed on measured utility and per-task normalized token cost, reranks the top decoded candidates in a single batched forward pass.
Jin-Xi Yu, Yubei Li, Eric Jiang et al.· 1 citation
Juris Policy Optimization (JPO), a post-training framework for structured legal reasoning in Chinese criminal judgment prediction, is proposed and experiments show that JPO consistently improves both judgment prediction and reasoning quality over supervised fine-tuning and reinforcement learning baselines.
Zhao-Lu Kang, Yan-Tao Liu, Tailong Luo et al.· 1 citation
This work presents Memory as a Controlled Process (MemCon), a framework that models memory operations as a Markov Decision Process and learns an online policy that adaptively decides when, what, and how much to retrieve, when to inject a distilled plan, and when to consolidate or forget.
Eric Jiang, Zhi Zhang, Yuchen Wu et al.· arXiv.org· 1 citation
It is argued that the next leap in AI4Math systems requires a decisive shift from predefined problem-solvers to research agents that can address frontier mathematical challenges with rigorous formal mathematical reasoning, highlighting core limitations of existing systems in serving as mathematical research agents.
E. Jiang, Xiao Liang, Yikai Zhang et al.· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.