LLM inference systems may vary batch composition, prompt chunking, prefill/decode execution, and KV-cache reuse, eviction, or recomputation. These optimizations should not affect system outputs. Production systems, including vLLM's batch-invariant mode and SGLang's deterministic mode, target this goal but lack a formal...
Jian-Xing Qin, Alexander Du, Dan-Feng Zhang et al.· 0 citations
This formative study establishes a methodological foundation and conceptual framework to guide subsequent validated, large-scale EHR-inconsistency analysis and proposes a graded ontology spanning strict contradiction and ambiguity, with a schema characterizing each flagged case by category, section, domain, and inconsi...
Jiankang Lu, Panyu Chen, Miriam M. Treggiari et al.· arXiv.org· 0 citations
LowRankArena is presented, a standardized evaluation platform for SVD-based LLM compression that unifies task versions, uniform-precision compression budgets, comparison regimes, and inference measurements, and provides a reproducible pipeline with over 3 TiB released compressed checkpoints.
Zi-Shan Shao, Li-Xun Zhang, Kang-Ning Cui et al.· 0 citations
A simple yet effective zero-shot ReAct-style framework built solely on iterative reasoning and a constrained action space defined by a typed Domain-Specific Language (DSL) of 15 relational operations, rather than free-form SQL generation.
Jian Lu, Haiwei Yu, Raymond M. Xiong et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.