TopGQ, an accurate post-training GNN quantization framework, alleviating redundant quantization overhead, is presented, and dual-axis scale absorption is proposed, which enables activation quantization along both the outer and inner dimensions by merging one into the adjacency matrix.
Dain Kwon, Kanghyun Choi, Hyeyoon Lee et al.· 0 citations
Decay-Aware State Compression (DASC), which derives retention horizons from model weights, selects long-horizon state units, and packs them into a ragged state checkpoint layout to integrate efficiently with tensor-parallel inference engines.
Yanzhi Yu, Ping-Wei Sun, Jian-Chao Tan et al.· 0 citations
LFPG-RL is developed and evaluated, which integrates link-flow propagation guidance (LFPG) into proximal policy optimization (PPO), and results support the contention that the method is a more efficient and accurate online OD demand calibration method compared to existing ones.
Tail-Replay is presented, a prefix caching mechanism that enables unconstrained token-level prefix reuse in hybrid large language models and is evaluated on three Gated DeltaNet-based hybrid models using the LongBench and RULER benchmarks.
Yi-Rui Liu, Ruoling Qi, Xuan'er Wu et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
This work discovers that certain attention heads exhibit sequential consistency in their attention patterns, which can be persistently identified using a coefficient-of-variation-based algorithm, and proposes CateKV, a hybrid KV cache method that retains only critical token information for consistent heads, thereby reducing KV cache size and computational overhead.
Hao-Yun Jiang, Hao-Lin Li, Jian-Wei Zhang et al.· International Conference on...· 2 citations
BCPPO (Bachelier-Inspired Constrained Proximal Policy Optimization), a proximal policy optimization (PPO) method, supports a practical balance among reward, caution around cost predictions that vary across trained critics, and policy-only deployment.
This work introduces a new architectural component that embeds structured inductive bias into deep learning: an attention mechanism operating over tensor-product representations (TPRs) that outperforms existing architectural components in combinatorial generalization.
Graph4BiLO is introduced, a graph neural network (GNN) approach for learning bilevel value functions from variable--constraint graph representations that obtains objective values comparable to Neur2BiLO across all tested sizes while avoiding size-specific neural networks.
Jessica D. Elrefaei, Kaixun Hua, Seungbae Kim et al.· 0 citations
The findings suggest world knowledge and task-directed ability can be learned in geometrically complementary forms, and that future post-training pipelines should consider how best to engineer the interface between them.
Rui-Ze Xu, Xiao Yu, Yuxin Tang et al.· 0 citations
This work conducts a comparative empirical study of five MU methods across symmetric, asymmetric, instance-dependent, and open-set noise on CIFAR-10, CIFAR-100, and the real-world noisy dataset Food-101N and finds that the appropriate unlearning strategy is conditioned on the noise structure.
It is suggested that domain-specific training matters more than model scale for PET/CT report error detection, supporting compact models as an accurate and computationally efficient approach to automated radiology report quality assurance.
Hermione Warr, Harry Anthony, Lilli J. Freischem et al.· 0 citations
INTERVenE is presented, a family of Transformer architectures whose input is an interval-based, knowledge-based temporal abstraction (KBTA), a token stream of named clinical concepts drawn from a curated medical ontology, rather than an unnamed bin index or a raw measurement triplet.