Skip to content

Category

artificial intelligence

6,497 papers

#artificial intelligence Preprint Aug 2026

Q-Strata: Hierarchical Bit Allocation for Mixed-Precision Quantization of Mixture-of-Experts LLMs

Q-Strata is proposed, a bi-level allocator that ranks within-block assignments with a cheap proxy and allocates across blocks with a model-level objective evaluated on the assembled quantized model, achieving lower WikiText2 perplexity than uniform-bitwidth GPTQ and the state-of-the-art MoE MPQ methods MxMoE and GEMQ in the low-bit regime.

Deokjae Lee, Si-Hun Chu, Hyun Oh Song · 0 citations
#artificial intelligence Preprint Aug 2026

Trajectory-Initialized Neural Double Q-Routing for Large-Scale Overhead Hoist Transport Systems

This work proposes Neural Double Q-routing, which replaces destination-indexed tables with a shared state--action value network, and achieves the lowest mean completion time among all compared methods in the six 150- and 200-OHT settings, whereas Dijkstra remains best in the three 100-OHT settings.

Chen-Feng Gu, Qiu-Sheng Zhao, An-Bang Liu et al. · 0 citations
#artificial intelligence Review Aug 2026

Tensor Methods for Language Models: From Token Representation to Training, Adaptation, Inference, Compression, and Interpretability

This survey provides a structured entry point to tensorized language models and clarifies when parameter savings can plausibly translate into memory efficiency, computational efficiency, or interpretability, and introduces a metric for the compression-realization gap between theoretical memory reduction and measured system-level speedup.

M. Tarasov, Salman Ahmadi-Asl, A. D. de Almeida et al. · 0 citations
#artificial intelligence Preprint Aug 2026

TopGQ: Fast GNN Post-Training Quantization Leveraging Topology Information

TopGQ, an accurate post-training GNN quantization framework, alleviating redundant quantization overhead, is presented, and dual-axis scale absorption is proposed, which enables activation quantization along both the outer and inner dimensions by merging one into the adjacency matrix.

Dain Kwon, Kanghyun Choi, Hyeyoon Lee et al. · 0 citations
#artificial intelligence Preprint Aug 2026

DASC: Decay-Aware State Compression for Hybrid Linear-Attention Serving

Decay-Aware State Compression (DASC), which derives retention horizons from model weights, selects long-horizon state units, and packs them into a ragged state checkpoint layout to integrate efficiently with tensor-parallel inference engines.

Yanzhi Yu, Ping-Wei Sun, Jian-Chao Tan et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Online Estimation of Dynamic Origin-Destination Matrices Using Reinforcement Learning with Link-Flow Propagation Guidance

LFPG-RL is developed and evaluated, which integrates link-flow propagation guidance (LFPG) into proximal policy optimization (PPO), and results support the contention that the method is a more efficient and accurate online OD demand calibration method compared to existing ones.

Donggyu Min, Dong-Kyu Kim · 0 citations
#artificial intelligence Conference Aug 2026

CateKV: On Sequential Consistency for Long-Context LLM Inference Acceleration

This work discovers that certain attention heads exhibit sequential consistency in their attention patterns, which can be persistently identified using a coefficient-of-variation-based algorithm, and proposes CateKV, a hybrid KV cache method that retains only critical token information for consistent heads, thereby reducing KV cache size and computational overhead.

Hao-Yun Jiang, Hao-Lin Li, Jian-Wei Zhang et al. · 2 citations
#artificial intelligence Preprint Aug 2026

BCPPO: Bachelier-Inspired Constrained Proximal Policy Optimization for Tail-Risk-Aware Safe Reinforcement Learning

BCPPO (Bachelier-Inspired Constrained Proximal Policy Optimization), a proximal policy optimization (PPO) method, supports a practical balance among reward, caution around cost predictions that vary across trained critics, and policy-only deployment.

Dong-Sheng Hou, Yanqiao Chen, Yu-Han Rui · 0 citations
#artificial intelligence Preprint Aug 2026

TPR-Attention for Combinatorial Generalization

This work introduces a new architectural component that embeds structured inductive bias into deep learning: an attention mechanism operating over tensor-product representations (TPRs) that outperforms existing architectural components in combinatorial generalization.

Melisa Civelekoğlu, Isabeau Prémont-Schwarz · 0 citations
#artificial intelligence Preprint Aug 2026

Graph4BiLO: Graph Neural Network Approximation for Bilevel Mixed-Integer Linear Optimization

Graph4BiLO is introduced, a graph neural network (GNN) approach for learning bilevel value functions from variable--constraint graph representations that obtains objective values comparable to Neur2BiLO across all tested sizes while avoiding size-specific neural networks.

Jessica D. Elrefaei, Kaixun Hua, Seungbae Kim et al. · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.