Skip to content

I/O Lower-Bound Theory and Reinforcement Learning for Efficient Neural Network Inference Optimization

Sep 2026 · ACM Transactions on Architecture and Code Optimization (TACO) · 0 citations · 24 references
Advanced Neural Network Applications

TL;DR

An automated optimization framework with a hierarchical two-layer tuning mechanism that synergizes theoretical I/O constraints with graph-level adaptive fusion while accounting for search overhead, the framework systematically explores high-performance execution patterns.

Abstract

As artificial intelligence models grow in complexity, optimizing neural network inference has become a critical challenge. Existing approaches often rely on manual, expert-driven tuning tailored to specific hardware, which lacks scalability across diverse architectures. In this paper, we propose an automated optimization framework with a hierarchical two-layer tuning mechanism. At the node level (intra-operator), we introduce an I/O lower-bound theory based on the Red-Blue Pebble game and the (X1, X2)-Partition theorem to guide tiling and memory-mapping configurations. At the graph level (inter-operator), we employ a reinforcement learning (RL) strategy to adaptively identify optimal operator fusion boundaries across network topologies. By synergizing theoretical I/O constraints with graph-level adaptive fusion while accounting for search overhead, the framework systematically explores high-performance execution patterns. For the TileAttn operator, the (X1, X2)-Partition theorem raises DRAM flow estimation accuracy from 77.5%-82.0% under X-Partition to 86.3%-94.8%. Our method reduces shared memory traffic by 10.79% on average, achieves the best performance in 65.3% of cross-platform cases and top-two in 86.1%, and its DQN-based fusion engine outperforms greedy strategies in 91.67% of scenarios. We further analyze the algorithm’s overhead and its amortization break-even points.

Read PDF

Similar papers

Preprint Aug 2026

Momba: Network Modernization Improves Multi-Objective Reinforcement Learning

Recent advances in neural network design are integrated: observation and feature normalization, weight normalization, and modeling of distributional returns with an entropy-regularized MORL algorithm, demonstrating that these changes substantially improve the quality of the produced solution sets without requiring majo...

Adam Štafa, Santeri Heiskanen, Petr Novotný et al. · 0 citations
Preprint Aug 2026

EvoRIC: Reinforcement Learning Fine-Tuned LLM-empowered RAN Intelligent Control Toward Autonomous O-RAN

The evolving RAN intelligent controller (RIC) (EvoRIC) framework is introduced, a hierarchical architecture that enables continuous evolution by leveraging a non-real-time RIC (non-RT RIC) for global model updates and a near-real-time RIC (near-RT RIC) for local execution, dynamically empowering LLMs with domain-specif...

Lingyan Bao, Jemin Lee, Tony Q. S. Quek · 0 citations
#artificial intelligence Preprint Sep 2026

Learning-to-Optimize as the Missing Architectural Layer of AI-Native Networks

This paper argues that the Learning-to-Optimize (L2O) represents the missing architectural layer between optimisation and AI-native intelligence, and establishes L2O as an architectural abstraction applicable across heterogeneous communication and computing systems.

G. Amati, Federica Mangiatordi, P. Salvo et al. · 0 citations
#machine learning Preprint Sep 2026

Verifying Neural Networks with Reinforcement Learning

Formal verification can play a key role in ensuring the reliability of Deep Neural Networks (DNNs) deployed in safety-critical systems. Modern DNN verifiers employ a branch-and-bound framework, which alternates between branching (splitting into smaller subproblems) and bounding (pruning subproblems) to efficiently expl...

Hai Duong, Thanh Le, Thanhvu Nguyen · 5 citations
#natural language process... Preprint Sep 2026

Expert-Space Exploration in MoE Reinforcement Learning

Reinforcement learning (RL) has become central to post-training of large language models. Recent advances in RL for Mixture-of-Experts (MoE) models have primarily focused on improving optimization stability and training efficiency, while treating the expert selection as a fixed component. Since routing determines the s...

Hong-Yi He, Zheng-Wen Lin, Xiao Liu et al. · 0 citations

Related blog posts

Microsoft Research Blog Jul 13, 2026

Verifying Rust cryptography in SymCrypt, from standards to code

Cryptographic code supports vital protections in modern computing systems. Learn how a new method helps verify code as developers write it while preserving speed and adaptability as it gets implemented and evolves. The post Verifying Rust cryptography in SymCrypt, from standards to code appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.