Skip to content
Conference

QXL - A Scalable Hardware Framework for Tabular Q-Learning Inference on RISC-V

Sep 2026 · IEEE International Conference on Application-Specific Systems, Architectures, and Processors · pp. 141-144 · 0 citations · 13 references

Abstract

Tabular Q-learning is a foundational reinforcement learning (RL) algorithm used in embedded decision-making, yet its inference phase suffers from latency-sensitive memory accesses, creating severe bottlenecks in real-time control systems. To address this, we present QXL, a hardware-software codesigned inference accelerator, seamlessly integrated as a custom instruction extension for RISC-V processors. Rather than employing complex, power-hungry dynamic caches, QXL utilizes an on-chip firmware profiling routine to track state visitation. The optimal actions for these high-traffic states are pre-loaded into a lightweight, configurable register scratchpad, bypassing main memory lookups during inference. We evaluate the framework's versatility across three diverse embedded domains: adaptive traffic signal control, stochastic inventory management, and spatial pathfinding. Hardware synthesis on a Nexys FPGA and a 45 nm ASIC demonstrates that QXL's internal datapath decouples from state-space complexity, allowing the logic to scale logarithmically in area. By successfully offloading the decision-making loop, QXL achieves up to a 60% reduction in inference cycle count and a $3.09 \times$ improvement in execution efficiency, establishing a deterministic, low-power architecture for real-time edge RL applications.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.