Volatility-Consistent Reward Shaping for Deep Reinforcement Learning Market Makers
Abstract
Deep reinforcement learning (DRL) is increasingly utilized for optimal execution and market making. However, standard DRL formulations typically rely on static inventory penalties to control risk. In this paper, we observe that applying a static inventory penalty may induce a volatility-inconsistent implicit risk preference. Specifically, as market variance increases during regime shifts, a static penalty inherently lowers the agent’s effective risk aversion, precisely when inventory exposure should be more tightly constrained. To address this, we provide a theoretical motivation rooted in the Hamilton-Jacobi-Bellman (HJB) optimal control framework to derive a volatility-consistent reward shaping mechanism. By mapping the continuous-time certainty equivalent of a Constant Absolute Risk Aversion (CARA) utility to the discrete-time Bellman objective, we demonstrate that dynamically scaling the inventory penalty with the instantaneous variance ensures coherent risk preferences across shifting market regimes.We evaluate this mechanism using an actor-critic architecture on granular limit order book (LOB) data. To ensure empirical rigor, the agent is backtested using a queue-reactive matching simulator across multiple historical macro-volatility events, including the FTX collapse, the SVB crisis, and the SEC-Binance lawsuit. We extensively compare our proposed method against static RL baselines and the classical Avellaneda-Stoikov model. Furthermore, we conduct comprehensive ablation studies on state representation and network architecture. Empirical results, aggregated over multiple random initializations, suggest that our volatility-consistent agent significantly reduces inventory concentration, accelerates convergence, and improves drawdown stability during stress periods. Our findings suggest that volatility-consistent reward shaping provides a more coherent, theoretically grounded, and highly robust framework for inventory-aware RL market making under non-stationary conditions.