Reinforcement Learning for Adaptive Market-Making with Inventory-Aware Risk Constraints
Abstract
It is essential for market-making strategies to optimally capture the spread, balance their inventory, and adjust to changing order-book dynamics. In this paper, we explore the use of Proximal Policy Optimization (PPO) to perform adaptive market making in a simulated limit-order-book setting using the MSFT data from the LOBSTER dataset, with 668,765 snapshots and 10 levels of market depth. The state is represented as fourteen normalised microstructure features, such as order imbalance, order flow imbalance, depth imbalance, rolling volatility, and normalized inventory. The agent chooses between three different quoting actions: widening, maintaining, or narrowing the spread, and receives rewards through a combination of mark-to-market profit/loss, spread capture, a quadratic inventory penalty, and risk penalties. Given the execution rules assumed by the simulator, PPO with an inventory penalty obtains a Sharpe ratio of 250.898 with a test-set PnL of $85,020.00, while PPO without an inventory penalty obtains a test-set PnL of $114,681.75 at a Sharpe ratio of 244.601. Through ablation experiments, the inventory penalty slightly increases the Sharpe ratio but decreases absolute profitability. All results are based on the simulated market environment and cannot be extrapolated to real-market environments, where actual performance cannot be guaranteed.