Skip to content
Conference

Reinforcement Learning for Adaptive Market-Making with Inventory-Aware Risk Constraints

Aug 2026 · 2026 International Conference on Secure Information Systems and Technologies (ICSIST) · pp. 1262-1269 · 0 citations · 26 references

Abstract

It is essential for market-making strategies to optimally capture the spread, balance their inventory, and adjust to changing order-book dynamics. In this paper, we explore the use of Proximal Policy Optimization (PPO) to perform adaptive market making in a simulated limit-order-book setting using the MSFT data from the LOBSTER dataset, with 668,765 snapshots and 10 levels of market depth. The state is represented as fourteen normalised microstructure features, such as order imbalance, order flow imbalance, depth imbalance, rolling volatility, and normalized inventory. The agent chooses between three different quoting actions: widening, maintaining, or narrowing the spread, and receives rewards through a combination of mark-to-market profit/loss, spread capture, a quadratic inventory penalty, and risk penalties. Given the execution rules assumed by the simulator, PPO with an inventory penalty obtains a Sharpe ratio of 250.898 with a test-set PnL of $85,020.00, while PPO without an inventory penalty obtains a test-set PnL of $114,681.75 at a Sharpe ratio of 244.601. Through ablation experiments, the inventory penalty slightly increases the Sharpe ratio but decreases absolute profitability. All results are based on the simulated market environment and cannot be extrapolated to real-market environments, where actual performance cannot be guaranteed.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.