Jul 2026· Proceedings of the 2026 International Symposium on AI-Enabled Information and Communication Systems· pp. 70-74· 0 citations· 11 references
TL;DR
DE-CPPO is introduced and represents transaction, conversion, inventory, competitor-price and macroeconomic sequences as components of a Markov state and shows that it can represent observed retail-demand dynamics.
Abstract
Consumer-market prices, demand, inventory and competitive signals change simultaneously; therefore, revenue-only pricing can lead to unstable adjustments, delayed purchases and reduced access for price-sensitive consumers. This paper introduces DE-CPPO and represents transaction, conversion, inventory, competitor-price and macroeconomic sequences as components of a Markov state. Dilated temporal encoding and gated fusion support Proximal Policy Optimisation (PPO), while safety projection limits the price level, adjustment rate and segment disparity. Under controlled retail conditions, DE-CPPO increased revenue by 13.6%, raised the demand-expansion index to 1.154, improved coverage to 82.4% and reduced stockouts to 3.1%. Ablation and stress tests show that temporal encoding, the expansion-oriented reward and safety projection all contribute to performance. A public-data transfer test involving 811 products over 52 weeks yielded a normalised mean absolute error of 0.199 and a root mean square error of 0.265 for the temporal encoder, providing further evidence that it can represent observed retail-demand dynamics.
Dynamic pricing in ride-sharing platforms must balance revenue generation with stable pricing decisions under changing demand and supply. This study aims to develop and evaluate a reinforcement learning-based dynamic pricing policy that maximizes expected revenue while reducing abrupt policy-level price adjustments. A...
Nur Alamsyah, Budiman, Almira Nurchawilah et al.· JITK (Jurnal Ilmu Pengetahua...· 0 citations
The findings reveal a clear methodological shift from rule-based and econometric approaches toward deep learning, multi-agent reinforcement learning, and simulation-driven decision systems, and show that data-intensive and platform-mediated sectors are becoming increasingly prominent in the development and application...
Dervis Ozay, M. Jahanbakht, Shouyi Wang· Journal of Theoretical and A...· 0 citations
This paper studies the algorithmic design of price competition in oligopolistic markets, with a focus on long-run market dynamics under reference price effects. We consider a sequential price competition framework with multiple sellers operating over a finite horizon, each lacking prior knowledge of the demand functi...
Yong-Ge Yang, Jian-Nan Ke, Cong Shi· Production and operations ma...· 0 citations
Automated market makers (AMMs) are a cornerstone of decentralised finance (DeFi). Constant product markets with concentrated liquidity, such as UniswapV3, are now a well-established design. In these markets, liquidity providers (LPs) face a sequential decision problem: they must decide when to rebalance their positions...
Georgios Chionas, Charalampos Kleitsikas, Stefanos Leonardos et al.· 0 citations
It is essential for market-making strategies to optimally capture the spread, balance their inventory, and adjust to changing order-book dynamics. In this paper, we explore the use of Proximal Policy Optimization (PPO) to perform adaptive market making in a simulated limit-order-book setting using the MSFT data from th...
Zubair Raja Ahamed Zaffarullah, Priyameet Kaur Keer, Rubani Singla et al.· 2026 International Conferenc...· 0 citations