REINFORCEMENT LEARNING-BASED DYNAMIC PRICING IN A STOCHASTIC DEMAND–SUPPLY ENVIRONMENT

Authors

  • Nur Alamsyah Universitas Informatika dan Bisnis Indonesia image/svg+xml
  • Budiman Universitas Informatika Dan Bisnis Indonesia
  • Almira Nurchawilah Universitas Informatika Dan Bisnis Indonesia
  • Wala Erpurini Universitas Jenderal Ahmad Yani

DOI:

https://doi.org/10.33480/jitk.v12i1.8286

Keywords:

Dynamic Pricing, Multiplier Stability, Proximal Policy Optimization, Reinforcement Learning, Revenue Optimization

Abstract

Dynamic pricing in ride-sharing platforms must balance revenue generation with stable pricing decisions under changing demand and supply. This study aims to develop and evaluate a reinforcement learning-based dynamic pricing policy that maximizes expected revenue while reducing abrupt policy-level price adjustments. A stochastic contextual environment was constructed from 1,000 historical ride records and evaluated using a leakage-safe 70/15/15 train-validation-test split. The agent was trained with Proximal Policy Optimization (PPO) using five discrete price adjustments from -10% to +10%. Expected revenue was combined with a multiplier-based stability penalty, where stability was measured from changes in the price multiplier rather than nominal price variation across heterogeneous rides. Across 30 paired test episodes, the PPO policy achieved a cumulative reward of 103,316.25 +/- 3,243.99 and expected revenue of 103,449.18 +/- 3,241.27, significantly exceeding static pricing (p < 0.001). Relative to rule-based surge pricing, PPO produced statistically indistinguishable cumulative reward (p = 0.808) while reducing multiplier volatility by 24.83%, mean absolute multiplier change by 24.21%, and action switch rate by 10.97% (all p < 0.001). These results indicate that PPO can preserve near-surge revenue while producing smoother dynamic pricing decisions within the simulated environment.

Downloads

Download data is not yet available.

References

[1] K. Haneefa and H. Singh, “Does AI matter in marketing? Unveiling the role of AI-driven marketing in quick grocery,” Int. Rev. Retail Distrib. Consum. Res., pp. 1–23, 2025, doi: https://doi.org/10.1080/09593969.2025.2526482.

[2] J. Nyangon, “Smart Grid Strategies for Tackling the Duck Curve: A Qualitative Assessment of Digitalization, Battery Energy Storage, and Managed Rebound Effects Benefits,” Energies, vol. 18, no. 15, p. 3988, 2025, doi: https://doi.org/10.3390/en18153988.

[3] N. Alamsyah, A. P. Kurniati, and others, “Airfare Fluctuation Analysis with Event and Sentiment Features by Stacking Ensemble Model,” in 2024 Ninth International Conference on Informatics and Computing (ICIC), IEEE, 2024, pp. 1–6. doi: 10.1109/ICIC64337.2024.10957538.

[4] L. Ming, T. Tunca, Y. Xu, and W. Zhu, “Market Formation, Pricing, and Value Generation in Ride-Hailing Services,” Manuf. Serv. Oper. Manag., vol. 27, no. 5, pp. 1551–1570, 2025, doi: https://doi.org/10.1287/msom.2022.0502.

[5] J. Li and B. Chen, “A Deep Q-Learning Optimization Framework for Dynamic Pricing in E-Commerce,” in Proceedings of the 2025 4th International Conference on Cyber Security, Artificial Intelligence and the Digital Economy, 2025, pp. 367–371. doi: https://doi.org/10.1145/3729706.37297.

[6] J. Y. Yang, “The Future of AI-Enabled Pricing,” in Reimagine Pricing: How AI is Changing Everything, Springer, 2025, pp. 131–142. doi: https://doi.org/10.1007/978-3-031-90418-9_5.

[7] N. Alamsyah, A. P. Kurniati, and others, “Event Detection Optimization Through Stacking Ensemble and BERT Fine-Tuning for Dynamic Pricing of Airline Tickets,” IEEE Access, 2024, doi: 10.1109/ACCESS.2024.3466270.

[8] Y. Kayikci, S. Demir, S. K. Mangla, N. Subramanian, and B. Koc, “Data-driven optimal dynamic pricing strategy for reducing perishable food waste at retailers,” J. Clean. Prod., vol. 344, p. 131068, 2022, doi: https://doi.org/10.1016/j.jclepro.2022.131068.

[9] T. A. Syed, H. Aslam, Z. A. Bhatti, F. Mehmood, and A. Pahuja, “Dynamic pricing for perishable goods: A data-driven digital transformation approach,” Int. J. Prod. Econ., vol. 277, p. 109405, 2024, doi: https://doi.org/10.1016/j.ijpe.2024.109405.

[10] R. Y. Chenavaz and S. Dimitrov, “Artificial intelligence and dynamic pricing: a systematic literature review,” J. Appl. Econ., vol. 28, no. 1, p. 2466140, 2025, doi: https://doi.org/10.1080/15140326.2025.2466140.

[11] X. Li, C. Schmidt, D. Gammelli, and F. Rodrigues, “Learning Joint Rebalancing and Dynamic Pricing Policies for Autonomous Mobility-on-Demand,” IEEE Trans. Intell. Transp. Syst., 2025, doi: 10.1109/TITS.2025.3582639.

[12] D. Wang, X. Chen, X. Liu, Y. Li, Z. Piao, and H. Li, “Dynamic Tariff Adjustment for Electric Vehicle Charging in Renewable-Rich Smart Grids: A Multi-Factor Optimization Approach to Load Balancing and Cost Efficiency,” Energies, vol. 18, no. 16, p. 4283, 2025, doi: https://doi.org/10.3390/en18164283.

[13] U. Y. Hasanah and R. Rino, “Pricing and Adaptation Strategies in Market Dynamics: A Systematic Literature Review,” Electron. J. Educ. Soc. Econ. Technol., vol. 6, no. 1, 2025, doi: https://doi.org/10.33122/ejeset.v6i1.497.

[14] L. Cheng, M. Li, C. Tan, P. Huang, M. Zhang, and R. Sun, “Computational game-theoretic models for adaptive urban energy systems: A comprehensive review of algorithms, strategies, and engineering applications,” Arch. Comput. Methods Eng., pp. 1–78, 2025, doi: https://doi.org/10.1007/s11831-025-10364-y.

[15] X. Guo and L. Zhang, “Dynamic Pricing Models in E-Commerce: Exploring Machine Learning Techniques to Balance Profitability and Customer Satisfaction,” IEEE Access, 2025, doi: 10.1109/ACCESS.2025.3563371.

[16] B. Jin and X. Xu, “Machine learning-based forecasts of residential property prices in Hangzhou city, Zhejiang province, China,” Neural Comput. Appl., vol. 37, no. 6, pp. 4971–4988, 2025, doi: https://doi.org/10.1007/s00521-024-10726-w.

[17] L. Zheng and L. Tan, “A decentralized scheme for multi-user edge computing task offloading based on dynamic pricing,” Peer--Peer Netw. Appl., vol. 18, no. 2, p. 91, 2025, doi: https://doi.org/10.1007/s12083-025-01904-1.

[18] A. Pagliaro, “Artificial intelligence vs. efficient markets: A critical reassessment of predictive models in the big data era,” Electronics, vol. 14, no. 9, p. 1721, 2025, doi: https://doi.org/10.3390/electronics14091721.

[19] M. Bichler, J. Durmann, and M. Oberlechner, “Algorithmic Pricing and Algorithmic Collusion: M. Bichler et al.,” Bus. Inf. Syst. Eng., pp. 1–9, 2025, doi: https://doi.org/10.1007/s12599-025-00965-z.

[20] P. Michailidis, I. Michailidis, and E. Kosmatopoulos, “Reinforcement Learning for Electric Vehicle Charging Management: Theory and Applications,” Energies, vol. 18, no. 19, p. 5225, 2025, doi: https://doi.org/10.3390/en18195225.

[21] S. Giannelos, “Reinforcement Learning in Energy Finance: A Comprehensive Review.,” Energ. 19961073, vol. 18, no. 11, 2025, doi: 10.3390/en18112712.

[22] P. Michailidis, I. Michailidis, C. R. Lazaridis, and E. Kosmatopoulos, “Traffic Signal Control via Reinforcement Learning: A Review on Applications and Innovations,” Infrastructures, vol. 10, no. 5, p. 114, 2025, doi: https://doi.org/10.3390/infrastructures10050114.

[23] Z. Zhou et al., “A Transformer-Based Reinforcement Learning Framework for Sequential Strategy Optimization in Sparse Data,” Appl. Sci., vol. 15, no. 11, p. 6215, 2025, doi: https://doi.org/10.3390/app15116215.

[24] A. Kavoosi, R. Tavakkoli-Moghaddam, H. Sajedi, N. Tajik, and K. Tafakkori, “Dynamic pricing and inventory control of perishable products by a deep reinforcement learning algorithm,” Expert Syst. Appl., vol. 291, p. 128570, 2025, doi: https://doi.org/10.1016/j.eswa.2025.128570.

[25] I. Poulaki, N. I. Koufodontis, and S. Papadimitriou, “Airline revenue management, distribution and passengers: market trends in a technology driven triangle,” Worldw. Hosp. Tour. Themes, vol. 17, no. 1, pp. 35–47, 2025, doi: https://doi.org/10.1108/WHATT-12-2024-0304.

[26] H. Xu, A. Zhang, Q. Wang, Y. Hu, F. Fang, and L. Cheng, “Quantum Reinforcement Learning for real-time optimization in Electric Vehicle charging systems,” Appl. Energy, vol. 383, p. 125279, 2025, doi: https://doi.org/10.1016/j.apenergy.2025.125279.

Downloads

Published

2026-08-11

How to Cite

[1]
“REINFORCEMENT LEARNING-BASED DYNAMIC PRICING IN A STOCHASTIC DEMAND–SUPPLY ENVIRONMENT”, jitk, vol. 12, no. 1, pp. 28–39, Aug. 2026, doi: 10.33480/jitk.v12i1.8286.

Most read articles by the same author(s)