2025 Bandit Algorithm for Online Learning in Inventory Problems with Lead TIMES
페이지 정보
작성자 관리자 작성일 26-07-15 12:42본문
- 개최지
- 미국
- 발표형식
- 구두
- 년도
- 2025
We study data-driven inventory management problems using a bandit algorithm. While many data-driven approaches have been successfully applied to inventory problems, there is a lack of truly online and universal algorithms. While prior work explored deep reinforcement learning approaches as a universal methodology for inventory problems, they lack the online learning capability that bandit algorithms offer. To address issues with applying bandit algorithms to inventory problems, our algorithm has following three key concepts: 1) low-risk exploitation through a virtual inventory system, 2) management of a candidate arm set, 3) efficient exploration via Thompson sampling. We evaluate the performance of our algorithm and obtain regret upper bounds.