Асосий контентга ўтиш
Мақола

Navigating operational trade-offs in open-pit mining: a comparative reinforcement learning framework for adaptive ore dispatch

Azamat UmirzokovDepartment of Mining Work, Tashkent State Technical University Named After Islam Karimov, Tashkent, Republic of UzbekistanAidar KuttybayevDepartment of Mining, Satbayev University, 22a Satpaev Str., Almaty, 050013, Republic of KazakhstanShukhratulla OchilovDepartment of Mining and Technology, University of Geological Sciences, Mirzo Ulugbek District, Tashkent, Republic of UzbekistanNurzod NosirovAlmalyk State Technical Institute, Almalyk, Republic of UzbekistanTursoat AmirovTashkent State Transport University, Tashkent, Republic of UzbekistanBakhodir KholievBukhara State Medical Institute Named After Abu Ali Ibn Sino, Bukhara, UzbekistanArystan KozhantovDepartment of Mining, Satbayev University, 22a Satpaev Str., Almaty, 050013, Republic of KazakhstanM R KarimovTashkent State Technical University Named After Islam Karimov, Tashkent, Republic of UzbekistanKhurshida NorovaTashkent State Technical University Named After Islam Karimov, Tashkent, UzbekistanUlugbek UrinovDepartment of Doctoral Studies and Scientific Research, DSc, Professor, Tashkent State Technical University, Tashkent, UzbekistanUmar Do’schanovUrgench State University, Urgench, Khorezm Region, 220100, UzbekistanUchkun EshonkulovDepartment of Geology and Mining, Karshi State Technical University, Karshi, Republic of Uzbekistan
2026en
ABI

Аннотация

Abstract Ore dispatch in open-pit mining is a multidimensional stochastic optimization problem that requires dynamic decision-making to balance competing objectives: maximizing throughput, maintaining grade consistency, and reducing equipment queuing. This study formulates and evaluates a comparative reinforcement learning (RL) framework that learns adaptive dispatch policies in a simulated Internet of Things (IoT)-enabled open-pit mine. Three deep RL algorithms are implemented and compared: value-based dueling double DQN, on-policy proximal policy optimization (PPO), and maximum-entropy soft actor-critic (SAC). The results show that, although all algorithms can learn viable policies, SAC exhibits higher sample efficiency and more stable convergence, reaching strong performance in fewer simulated episodes. A detailed policy analysis reveals that the SAC agent learns an advanced non-intuitive strategy that improves grade control and yields balanced performance across objectives. Still, it does not consistently outperform a simple shortest-queue heuristic on all key performance indicators (KPIs), particularly throughput and queue minimization. This discrepancy exposes an agent–objective alignment problem and indicates that the hand-crafted reward function does not fully capture the underlying business priorities. The proposed framework, therefore, serves not only as an optimization tool but also as a diagnostic mechanism for exploring multi-objective trade-offs and revealing misalignment between engineered rewards and high-level KPIs, providing a practical basis for designing and validating autonomous dispatch systems prior to real-world deployment.

Ҳали таржима қилинмаган

Идентификаторлар

Иқтибослар ва манбалар