Asosiy kontentga oʻtish
Maqola

Reinforcement learning–enhanced population-based multi-objective ore blending with iot-observed grade uncertainty in open-pit mines

Azamat UmirzoqovDepartment of Mining Work, Tashkent State Technical University, Tashkent, Republic of UzbekistanShokhjakhon AbdufattokhovDepartment of Information Technologies, Tashkent International University of Education, Tashkent, UzbekistanRyumduk OhDepartment of Software & IT Convergence, Korea National University of Transportation, Chungju-si, South KoreaShuhratulla OchilovDepartment of Mining and Technology, University of Geological Sciences, Tashkent, Mirzo Ulugbek district, Republic of UzbekistanAidar KuttybayevDepartment of Mining, Satbayev University, 22a Satpaev str, Almaty, 050013, Republic of KazakhstanМахфуза ТухтаеваDepartment of Mining and Technology, University of Geological Sciences, Mirzo Ulugbek district, Olimlar Str., 64, Tashkent, 100164, Republic of UzbekistanKazi BizhanovEngineeringAstana Design and Research Center, Represented by Director Kazi Balkenovich, GREY PLAZA Business Center, Yesil district, Kerey Zhanibek Khandar Street 32Astana, Ekibastuz, 010000, Republic of UzbekistanArystan KozhantovDepartment of Mining, Satbayev University, 22a Satpaev str, Almaty, 050013, Republic of KazakhstanSuhbat NorinovThe Almalyk branch of the National Technological Research University “MISIS”, Almalyk City, 110100, UzbekistanGalymzhan SamenovEurasian National University named after L.N. Gumilev, Astana, Republic of KazakhstanAinash KainazarovaDepartment of Mining, Ekibastuz Engineering and Technical Institute named after Academician K.Satpayev, 54 «а» Energetikov St., Ekibastuz, 141208, Republic of Kazakhstan
2026en
ABI

Annotatsiya

Abstract Open-pit ore blending must simultaneously control grade-target compliance, economic value, and operating cost with noisy, partial-grade observations from the Mining 4.0 sensing infrastructure. Building on recent population-based multi-objective reinforcement learning (MORL) architectures for ore blending, this work presents an in-depth empirical study of sequential decision-making under IoT-observed grade uncertainty, involving three competing objectives: grade deviation, net present value (NPV), and cost. To overcome preference collapse and insufficient Pareto coverage in single-policy MORL, we systematically extend the specialist-bank approach by introducing uncertainty-weighted multi-objective gradients, a curriculum noise annealing protocol tied to sensor characteristics, and plant-level operational constraints including crusher throughput and haulage costs. We evaluate this extended architecture against single-policy MORL, rolling-horizon evolutionary control, and static evolutionary baselines on open benchmark mine instances under multiple uncertainty levels. The population MORL achieves a higher peak NPV than static NSGA-II on the primary mine instance, matches the performance of rolling-horizon evolutionary control at three orders of magnitude lower inference latency, and generalizes to an external mine instance (Newman1) without retraining, demonstrating practical transfer to a different orebody distribution. These results suggest that population-based MORL is a viable path to realizing real-time, preference-aware blending control in the face of sensor uncertainty and highlight the importance of empirical, constraint-aware analysis when deploying these architectures alongside offline evolutionary optimization.

Hali tarjima qilinmagan

Identifikatorlar

Iqtiboslar va manbalar