Reinforcement learning–enhanced population-based multi-objective ore blending with iot-observed grade uncertainty in open-pit mines
Annotatsiya
Abstract Open-pit ore blending must simultaneously control grade-target compliance, economic value, and operating cost with noisy, partial-grade observations from the Mining 4.0 sensing infrastructure. Building on recent population-based multi-objective reinforcement learning (MORL) architectures for ore blending, this work presents an in-depth empirical study of sequential decision-making under IoT-observed grade uncertainty, involving three competing objectives: grade deviation, net present value (NPV), and cost. To overcome preference collapse and insufficient Pareto coverage in single-policy MORL, we systematically extend the specialist-bank approach by introducing uncertainty-weighted multi-objective gradients, a curriculum noise annealing protocol tied to sensor characteristics, and plant-level operational constraints including crusher throughput and haulage costs. We evaluate this extended architecture against single-policy MORL, rolling-horizon evolutionary control, and static evolutionary baselines on open benchmark mine instances under multiple uncertainty levels. The population MORL achieves a higher peak NPV than static NSGA-II on the primary mine instance, matches the performance of rolling-horizon evolutionary control at three orders of magnitude lower inference latency, and generalizes to an external mine instance (Newman1) without retraining, demonstrating practical transfer to a different orebody distribution. These results suggest that population-based MORL is a viable path to realizing real-time, preference-aware blending control in the face of sensor uncertainty and highlight the importance of empirical, constraint-aware analysis when deploying these architectures alongside offline evolutionary optimization.
Hali tarjima qilinmagan