Skip to main content
Article

Interpretable Machine-Learning Framework for Engineering Optimization of Microalgal Biofuel Yield Across Cultivation Conditions

Basem Abu ZneidFaculty of Engineering, Hourani Center for Applied Scientific Research, Al-Ahliyya Amman University, Amman, JordanRenuka Jyothi. SDepartment of Biotechnology and Genetics, School of Sciences, JAIN (Deemed to be University), Bangalore, Karnataka, IndiaNoor Mazin BasheerDepartment of Medical Laboratory Technics, College of Health and Medical Technology,Alnoor University, mosul, IraqJayshree NelloreDepartment of Biotechnology, Sathyabama Institute of Science and Technology, Chennai, Tamil Nadu, IndiaVipasha SharmaDepartment of Biotechnology, University Institute of Biotechnology, Chandigarh University, Mohali, Punjab, IndiaSiya SinglaCentre for Research Impact & Outcome, Chitkara University Institute of Engineering and Technology, Chitkara University, Rajpura, 140401, Punjab, IndiaSardor SabirovDepartment of General Professional Sciences, Mamun University, Uzbekistan, KhivaRasul UsmanovDepartment of Chemistry, Urgench State University, Urgench, UzbekistanAhmad KhalidYemenia University
2026en
ABI

Abstract

Optimizing microalgal biofuel yield remains a highly complex engineering challenge due to the intricate, non-linear dependencies among diverse cultivation environments and biological parameters. To provide actionable decision support for harvest scheduling and bioreactor control, this study leverages a comprehensive dataset of 2,950 fully retained experimental records, utilizing statistical leverage diagnostics strictly for homogeneity verification without excluding any operational extremes. By evaluating multiple predictive architectures through rigorous independent testing, the Random Forest model demonstrated superior generalization capability, achieving a test-set R 2 of 0.918 and an average absolute relative error (AARE) of 29.7%. Advanced SHAP further revealed that time and type fundamentally drive predictive variance, while highlighting critical non-linear engineering trade-offs regarding optimal nutrient dosing and light intensity. To ensure robust generalization and mitigate data leakage, a rigorous Leave-One-Source-Out (LOSO) cross-validation strategy was implemented. This grouped validation confirmed that the Random Forest model maintained the highest predictive robustness across independent external data sources.

Identifiers

Citations and references

Cited by 017 references