Skip to main content
Article

Predictive modeling of oral squamous cell carcinoma risk using a multi-parametric machine learning framework integrating clinical, sociodemographic, and salivary biomarker data

Usamah SayedFaculty of Allied Medical Sciences, Hourani Center for Applied Scientific Research, Al-Ahliyya Amman University, Amman, JordanAymn Hassan RashidDepartment of Computer Technology Engineering, College of Technical Engineering, The Islamic University, Najaf, IraqGafur Abdula kimovCentre for Research Impact and Outcome, Chitkara University, Rajpura, Punjab, IndiaSara Qhassan IbrahimDepartment of Medical Devices Engineering, Engineering Technical College, Al-Noor University, Mosul, 41012, IraqEshkobilov OzodbekDepartment of Basic Medical Sciences, Termez University of Economics and Service, Termez, UzbekistanDilbar UrazbaevaDepartment of Psychology and Medicine, Mamun University, Khiva, UzbekistanVipasha SharmaDepartment of Biotechnology, University Institute of Biotechnology, Chandigarh University, Mohali, Punjab, IndiaAbdolali Yarahmadi KandahariFaculty of Engineering, Kandahar University, Kandahar, Afghanistan
2026en
ABI

Abstract

Abstract Oral Squamous Cell Carcinoma (OSCC) represents a major public health challenge due to diagnostic delays and the invasive nature of conventional biopsy techniques, underscoring the need for accurate, non-invasive prognostic tools. This study develops and validates a multi-parametric machine learning framework that integrates sociodemographic profiles, lifestyle behaviors, and salivary biomarkers to quantify individual OSCC risk probabilities. A curated dataset of 4,200 clinical samples was analyzed across eight predictive algorithms following rigorous preprocessing, including complete case analysis and feature scaling. Inputs encompassed demographic variables and salivary analytes such as IL-6, MMP-9, and antioxidant capacity. Models were optimized using fivefold cross-validation and evaluated with statistical metrics. The Multi-Layer Perceptron Artificial Neural Network (MLP–ANN) demonstrated superior performance, achieving a testing R2 of 0.997 with minimal relative error, outperforming ensemble and instance-based classifiers. To address interpretability, SHapley Additive exPlanations (SHAP) were applied, revealing alcohol consumption and smoking as dominant global drivers of malignancy risk. The model also captured complex non-linear biological interactions: elevated salivary MMP-9 and IL-8 were strong positive contributors, while antioxidant capacity exhibited a protective inverse relationship against oxidative stress. Dependency plots further highlighted saturation thresholds in carcinogen exposure, offering granular insights into how cumulative lifestyle factors interact with inflammatory biomarkers. Overall, the proposed MLP–ANN framework provides a transparent, high-fidelity tool for early OSCC detection, enabling non-invasive, personalized risk stratification and preventative intervention.

Identifiers

Citations and references

Cited by 00 references