Predictive modeling of oral squamous cell carcinoma risk using a multi-parametric machine learning framework integrating clinical, sociodemographic, and salivary biomarker data
Аннотация
Abstract Oral Squamous Cell Carcinoma (OSCC) represents a major public health challenge due to diagnostic delays and the invasive nature of conventional biopsy techniques, underscoring the need for accurate, non-invasive prognostic tools. This study develops and validates a multi-parametric machine learning framework that integrates sociodemographic profiles, lifestyle behaviors, and salivary biomarkers to quantify individual OSCC risk probabilities. A curated dataset of 4,200 clinical samples was analyzed across eight predictive algorithms following rigorous preprocessing, including complete case analysis and feature scaling. Inputs encompassed demographic variables and salivary analytes such as IL-6, MMP-9, and antioxidant capacity. Models were optimized using fivefold cross-validation and evaluated with statistical metrics. The Multi-Layer Perceptron Artificial Neural Network (MLP–ANN) demonstrated superior performance, achieving a testing R2 of 0.997 with minimal relative error, outperforming ensemble and instance-based classifiers. To address interpretability, SHapley Additive exPlanations (SHAP) were applied, revealing alcohol consumption and smoking as dominant global drivers of malignancy risk. The model also captured complex non-linear biological interactions: elevated salivary MMP-9 and IL-8 were strong positive contributors, while antioxidant capacity exhibited a protective inverse relationship against oxidative stress. Dependency plots further highlighted saturation thresholds in carcinogen exposure, offering granular insights into how cumulative lifestyle factors interact with inflammatory biomarkers. Overall, the proposed MLP–ANN framework provides a transparent, high-fidelity tool for early OSCC detection, enabling non-invasive, personalized risk stratification and preventative intervention.
Перевод пока недоступен