A Heterogeneity-Aware Federated Learning Framework for Multimodal Nanotoxicological Risk Assessment
Abstract
Accurate prediction of engineered nanomaterial (ENM) toxicity is hampered by the fragmented and privacy-sensitive distribution of nanotoxicological datasets across regulatory agencies and research institutions, precluding the centralised data pooling required by conventional machine learning approaches. This study presents a federated learning framework in which K = 4 institutional nodes collaboratively train a shared toxicity prediction model without transmitting raw records. The local model is a multi-branch neural network that jointly processes physicochemical descriptors and gene expression profiles through modality-specific branches integrated via a temperature-scaled cross-modal attention mechanism (τ = 0.70). Global aggregation extends standard Federated Averaging with Kullback–Leibler divergence heterogeneity normalisation, correcting for systematic bias under non-IID institutional data distributions, with convergence guarantees formally characterised. Privacy is enforced through cryptographic secure aggregation with zero accuracy cost, supplemented by optional (ε, δ)-differential privacy. Evaluated on a consolidated corpus of 42,350 nanotoxicological records, the framework achieved accuracy of 89.6%, AUROC of 0.925, F1-score of 0.885, and Matthews Correlation Coefficient of 0.781 on a held-out test set of 6,280 records, grouped by unique nanomaterial identity to prevent data leakage, while keeping all raw data local and maintaining a false negative rate of 12.8% (false positive rate recomputation under the corrected split is in progress). The multimodal design improved F1-score by 12.8% over unimodal baselines, and total inter-node communication overhead was approximately 1.0 GB across 30 rounds. These findings demonstrate that federated learning can deliver strong predictive performance for nanotoxicological risk assessment without compromising data sovereignty.