Overview
A machine learning framework trained on SepsisTensor v1 and restricted to patients with confirmed sepsis and to gene expression alone, without severity scores or demographics. Differential expression, cross-cohort consistency filtering, and XGBoost-based recursive feature elimination selected 36 genes.
The external cohort was kept out of harmonization, feature selection, calibration, and threshold selection. Discrimination there was modest, and the study presents the panel as a molecular baseline for multimodal models rather than a standalone clinical tool.
ComBat harmonization · RFECV · XGBoost · SHAP · Isotonic calibration · Bootstrap inference
- Patients
- 1,636 from seven GEO cohorts
- Internal AUROC
- 0.81 ± 0.02 (five-fold CV)
- External AUROC
- 0.66 (95% CI 0.60–0.72)
- Leave-one-cohort-out
- Pooled AUROC 0.69, I² = 42.3%
- External Brier score
- 0.208 → 0.177 after isotonic calibration
Figures



