This approach improves risk assessment in underwriting using machine learning and EHR data, suggesting more precise predictions.
Traditional health insurance underwriting methods rely heavily on actuarial models based on static demographic and historical cost data, limiting their ability to reflect individual health risks accurately. This study proposes a machine learning-based framework to improve personalized risk stratification by leveraging claims data, electronic health records (EHR), and lifestyle indicators. The framework integrates eXtreme Gradient Boosting (XGBoost) with a feedforward neural network (FNN) comprising three hidden layers and incorporating ReLU activation, dropout regularization, and batch normalization. The hybrid model was trained and evaluated on a real-world dataset containing over anonymized member records from a large U.S. insurer. It achieved an AUC-ROC of 0.79 significantly outperforming traditional baseline methods. Model interpretability was addressed using SHAP to identify key risk drivers. This journal outlines an approach that supports dynamic, data-driven underwriting decisions while maintaining compliance and transparency. These results demonstrate that machine learning can enhance accuracy, efficiency, and fairness in health insurance risk assessment.
No takes yet. Share an insight, caveat, or question.
Kalyanasundaram et al. (2025) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: