Highlights
-
•
Identified 5 comorbidity-based clusters in over 1 million heart failure patients.
-
•
Cluster 5 with highest comorbidity burden had increased 30-day readmission risk.
-
•
XGBoost showed best performance predicting short-term HF outcomes (AUROC 0.76).
-
•
Age and Charlson Comorbidity Index were key predictors across models.
-
•
Comorbidity-based clustering can improve risk stratification and personalized care.
Heart failure (HF) is a major global health burden, and complex comorbidity patterns can worsen clinical outcomes and complicate patient care. This study aimed to identify distinct comorbidity-based clusters among HF patients and evaluate their associations with short-term clinical outcomes. We analyzed electronic health records from 1,010,573 HF patients in China between 2021 and 2024 and classified into 5 distinct clusters using the Clustering Large Applications (CLARA) algorithm. Cluster 5, characterized by the highest comorbidity burden, was associated with an increased risk of 30-day readmission (adjusted OR: 1.29, 95% CI: 1.25 to 1.33), whereas Clusters 2 and 3 demonstrated lower risks compared with the reference group. XGBoost achieved the best predictive performance among multiple machine learning models (area under the receiver operating characteristic curve 0.76; Brier score 0.17). Age and Charlson Comorbidity Index score were the most influential predictors, and features derived from the comorbidity clusters provided additional predictive value. In conclusion, these findings demonstrate substantial heterogeneity among HF patients, highlight the clinical relevance of comorbidity-based clustering, and suggest its potential to improve risk stratification and personalized care strategies.
Heart failure (HF) is a significant global health burden, affecting more than 64 million individuals worldwide, with prevalence increasing due to population aging and rising cardiovascular risk factors. In China, approximately 13 million adults are affected, representing 1.3% of the adult population, with disease burden increasing sharply with age and disproportionately affecting men over 60 years. Comorbid conditions are highly prevalent, with over 70% of hospitalized patients having at least 1 comorbidity, and higher Charlson Comorbidity Index (CCI) scores are strongly associated with adverse outcomes, including mortality and hospital readmission. Despite therapeutic advances, the prevalence and impact of comorbidities have remained largely unchanged, and accumulating evidence indicates that comorbidities cluster in distinct patterns, contributing to substantial heterogeneity in clinical course and treatment response. Recent studies highlight the growing burden of HF-related hospitalizations and rising mortality, particularly as the prevalence of comorbidities increases. , In this study we applied the Clustering Large Applications (CLARA) algorithm to identify clinically interpretable subgroups based on comorbidity patterns, , and the XGBoost algorithm to predict clinical outcomes. By combining these methods, we sought to characterize comorbidity-driven heart failure phenotypes and enhance risk stratification for personalized management.
Methods
Study design and patient enrollment
This retrospective cohort study used the Jiangsu Population Health Records Big Data Platform, a province-wide database established in 2012 that integrates standardized information from 218 tertiary hospitals, including inpatient and outpatient encounters, prescriptions, and procedures, with diagnoses coded by the International Classification of Diseases, 10th Revision (ICD-10). The study period was January 1, 2021, to July 31, 2024. The database has supported research on chronic disease and population health.
Patients aged 18 years or older with a confirmed diagnosis of HF during the study period were eligible. HF was identified primarily through ICD-10 codes, with standardized Chinese diagnostic terminology applied as a complementary approach to ensure complete ascertainment ( Supplementary Table 2 ). The combined use of ICD coding and diagnosis text or keywords has been described in previous Chinese studies, providing precedent for applying this supplementary approach in our analysis. The index date was the earliest diagnosis, corresponding to the discharge date for hospitalized patients or the visit date for outpatients. The period preceding this date was defined as the baseline interval. Patients were excluded if demographic information was incomplete or if they were younger than 18 years at the index date ( Figure 1 ). The study adhered to the Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) guidelines ( Supplementary Table 1 ).
Study flow diagram for cohort selection and exclusion criteria. Flowchart illustrating the inclusion and exclusion criteria for the heart failure cohort ( n = 1,010,573).
Baseline characteristics
Age, sex, and year of diagnosis were collected at the index date. Comorbidities were identified from inpatient and outpatient records based on ICD-10 codes, including myocardial infarction, peripheral vascular disease, cerebrovascular disease, dementia, chronic pulmonary disease, connective tissue disease, peptic ulcer disease, mild liver disease, hemiplegia, moderate to severe renal disease, tumor without metastasis, moderate to severe liver disease, metastatic solid tumors, AIDS, and diabetes mellitus. Both the CCI and the age-adjusted Charlson Comorbidity Index (ACCI), the latter assigning an additional point per decade of age from 50 years upwards (maximum 4 points), were calculated ( Supplementary Table 2 ).
Clinical outcome assessment
The follow-up period began on the index date of HF diagnosis and continued until the first occurrence of a prespecified clinical endpoint or the end of the study period on July 31, 2024. The predefined outcomes included HF-related hospital readmission within 30 days and, following a 90-day clinically stable period, hospital readmission, emergency department visit, or outpatient initiation of diuretic therapy due to worsening HF. These outcomes were selected to capture both short-term adverse events and clinical deterioration after initial stabilization. Diuretic use was identified from prescription records using a predefined list of medications determined through name-based keyword matching, including both generic and brand names ( Supplementary Table 3 ).
Cluster analysis
CLARA was selected for its computational efficiency in large-scale datasets and its robustness to noise and outliers. The algorithm applied k-modes clustering with Manhattan distance (L1 norm) to quantify dissimilarity, a metric particularly appropriate for categorical data because it captures absolute differences between features. All baseline comorbidities were included as individual binary features to capture the full spectrum of disease burden. This data-driven approach allowed us to identify clinically meaningful comorbidity patterns without imposing prior assumptions, thereby preserving the granularity and integrity of the dataset. The optimal number of clusters was determined through systematic evaluation of internal validation indices, including the average silhouette coefficient, Calinski–Harabasz (CH) index, and Dunn index, to ensure both statistical robustness and clinical interpretability of the clustering solution.
Machine learning model analysis
Three supervised machine learning algorithms were used to predict clinical outcomes. These included XGBoost, Random Forest, and LASSO logistic regression, selected for their respective advantages in modeling nonlinear patterns, ensemble prediction, and feature selection. Model performance was evaluated using F1-score, sensitivity, specificity, area under the receiver operating characteristic curve (AUROC), area under the precision-recall curve (AUPR), and Brier score. XGBoost showed the best overall performance and was selected as the final model. To evaluate the contribution of clustering-derived features, a comparison was conducted between a baseline XGBoost model and a version incorporating cluster membership as an additional predictor.
Statistical analysis
Statistical analyses were conducted using baseline patient characteristics, cluster assignments, and clinical outcome information. The integration of clustering results with machine learning outputs enabled a detailed evaluation of the associations between comorbidity patterns and adverse clinical events. The primary outcomes were 30-day readmission and subsequent clinical deterioration following a 90-day stable period. For each cluster, the number and proportion of patients with adverse outcomes were calculated. To estimate the association between cluster membership and clinical outcomes, crude odds ratio (OR) with corresponding 95% CIs were first computed. Adjusted OR (aOR) were then calculated after controlling for potential confounders, including age and sex. To further assess the predictive value of clustering, cluster-based features were integrated into the final XGBoost model. Model performance was compared before and after the inclusion of clustering features to evaluate improvements in predictive accuracy. All statistical analyses were performed using R software version 4.2.1. To provide an overview of the study design and analytical workflow, a central figure summarizing the data source, feature extraction, clustering approach, predictive modeling, and clinical implications is presented in Figure 2 .
Central figure for comorbidity-based clustering and clinical outcome prediction in heart failure. Overview of the workflow from cohort selection and feature extraction to clustering, machine learning evaluation, and diverse clinical outcomes supporting personalized management.
Results
Cluster analysis results
Internal validation metrics including the average silhouette coefficient, CH index, and Dunn index were used to identify the optimal number of clusters using the CLARA algorithm. While the CH index (2141.95) decreased slightly with 5 clusters, both the average silhouette coefficient (0.25) and Dunn index (0.21) remained relatively high, indicating a favorable balance between intracluster cohesion and intercluster separation ( Supplementary Table 4 ). This combination indicated a favorable balance between cohesion and separation. The distribution of comorbidity burden across clusters further supported clinical interpretability.
To assess the stability of the solution, additional cluster numbers were also tested. However, with the variation in the number of clusters, no significant improvements in clinical granularity or predictive accuracy were observed. Although higher cluster numbers identified smaller subgroups, they did not enhance clinical insights or the prediction of adverse outcomes. Therefore, based on both statistical performance and clinical relevance, the 5-cluster solution was deemed optimal.
Baseline characteristics by cluster
The final cohort included 1,010,573 HF patients. Table 1 presents baseline characteristics by cluster. Mean age varied markedly across clusters, ranging from 50.1 (8.6) to 87.2 (33.2) years. The proportion of female patients also differed significantly, with the highest representation in Cluster 1 (50.9%) and the lowest in Cluster 3 (34.7%). Across all clusters, cerebrovascular disease was the most common comorbidity, reaching a maximum prevalence of 62.2%, followed by chronic kidney disease and diabetes mellitus. Clusters 4 and 5 demonstrated notably higher overall comorbidity burdens compared to the remaining clusters. The mean (SD) CCI scores ranged significantly from 3.8 (2.5) to 9.6 (2.2) across clusters (all p <0.001).
Table 1
Baseline characteristics of HF patient clusters identified using CLARA clustering
| Characteristics | Cluster 1 | Cluster 2 | Cluster 3 | Cluster 4 | Cluster 5 | p-value |
|---|---|---|---|---|---|---|
| N = 285,916 | N = 228,793 | N = 175,635 | N = 237,551 | N = 82,678 | ||
| Age, years | ||||||
| Mean (SD) | 82.9 (19.5) | 67.9 (4.3) | 50.1 (8.6) | 71.9 (5.2) | 87.2 (33.2) | <0.001 |
| 18-49 | 0 | 0 | 60,802 (34.6) | 0 | 0 | |
| 50-64 | 0 | 52,568 (23.0) | 114,833 (65.4) | 26,810 (11.3) | 0 | |
| 65-80 | 102,661 (35.9) | 176,225 (77.0) | 0 | 208,200 (87.6) | 5,567 (6.7) | |
| >80 | 183,255 (64.1) | 0 | 0 | 2,541 (1.1) | 77,111 (93.3) | |
| Sex | ||||||
| Male | 140,491 (49.1) | 128,556 (56.2) | 114,619 (65.3) | 135,626 (57.1) | 45,817 (55.4) | <0.001 |
| Female | 145,425 (50.9) | 100,237 (43.8) | 61,016 (34.7) | 101,925 (42.9) | 36,861 (44.6) | |
| Year of diagnosis | ||||||
| 2021 | 78,534 (27.5) | 66,861 (29.2) | 45,761 (26.1) | 53,495 (22.5) | 17,202 (20.8) | <0.001 |
| 2022 | 74,252 (26.0) | 59,255 (25.9) | 45,318 (25.8) | 60,239 (25.4) | 21,269 (25.7) | |
| 2023 | 86,278 (30.2) | 67,265 (29.4) | 54,757 (31.2) | 78,240 (32.9) | 27,333 (33.1) | |
| 2024 | 46,852 (16.4) | 35,412 (15.5) | 29,799 (17.0) | 45,577 (19.2) | 16,874 (20.4) | |
| CCI | ||||||
| Mean (SD) | 5.4 (1.2) | 4.3 (1.2) | 3.8 (2.5) | 8.6 (2.1) | 9.6 (2.2) | <0.001 |
| 0-2 | 0 | 8,122 (3.5) | 63,721 (36.3) | 0 | 0 | |
| 3-5 | 179,078 (62.6) | 187,661 (82.0) | 73,561 (41.9) | 0 | 0 | |
| 6-8 | 103,645 (36.3) | 32,391 (14.2) | 30,648 (17.4) | 139,149 (58.6) | 28,706 (34.7) | |
| ≥9 | 3,193 (1.1) | 619 (0.3) | 7,705 (4.4) | 98,402 (41.4) | 53,972 (65.3) | |
| ACCI | ||||||
| Mean (SD) | 9.1 (1.3) | 6.6 (1.3) | 4.5 (2.7) | 11.3 (2.2) | 13.5 (2.1) | <0.001 |
| 0-2 | 0 | 0 | 36,973 (21.1) | 0 | 0 | |
| 3-5 | 0 | 52,536 (23.0) | 81,652 (46.5) | 0 | 0 | |
| 6-8 | 110,422 (38.6) | 165,820 (72.5) | 44,330 (25.2) | 2,899 (1.2) | 0 | |
| ≥9 | 175,494 (61.4) | 10,437 (4.6) | 12,680 (7.2) | 234,652 (98.8) | 82,678 (100.0) | |
| Comorbidity count | ||||||
| 0-1 | 199,314 (69.7) | 182,160 (79.6) | 99,029 (56.4) | 4,934 (2.1) | 0 | <0.001 |
| 2-3 | 85,305 (29.8) | 46,611 (20.4) | 58,748 (33.4) | 136,324 (57.4) | 37,245 (45.0) | |
| ≥4 | 1,297 (0.5) | 22 (0.0) | 17,858 (10.2) | 96,293 (40.5) | 45,433 (55.0) | |
| Insurance | ||||||
| Medical insurance for urban workers | 73,399 (25.7) | 57,565 (25.2) | 43,383 (24.7) | 65,862 (27.7) | 24,453 (29.6) | <0.001 |
| Medical insurance for urban residents | 53,178 (18.6) | 36,925 (16.1) | 18,544 (10.6) | 35,408 (14.9) | 11,774 (14.2) | |
| New rural cooperative medical care | 1,950 (0.7) | 1,000 (0.4) | 506 (0.3) | 967 (0.4) | 346 (0.4) | |
| Poverty relief | 213 (0.1) | 125 (0.1) | 77 (0.0) | 78 (0.0) | 64 (0.1) | |
| Commercial medical insurance | 969 (0.3) | 619 (0.3) | 274 (0.2) | 443 (0.2) | 151 (0.2) | |
| All-state expense | 233 (0.1) | 121 (0.1) | 149 (0.1) | 120 (0.1) | 332 (0.4) | |
| All expenses paid | 18,364 (6.4) | 15,577 (6.8) | 14,104 (8.0) | 11,649 (4.9) | 4,502 (5.4) | |
| Other social insurance | 829 (0.3) | 439 (0.2) | 233 (0.1) | 296 (0.1) | 156 (0.2) | |
| Baseline comorbidities | ||||||
| Myocardial infarction | 18,475 (6.5) | 16,303 (7.1) | 24,379 (13.9) | 33,202 (14.0) | 10,014 (12.1) | <0.001 |
| Peripheral vascular disease | 18,758 (6.6) | 11,422 (5.0) | 12,514 (7.1) | 48,306 (20.3) | 18,277 (22.1) | |
| Dementia | 4,820 (1.7) | 860 (0.4) | 783 (0.4) | 5,418 (2.3) | 6,091 (7.4) | |
| Cerebrovascular disease | 116,356 (40.7) | 52,066 (22.8) | 34,904 (19.9) | 137,274 (57.8) | 51,390 (62.2) | |
| Rheumatic tissue | 5,781 (2.0) | 7,696 (3.4) | 13,275 (7.6) | 21,597 (9.1) | 4,857 (5.9) | |
| Ulcer disease | 4,159 (1.5) | 3,097 (1.4) | 3862 (2.2) | 10,316 (4.3) | 3,586 (4.3) | |
| Type 1 diabetes | 166,25 (5.8) | 14,972 (6.5) | 42,611 (24.3) | 153,900 (64.8) | 50,349 (60.9) | |
| Type 2 diabetes | 9,685 (3.4) | 9,987 (4.4) | 38,637 (22.0) | 147,083 (61.9) | 47,344 (57.3) | |
| Chronic pulmonary disease | 61,720 (21.6) | 27,682 (12.1) | 7,638 (4.3) | 48,722 (20.5) | 30,084 (36.4) | |
| Mild liver disease | 13,063 (4.6) | 12,643 (5.5) | 23,846 (13.6) | 41,248 (17.4) | 13,096 (15.8) | |
| Hemiplegia or paraplegia | 1,172 (0.4) | 471 (0.2) | 947 (0.5) | 4,055 (1.7) | 1,401 (1.7) | |
| Chronic kidney disease | 34,162 (11.9) | 17,704 (7.7) | 47,491 (27.0) | 91,800 (38.6) | 48,998 (59.3) | |
| Tumor | 13,820 (4.8) | 8,538 (3.7) | 8,912 (5.1) | 37,805 (15.9) | 14,781 (17.9) | |
| Moderate to severe liver disease | 4,570 (1.6) | 3,455 (1.5) | 11,190 (6.4) | 20,986 (8.8) | 9,734 (11.8) | |
| Metastatic solid tumor | 588 (0.2) | 591 (0.3) | 2,029 (1.2) | 8,235 (3.5) | 2,736 (3.3) | |
| AIDS | 7 (0.0) | 28 (0.0) | 179 (0.1) | 159 0.1) | 16 (0.0) | |
Stay updated, free articles. Join our Telegram channel
Full access? Get Clinical Tree