ABSTRACT
Aim
This study aims to use routinely collected health data and trial emulation methodology to inform the design of a pragmatic randomized controlled trial (RCT) in people requiring multivessel coronary revascularization with severe symptomatic multivessel disease and high-risk characteristics, typically underrepresented in previous RCTs.
Methods
Hospital episode statistics (HES) linked to Office for National Statistics will be the main data source. The study population is patients who require multivessel myocardial revascularization with at least one of the following high-risk characteristics: age >75 years, female, diagnosed with acute coronary syndrome, heart failure, chronic kidney disease, peripheral vascular disease, or intermediate frailty risk. The intervention procedure is coronary artery bypass grafting (CABG) and the control (reference) is percutaneous coronary intervention (PCI). Outcomes include all-cause and cardiovascular (CV) death, CV hospitalization, major adverse cardiovascular events, and major vascular complications or bleeding within 5 years of the index procedure. This study includes 3 stages of statistical analyses: (1) latent class analysis (LCA) to identify mutually exclusive patient clusters (latent classes) representing different clinical phenotypes, (2) instrumental variable analysis (IVA) to estimate the average treatment effect (ATE) in the whole population and each patient cluster; and (3) repeating stage 2 in an emulated trial population obtained by matching the HES population with individual participant data from an RCT. We will then co-design the protocol for a definitive clinical trial in partnership with patients, public, and stakeholders.
Discussion
This study introduces a novel, stepwise data science framework that integrates machine learning (unsupervised learning through LCA), causal inference, and trial emulation methods applied in big data, to design a future stratified and adaptive RCT of CABG versus PCI in high-risk patients. Our proposed approach fosters new collaborations among data scientists, trial methodologists, clinicians, and patient and public representatives in complex trial designs for diverse, high-risk populations. This study represents a new framework for co-production in trials of cardiovascular interventions, which offers a scalable model and has the potential to transfer to other disease areas.
Clinical trial registration
URL : https://www.clinicaltrials.gov/study/NCT05853536 . Unique identifier : NCT05853536.
Background
Treatment guidelines for myocardial revascularization in severe symptomatic multivessel coronary artery disease (CAD) recommend coronary artery bypass grafting (CABG) or percutaneous coronary intervention (PCI) based on the severity and extent of coronary atheroma, and the presence or absence of comorbidities such as diabetes and left ventricular dysfunction. , Older age (>75 years), female sex, ethnicity, comorbidities such as chronic kidney disease (CKD), peripheral vascular disease (PVD), and frailty, or presentation with chronic versus acute coronary syndrome (ACS), are additional factors considered in treatment decisions that are important determinants of clinical outcomes. However, existing randomized controlled trials (RCTs) of revascularization either excluded or recruited low numbers of participants with these high-risk characteristics. , Guidelines make no specific references to these high-risk groups beyond recommendations for PCI in people at high surgical risk. Management of these patients in the absence of evidence leads to unwarranted variation in treatment choices and outcomes.
An RCT of CABG versus PCI in high-risk groups presents several design challenges. First, high-risk groups demonstrate significant overlap or interdependence Trials restricted to individual risk factors, as in the case of diabetes or ischemic left ventricular systolic dysfunction (iLVSD) for example, do not address this clinical complexity. Multiple 2-arm parallel group trials that consider each of these groups individually would take many years to complete. An all-comers trial designed to detect the minimal clinically important difference (MCID) for outcomes in individual high-risk subgroups would need an extremely large sample size to have adequate statistical power. Furthermore, due to the degree of overlap, trial analyses demonstrating futility or effectiveness for 1 high-risk characteristic may affect equipoise in other groups where they are overrepresented. Second, people with different high-risk characteristics may have different rates of outcome accrual, reflecting different disease trajectories, as well as different interactions with the mode of revascularisation. , Third, as the number of high-risk characteristics increases, the size of the treatment effect of the more invasive CABG intervention required to change practice may also differ.
This protocol describes a new framework to address these challenges using a stepwise data science approach. First, we aim to identify nonoverlapping mutually exclusive clusters of patients with similar high-risk characteristics for patient stratification, representing distinctive meaningful clinical phenotypes. Second, we will use causal inference analyses to estimate the treatment effects between the 2 modes of revascularization (CABG vs PCI) for different clinical outcomes for the whole patient population, as well as each clinical phenotype. Third, we will repeat this analysis in a population matched to individual participant data (IPD) from a recent high-quality trial (the ERICCA trial in people with high-risk characteristics undergoing surgical revascularization to estimate the treatment effects in an emulated trial population. Finally, we will use all the findings from the previous 3 stages to co-design an efficient, pragmatic RCT of CABG versus PCI in these high-risk populations in partnership with patients, the public, and stakeholders.
Methods
Study design and data source
This is a retrospective observational population-based cohort study and a trial emulation study, using routinely collected health data from the Hospital Episode Statistics Admitted Patient Care (HES-APC) dataset from April 1, 2007 to March 31, 2020, capturing all NHS hospital admissions in England. The APC dataset includes detailed information of patient demographics, diagnoses, external causes/injuries, operations, bed days, admission method, time waited, specialty, provider, and Adult Critical Care of elective hospital episodes. Patients can opt out of sharing their data and health records for research and planning, or opt back in again, at any time. Our HES datasets are linked to the civil registration of deaths dataset from the Office for National Statistics (ONS) through a unique encrypted identifier based on individual patient’s NHS number for ascertainment of the death outcomes.
Study population, inclusion and exclusion criteria
The study population was adult patients admitted to National Health Service (NHS) hospitals in England between April 1, 2009 and March 31, 2015 (6 fiscal years) who underwent isolated CABG or multivessel PCI, which allowed us to have a complete 5-year follow-up (March 31, 2020), representing routine healthcare before COVID-19. Diagnoses and procedures performed in all hospital episodes within 2 years prior to the index episodes will be used to establish patients’ medical histories. Patients need to have at least one of the following 7 high-risk characteristics to be included in this study:
-
1)
Age >75 years at the index procedure
-
2)
Female sex
-
3)
Diagnosed with acute coronary syndrome
-
4)
Diagnosed with heart failure (HF)
-
5)
Diagnosed with chronic kidney disease; defined as kidney damage or a glomerular filtration rate <60 mL/min/1.73 m 2, persisting for 3 months or more
-
6)
Diagnosed with peripheral vascular disease
-
7)
Intermediate hospital frailty risk score : 5 to 15
Exclusion criteria are:
-
1)
Previous cardiac surgery.
-
2)
Previous PCI 2 years prior to the index procedure during the study period.
-
3)
Index procedures other than isolated CABG or multivessel PCI.
-
4)
Nonmultivessel PCI, as we expect that surgery would be less likely to be considered or indicated in these patients.
-
5)
Pregnancy (maternal admission, either in the preceding, or the year following, the index intervention);
-
6)
Hospitalization for bleeding within 2 years before the index procedure.
-
7)
Patients with high frailty risk defined by the Hospital Frailty Risk Score (HFRS) >15.
Intervention and control procedures
Both PCI and CABG are clinically acceptable procedures for multivessel revascularization. For the purposes of the analyses, we will consider CABG to be the intervention and PCI to be the control. The index episode is defined as the hospital episode in which revascularization (CABG or multivessel PCI) was performed for each patient. To identify a cohort where multivessel coronary disease was treated percutaneously and where CABG might be a viable alternative, we will include only patients undergoing multivessel PCI in the PCI group. Multivessel PCI will be defined by either (1) a single PCI involving multiple coronary arteries (OPCS-4 K492, where OPCS-4 is a classification coding system for operations, procedures and interventions performed in NHS hospitals) or deployment of 3 or more stents (OPCS-4 K752, K754), or (2) a staged PCI where the patient underwent more than 1 PCI intervention (OPCS-4 K49, K50, K75) within 90 days of their first procedure. This definition aims to capture patients with multivessel CAD treated with PCI and aligns broadly with contemporary trials comparing CABG versus PCI for multivessel revascularization. In this study, the intervention is CABG (OPCS-4 codes: K40-K46) and the control is multivessel PCI (OPCS-4: K49-K50 and K75), where OPCS-4 is a statistical classification for clinical coding of hospital interventions and procedures undertaken in the NHS.
Primary and secondary outcomes
The composite primary outcome for this study is all-cause death or cardiovascular hospitalization within 5 years of revascularization. Secondary outcomes include:
-
a.
All-cause death within 5 years (confirmed by ONS deaths registry).
-
b.
Cardiovascular death within 5 years (based on ICD-10 codes).
-
c.
Cardiovascular hospitalization within 5 years.
-
d.
Major adverse cardiovascular events (MACE) including myocardial infarction (MI), stroke, ACS or HF within 5 years.
-
e.
Major vascular complication or hospitalization for bleeding within 5 years, as a safety endpoint.
-
f.
Hospitalization outcomes will be identified using the diagnoses recorded in HES episodes, while death will be ascertained from HES and ONS.
Covariates
HES-APC includes demographic, geographic, diagnostic (ICD-10 coded), and procedural (OPCS-4 coded) data, with an accuracy of 96% for diagnostic coding and 97% for procedural coding Patient demographics, including age, sex, ethnicity, and the index of multiple deprivation (IMD), will be extracted from the index episode. The IMD includes 7 distinct domains (income, employment, education, skills and training, health deprivation and disability, crime, barriers to housing and services, and living environment), measuring relative deprivation of the patients based on their residential postcodes in England. Clinical features, such as diabetes, hypertension, lipidemia, CKD, stroke, MI, PVD, HFRS, and Charlson comorbidity index will be extracted within 2 years of the index procedure.
Missing data
Date of birth and sex are well documented in HES. A small number of patients without sex information will be excluded from the analysis. The high-risk clinical characteristics are based on ICD-10 codes in the diagnosis field (up to 20 diagnoses). We assume patients did not have the comorbidities if relevant ICD-10 codes are absent. Multiple imputation with chained equations will be used to impute missing data in IMD (predictive mean matching for continuous variable) and ethnicity (multinomial logistic regression for categorical variable) using covariates and outcome variables stated above in the imputation model with 10 imputations.
Planned statistical analyses for the 4 research objectives
Stage 1: Latent class analysis (LCA)
We hypothesize the study population can be classified as patient subgroups, each with distinctive clinical patterns. However, the number of subgroups is unknown. We will use latent class analysis (LCA) to identify classes/clusters using the observed information, ie, the 7 prespecified binary variables (age, sex, ACS, HF, CKD, PVD, and frailty), a type of unsupervised learning. Latent class model is a probability-based unsupervised learning method, using maximum likelihood to estimate the parameters The association between the observed characteristics and the latent classes is probabilistic. The model can generate posterior probabilities for class memberships in all identified classes, while cluster analysis cannot yield posterior probabilities. Individuals will be assigned to the most likely class based on the largest estimated posterior probability. Noticeably, membership of individual patients to classes/clusters is estimated with a certain level of uncertainty (“error”). Therefore, the use of LCA is usually more exploratory rather than confirmatory. The output of LCA is patient profiles and disease phenotypes in different classes/clusters. We will use different visualization graphs to support the interpretation of the meaning of the identified clusters.
The identified latent classes will be mutually exclusive and collectively exhaustive. The classes are different from each other and have their own features. Each patient belongs to just 1 class/cluster (classification). Individuals with the same high-risk characteristics will have the same probability and be in the same class. Ideally, we would like the model to maximize the intragroup homogeneity and intergroup heterogeneity. We will explore the solutions between 2 and 9 latent classes, and discuss the results in a research team including statisticians, clinicians (interventional cardiologists and cardiac surgeons), and trialists. We will use the following criteria to decide the optimal number of latent classes for patient stratification:
1. Information criteria: Akaike information criterion (AIC) and Bayesian information criterion (BIC) account for model fit (log-likelihood values) and model complexity (penalizing more complex models with a larger number of parameters in the model). Lower AIC and BIC values indicate better model fit.
2. Clinical consensus: The identified latent classes need to make sense in clinical practice through discussion with clinicians, representing certain types of patients that cardiologists and cardiac surgeons will recognize in the clinic.
3. Practical consideration for potential translating the finding to a 2-arm, parallel, multiple strata RCT: The number of latent classes/clusters and the sample size in each class/cluster should be feasible for participant recruitment.
The team will reach a consensus on the best solution from LCA to carry forward for patient stratification to the analyses using instrumental variable analysis and trial emulation. Ideally, the number of latent classes is manageable in a trial, with a reasonable sample size in each class.
Stage 2: Instrumental variable analysis (IVA)
We hypothesize that different phenotypes (latent classes) would have heterogeneous treatment effects, which will be tested using instrumental variables analysis (IVA). IVA is a 2-stage regression to address the risk of bias from both measured and unmeasured confounding and to estimate the Average Treatment Effects (ATEs) for the overall population and subpopulations (the identified latent classes in Stage 1). There are some key assumptions for instrumental variables (IV) : (1) relevance—the instrument must be associated with the intervention; (2) independence—the instrument and the outcome have no uncontrolled common causes; (3) the instrument must only affect the outcome through the intervention, which cannot be evaluated empirically Physician preference has been widely used as an IV. , Clinical commissioning group (CCG), existed between 2013 and 2022, was a type of NHS organization which planned and commissioned most hospital and community health services for their local populations in respective geographical regions across England. Previous empirical studies ,, have demonstrated regional variation of patients receiving CABG or PCI, and used standardized CABG rate as IV, which will be the IV in this study.
We will use 2-stage residual inclusion (2SRI) ,, for IV estimation in a survival context (time-to-event data). The first stage regression is a probit regression of the binary treatment options on the IV with relevant sociodemographic and clinical variables. In the second stage, the clinical outcomes (eg, 5-year all-cause mortality) will be regressed on the treatment procedure (CABG or PCI), adjusting for the same sociodemographic and clinical variables, and the generalized residuals from the first stage model. To recognize statistical uncertainty in the estimates of treatment effects, bootstrapping will be used to obtain the standard errors of the 2SRI estimator An empirical study, using similar methods in the English clinical practice research datalink data, bootstrapped 500 times and stratified the bootstrap resampling by CCG, treatment group, death and censoring status to maintain the structure of the original population across replicates. We will do the same in this study if feasible in computational power.
We will use Cox proportional hazards models for all outcomes whose data type is time-to-event to account for censoring, as well as competing risk regression (the Fine–Gray model) for hospitalization or MACE outcomes recorded in HES within 5 years after the index episode, with death as a competing risk for the hospitalization events, as a sensitivity analysis. We will estimate the hazard ratios (HR and bootstrap confidence interval, CI) when using Cox regression and subhazard ratios (SHR and bootstrap CI) for other adverse hospitalization outcomes using the Fine–Gray model.
Stage 3: Trial emulation in a high-risk patient population
To estimate the ATE in the target trial population (patients at high-risk), we will match the HES population to an existing trial recruiting high-risk patients (the ERICCA trial ) using prespecified baseline characteristics and comorbidities in the emulated trial modelling framework. This ERICCA trial evaluated remote ischemic preconditioning in 1,612 CABG cases across 18 UK centers. This will derive an emulated patient cohort more accurately reflecting an actual trial population. We will follow Hernán and Robins’ suggestions of using big data to emulate a target trial, with the key elements in the trial specified in Table . We will use propensity score matching (PSM) to match patients in the HES dataset to individual participants (IPD) in the ERICCA trial to obtain 2 matched cohorts: (1) HES CABG–> ERICCA and (2) HES PCI–> ERICCA cohorts. The matching is based on the “ participant’s baseline high-risk profile, ” as this is the population we are likely to recruit if we run such a trial and use the same facility and infrastructure in the UK. Clinically important covariates available in both ERICCA and HES datasets, such as age, sex, and comorbidities (diabetes, hypertension, CKD, previous MI, and PVD) will be used in PSM, and evaluated whether the balance among the covariates is achieved or not after matching. We will repeat the IVA (Stage 2) to estimate the ATE from the emulated CABG (1) and PCI (2) cohorts and in each cluster. The differences between the treatment effects in the emulated cohort (ERICCA matched to the HES dataset) and the whole HES patient population will provide an estimate of the likely generalizability of the target trial to the overall target population (external validity), as well as a better estimate of likely target recruitment rates. A diagram for the research process from Stages 1 to 3 is in Figure .
Table
Target trial and emulated trial for a superiority trial of CABG versus complex PCI in patients with high-risk characteristics.
| Target trial | Emulated trial | |
|---|---|---|
| Trial design | Parallel, stratified | Parallel, stratified |
| Blinding | Open label | Open label |
| Setting | 30 cardiac centers in the UK | Patients received NHS care in England |
| Eligibility criteria of participants |
Inclusion criteria: ALL of the following
|
Inclusion criteria: ALL of the following
|
| Stratification | Strata will be assigned for the main trial based on the results of the emulated study. | Strata in the emulated trial will be as defined in the section above (latent class analysis). |
| Comparative populations |
Intervention: CABG
Control (reference): Complex PCI |
Patients undergoing CABG (OPCS-4 K40-K46) will be compared with patients undergoing high-risk complex PCI procedures.Complex PCI will be defined as either:
|
| Treatment allocation | Randomization in 1:1 Ratio—computer generated | Randomization will be emulated by harnessing the naturally occurring regional variation in treatment practices across hospitals. Using the regional surgical rate as an instrument variable, we will conduct instrument variable analysis which controls for known and unknown confounders to estimate treatment effects. |
| Recruitment | Patients will be identified in outpatient clinics, wards, and in heart team meetings. | Patients with diagnosis of high-risk characteristics, hospitalized between 2009 and 2015, and undergoing CABG or complex PCI. Index episodes and all hospital admissions within 2 y prior to the index episode will be used to establish patient characteristics and eligibility. |
| Follow up | Follow-up will begin at randomization until 5 y. | Follow-up after the index intervention (CABG or complex PCI) for a minimum of 5 y. |
| Primary outcome | Composite of all-cause mortality and cardiovascular hospitalizations at 5 y (will co-design with patient and public representatives and stakeholders). |
Composite of all-cause mortality and cardiovascular hospitalizations (ICD I00-I99) identified based on primary diagnosis at 5 y.
Sensitivity analysis to be conducted using both primary and secondary diagnosis. |
| Secondary outcome |
|
|
| Statistical analysis | Cox or Fine–Gray model regresses outcomes on randomized group in the overall study population and each individual stratum. | Instrumental variable analysis (IVA) regresses outcomes on treatment type (CABG vs PCI) and covariates, using regional (CCG) surgical rates as the instrument variable for treatment type. |
Stay updated, free articles. Join our Telegram channel
Full access? Get Clinical Tree