Modeling of case-cohort data by multiple imputation : application to cardio-vascular epidemiology

The weighted estimators generally used for analyzing case-cohort studies are not fully efficient. However, case-cohort surveys are a special type of incomplete data in which the observation process is controlled by the study organizers. So, methods for analyzing Missing At Random (MAR) data could be appropriate, in particular, multiple imputation, which uses all the available information and allows to approximate the partial maximum likelihood estimator.This approach is based on the generation of several plausible complete data sets, taking into account all the uncertainty about the missing values. It allows adapting any statistical tool available for cohort data, for instance, estimators of the predictive ability of a model or of an additional variable, which meet specific problems with case-cohort data. We have shown that the imputation model must be estimated on all the completely observed subjects (cases and non-cases) including the case indicator among the explanatory variables. We validated this approach with several sets of simulations: 1) completely simulated data where the true parameter values were known, 2) case-cohort data simulated from the PRIME cohort, without any phase-1 variable (completely observed) strongly predictive of the phase-2 variable (incompletely observed), 3) case-cohort data simulated from de NWTS cohort, where a phase-1 variable strongly predictive of the phase-2 variable was available. These simulations showed that multiple imputation generally provided unbiased estimates of the risk ratios. For the phase-1 variables, they were almost as precise as the estimates provided by the full cohort, slightly more precise than Breslow et al. calibrated estimator and still more precise than classical weighted estimators. For the phase-2 variables, the multiple imputation estimator was generally unbiased, with a precision better than classical weighted estimators and similar to Breslow et al. calibrated estimator. The simulations performed with the NWTS cohort data provided less satisfactory results for the effects where the phase-2 variable was involved: the multiple imputation estimators were slightly biased and less precise than the weighted estimators. This can be explained by the interactions terms involving the phase-2 variable in the analysis model and the necessity of estimating specific imputation models in different strata not including sometimes enough cases to satisfy the asymptotic conditions. We advocate the use of multiple imputation for improving the precision of the risk ratios estimates while making sure they are similar to the weighted estimates.Our simulations also showed that multiple imputation provided estimates of a model predictive value (Harrell's C) or of an additional variable (difference of C indices, NRI or IDI) similar to those obtained from the full cohort.

Data and Resources

Additional Info

Field Value
Source https://theses.hal.science/tel-00779739
Author Marti Soler, Helena, Marti Soler
Maintainer CCSD
Last Updated May 15, 2026, 00:07 (UTC)
Created May 15, 2026, 00:07 (UTC)
Identifier NNT: 2012PA11T022
Language fr
Rights https://about.hal.science/hal-authorisation-v1/
contributor Centre de recherche en épidémiologie et santé des populations (CESP) ; Université de Versailles Saint-Quentin-en-Yvelines (UVSQ)-Université Paris-Sud - Paris 11 (UP11)-Assistance publique - Hôpitaux de Paris (AP-HP) (AP-HP)-Hôpital Paul Brousse ; AP-HP. Université Paris Saclay-AP-HP. Université Paris Saclay-Institut National de la Santé et de la Recherche Médicale (INSERM)
creator Marti Soler, Helena, Marti Soler
date 2012-05-04T00:00:00
harvest_object_id 533403b3-3943-436e-95b0-512bbc054887
harvest_source_id 3374d638-d20b-4672-ba96-a23232d55657
harvest_source_title test moissonnage SELUNE
metadata_modified 2026-03-30T00:00:00
set_spec type:THESE