Direct and Unbiased Multiple Imputation Methods for Missing Values of Categorical Variables
[摘要] Missing data is a common problem in statistical analyses. To make use of information in data with incomplete observation, missing values can be imputed so that standard statistical methods can be used to analyze the data. Variables with missing values are often categorical and the miss ing pattern may not be monotone. Currently, commonly used imputation methods for data with a non-monotone missing pattern do not allow di rect inclusion of categorical variables. Categorical variables are converted to numerical variables before imputation. For many applications, the imputed numerical values for those categorical variables must then be converted back to categorical values. However, this conversion introduces bias which can seriously affect subsequent analyses. In this paper, we propose two direct imputation methods for categorical variables with a non-monotone missing pattern: the direct imputation approach incorporated with the expectation maximization algorithm and the direct imputation approach incorporated with a new algorithm: the imputation-maximization algorithm. Simulation studies show that both methods perform better than the method using vari able conversion. An application to real data is provided to compare the direct imputation method and the method using variable conversion.
[发布日期] [发布机构]
[效力级别] [学科分类] 土木及结构工程学
[关键词] bias;categorical variable;HIV [时效性]