Improving accuracy of missing data imputation in data mining

Abstract = 38 times | PDF = 36 times

##plugins.themes.bootstrap3.article.main##

Nzar A. Ali Zhyan M. Omer

Abstract

In fact, raw data in the real world is dirty. Each large data repository contains various types of anomalous values that influence the result of the analysis, since in data mining, good models usually need good data, databases in the world are not always clean and includes noise, incomplete data, duplicate records, inconsistent data and missing values. Missing data is a common drawback in many real-world data sets. In this paper, we proposed an algorithm depending on improving (MIGEC) algorithm in the way of imputation for dealing missing values. We implement grey relational analysis (GRA) on attribute values instead of instance values, and the missing data were initially imputed by mean imputation and then estimated by our proposed algorithm (PA) used as a complete value for imputing next missing value.We compare our proposed algorithm with several other algorithms such as MMS, HDI, KNNMI, FCMOCS, CRI, CMI, NIIA and MIGEC under different missing mechanisms. Experimental results demonstrate that the proposed algorithm has less RMSE values than other algorithms under all missingness mechanisms.

Keywords

Data mining; Missing value; Missing value ; Data preprocessing

References

[1] U. Fayyad, G. Piatetsky-Shapiro, P. Smyth ,(1996) “From data mining to knowledge discovery”, American Association for Artificial Intelligence, San Francisco, Vol. 17, No. 3.
[2] R. Nisbet , J. Elder, G. Miner , (2009) “Handbook of Statistical Analysis and Data Mining Applications” . Academic Press, Boston.
[3] J. Han and M. Kamber, (2011) “Data Mining: Concepts and Techniques”, Morgan Kaufmann,San Francisco .
[4] J. Tian , B. Yu , D. Yu , Sh. Ma , (2014) “A hybrid multiple imputation algorithm using Gray System Theory and entropy based on clustering” Appl Intell 40:376–388, DOI 10.1007/s10489-013-0469-x, Springer Science+Business Media New York.
[5] X.Y. Zhou , J. S. Lim , (2014) “Replace Missing Values with EM algorithm based on GMM and Naïve Bayesian” International Journal of Software Engineering and Its Applications Vol.8, No.5, pp.177-188.

##plugins.themes.bootstrap3.article.details##