Dissertation > Excellent graduate degree dissertation topics show

The Class-Mean Method and Its Extensions to Handling Incomplete Data in Data Mining

Author: JiYanXia
Tutor: MaJiangHong
School: Chang'an University
Course: Applied Mathematics
Keywords: Data Mining Incomplete Data Class mean imputation Weighted Adjustment Act Fuzzy techniques Membership
CLC: TP311.13
Type: Master's thesis
Year: 2010
Downloads: 144
Quote: 1
Read: Download Dissertation

Abstract


With the increasingly high demand for the continuous development of the data collection and storage technologies as well as access to information, data mining is an effective tool that has become indispensable to many industries. In the entire process of data mining, data preprocessing is undoubtedly a very critical step, because only the use of reliable, accurate data in order to obtain satisfactory mining results. According to incomplete statistics, data pre-processing account for about 60% of the workload of the entire mining process. Incomplete Data is the most basic problem encountered in the data preprocessing Moreover, few real-world database is complete, no missing, so, how to deal with the missing incomplete data has become a domestic and international concern to the hot research problem solved. Incomplete data processing methods can be divided into the deleted tuple, filled missing data, and does not handle three categories. Han and Zhang, from the use of the frequency and extent of research and other aspects of view, the most commonly used method of handling missing values ??data filling method, this method is both technically and in theory should be well attention and study. Incomplete data for data mining research can be a simple and effective approach filling in missing data. First of all, on the common missing data to fill the system summary and comparison Secondly, for the lack of class mean method, efforts be improved proposed the weighted the class mean adjustment method based on fuzzy class mean imputation. The main research work are as follows: 1, through the introduction of data mining knowledge and elaborate mining process, data preprocessing is one of the most critical aspects of the data mining and stages; incomplete data in the stage processing problems detailed discussion, including its causes, prevention beforehand, post-processing, etc.; principle and pointed out that the advantages and disadvantages of various methods of analysis of several commonly used afterwards padding method of the various methods. 2, for simple fill strategy commonly used in the class mean France main drawback of our improved method: for each group gives different weights to the weighted adjustment makes fill results as close to the true value, and given different weights weight can make each objects to fill the values ??are not the same, and ultimately achieve the purpose of the degree of variability improve variable. In addition, for the weighted adjusted Law rights re compare subjective and is proposed based on fuzzy class mean imputation, this method can overcome the two drawbacks. Three methods using the R software simulation by experimental results verify the feasibility and effectiveness of the methods mentioned two improved.

Related Dissertations

  1. A Study on Healthcare Product Marketing Based on Data Mining Technology,F426.72
  2. Gao Zhong-ying academic thought and experience and use of Bufei Decoction treatment of common diseases of the respiratory system drug law,R249.2
  3. Bing- thick academic thought and clinical experience and empirical studies apply to turtle soups treatment of chronic kidney disease,R249.2
  4. The Design and Implementation of Bicluster Data Analyzing Software,TP311.52
  5. Research on Clustering Algorithm Based on Mutation Particle Swarm Optimization,TP18
  6. Research on Fuzzy C-Mean Clustering Algorithm Based on Particle Swarm Optimization and Shuffled Frog Leaping Algorithm,TP18
  7. Research on Clustering Algorithm Based on Genetic Algorithm and Rough Set Theory,TP18
  8. Based on data mining research tax audit case selection,F812.42
  9. Community-oriented education, personalized learning system and its implementation,TP391.6
  10. Association rule mining based Intrusion Detection System Research and Implementation,TP393.08
  11. Data warehouse technology in the banking customer management systems research and implementation,TP315
  12. Design to E-learning System in Senior Vocational School Base on Moodle,TP311.52
  13. Design and Development of Teaching Quality Assessment System Based on Data Mining,TP311.13
  14. The Application of Association Rules Algorithm in Higher Vocational Colleges’ Endorsement of Impoverished Students,G717
  15. Based on Data Mining Technologies in Urban Water Supply Analysis and Decision,F299.24;F224
  16. Research on Application of Data Mining Technology in Degree of Satisfaction Analysis of Television Customers,TP311.13
  17. Web Usage Mining and the Research of Personalized Recommendation,TP311.13
  18. Data Mining of Application in the School Management and Training Students,TP311.13
  19. Research on Employment Monitoring System of University Graduate,G647.38
  20. Design and Implementation for Decision Support System of Drug Administration Based on Data Warehouse,TP311.13
  21. A Research on the Credit Card Client Activating and Response Extent Based on Data Mining,F832.2

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer software > Program design,software engineering > Programming > Database theory and systems
© 2012 www.DissertationTopic.Net  Mobile