Dissertation > Excellent graduate degree dissertation topics show

No amount of clustering algorithm for large data set the number of

Author: LiuChao
Tutor: WangWei
School: Nanjing Normal University
Course: Educational Technology
Keywords: Data Mining Clustering AP clustering K-Means clustering
CLC: TP311.13
Type: Master's thesis
Year: 2012
Downloads: 730
Quote: 0
Read: Download Dissertation

Abstract


With the development of computer science and technology, especially computer networks, more and more people are faced with a flood of information. In particular, due to the application of database, a large number of data have been accumulated in various fields and industries. The problem "a wealth of data and poor knowledge" is becoming increasingly prominent. In recent decades, knowledge discovery (rule extraction, data mining, machine learning, etc.) came into being receiving artificial intelligence scholars’extensive attention. A variety of different methods are proposed.The so-called data mining, data is from a large number of disordered in the discovery of hidden, effective, valuable, understandable model, and then discover useful knowledge, and come to the time trends and associations, to provide users with problem solving level decision support capabilities. At the same time, clustering as one of the main methods of data mining, more and more cause for concern. In the knowledge discovery tasks, people often have to face large amounts of data processing tasks, especially with the growing network information and the complex areas such as financial data, medical diagnostics, satellite-data and so on. We are now facing the handling of objects frequently up to millions, tens of millions. The computer’s processing power often appears to lack. Large amounts of data will bring a lot of difficulties in knowledge acquisition methods in knowledge discovery.This article describes the clustering method and its principles, and its limitations and advantages of the analysis, trying to integrate different clustering algorithms ideas, to utilization of the advantages of a particular algorithm, it not only can handle the amount of data, but also can not need to preset the number of categories, through this we can improve the clustering accuracy and reduce the clustering instability. Through theoretical analysis and experimental, the results show that the original AP algorithm can’t solve the problem of the large amount of data. Through the integration of the original AP clustering algorithm and K-Means clustering algorithm, we proposed the KMAP clustering algorithm. Through theoretical analysis and experimental, the new improved KMAP algorithm not only can solve the problem of the original AP clustering algorithm can’t handle large data, and increase its scope of application, but also resolve the K-Means clustering algorithm’s instability which caused by the order of the input dataset. Because of KMAP’s "K" value is not easy to determine, we proposed the KCAP clustering algorithm which is proposed to reduce the "K" value on the KMAP, and through this we can make the KMAP algorithm not need to preset the number of categories.

Related Dissertations

  1. Research and Implementation of Mining Implicit User Interest,TP311.13
  2. A Study on Healthcare Product Marketing Based on Data Mining Technology,F426.72
  3. Gao Zhong-ying academic thought and experience and use of Bufei Decoction treatment of common diseases of the respiratory system drug law,R249.2
  4. Bing- thick academic thought and clinical experience and empirical studies apply to turtle soups treatment of chronic kidney disease,R249.2
  5. Research and Improvement on K-Means Clustering Algorithm,TP311.13
  6. The Load Research and Comprehensive Evaluation on the Agricultural Non-Point Source Pollution in Nantong,X592
  7. The Design and Implementation of Bicluster Data Analyzing Software,TP311.52
  8. Research on Clustering Algorithm Based on Mutation Particle Swarm Optimization,TP18
  9. Research on Fuzzy C-Mean Clustering Algorithm Based on Particle Swarm Optimization and Shuffled Frog Leaping Algorithm,TP18
  10. Research on Clustering Algorithm Based on Genetic Algorithm and Rough Set Theory,TP18
  11. Based on data mining research tax audit case selection,F812.42
  12. Community-oriented education, personalized learning system and its implementation,TP391.6
  13. Association rule mining based Intrusion Detection System Research and Implementation,TP393.08
  14. Data warehouse technology in the banking customer management systems research and implementation,TP315
  15. Investments Projects Evaluation Supporting System Based on Optimizing Model of Industrial Parameters,F283
  16. The Research and Application of Data Mart in the Telecommunication Business Analysis,TP311.13
  17. Application of Data Mining in the Analysis of Higher Vocational Colleges’ Achievement,TP311.13
  18. Traffic Classification Using Stream Data Mining Algorithm,TP393.06
  19. Research and Design of Individualized Online Teaching-aided System Based on Data Mining,TP311.13
  20. Design and Implementation of Course Assessment and Analysis of Decision System Based on Data Mining,TP311.13
  21. Design to E-learning System in Senior Vocational School Base on Moodle,TP311.52

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer software > Program design,software engineering > Programming > Database theory and systems
© 2012 www.DissertationTopic.Net  Mobile