Dissertation > Excellent graduate degree dissertation topics show

Research on the Improvement of C-means Clustering Algorithm

Author: WangHui
Tutor: ShenShiLei
School: Henan University
Course: Applied Computer Technology
Keywords: Clustering Analysis K-means Algorithm Fuzzy C-means Algorithm Feature Weighting
CLC: TP311.13
Type: Master's thesis
Year: 2011
Downloads: 281
Quote: 3
Read: Download Dissertation

Abstract


As a non-supervised learning method, clustering analysis is an important analysis tool and method for data processing in data mining. The purpose of clustering is to partition a data set without any category tags into several clusters according to some rule, where the objects within the same cluster have high similarity, but the objects between clusters are very dissimilar. Clustering has been widely applied in many fields such as commerce, finance, image processing, information retrieve and so forth. Clustering analysis mainly focuses on how to design the practical clustering algorithm with better performance.C-means clustering algorithm is a typical partition method, including K-means (hard C-means) algorithm and Fuzzy C-means (FCM) algorithm. This algorithm is simple to implement, fast and effective. However, it has its own inherent disadvantages: the number of clusters should be given before clustering which is difficult for the inexperienced users; it always cannot get the global optimum value because of easily trapping into a local optimum; the choice of the initial centers has a great influence on the clustering result; the algorithm is sensitive to the isolated points and noise, etc. Optimizing deeply clustering algorithms will not only help to perfect its theory, but also its popularization and application.In this paper, we study intensively the classic K-means clustering algorithm and fuzzy C-means clustering algorithm, and sum up their strengths and weaknesses. To overcome the shortcomings such as the sensitivity to the initial clustering centers, the bigger computational cost etc., based on K-means algorithm and fuzzy C-means algorithm, two improved clustering algorithms are proposed in this paper. The researches and contributions can be summed up in two aspects as follows:(1) The traditional K-means algorithm sensitive to the initial clustering center, the isolated points and noise leads to the instability and the low accuracy clustering results. For the purpose of eliminating these shortcomings, based on attribute selection and feature weighting methods, an improved weighted K-means clustering algorithm is proposed. In the presented algorithm, firstly we preprocess the data set and the original clustering center is obtained according to the data sample distribution. Then the improved K-means algorithm using the feature weights is designed. Experiment results have shown that the proposed clustering algorithm can produce high quality clustering steadily. (2) The existing studies have proved that the first component of the color feature set can effectively describe the color pixels feature. To reduce the computational cost, the first component of color feature set is chosen as the one-dimensional eigenvector. It is used as the pixels in the image segmentation method based on the fuzzy C-means clustering algorithm. Based on the rough set theory, the cluster number and the initial cluster centers are obtained. Feature distance, which is suitable for any structure of eigenvector space, is used instead of Euclidian distance to overcome the influence caused by the structure of eigenvector space. Then an improved fuzzy C-means clustering algorithm is introduced to cluster the sample data. Experimental results show that the presented image segmentation method can effectively improve the precision and accuracy of image segmentation, and has small computational cost and fast convergence speed.

Related Dissertations

  1. Research and Improvement on K-Means Clustering Algorithm,TP311.13
  2. Evaluation of Photosynthetic Efficiancy of Seedlings of the Hybrid Progenies (F1) in Peach,S662.1
  3. The Load Research and Comprehensive Evaluation on the Agricultural Non-Point Source Pollution in Nantong,X592
  4. Research on K-means Optimization Clustering Algorithm,TP311.13
  5. Study on Photosynthetic Characteristics of Peach Based on Heterosis of Assimilation Capacity,S662.1
  6. Web Usage Mining and the Research of Personalized Recommendation,TP311.13
  7. Term Weight-Based Chinese Text Classification Algorithm,TP391.1
  8. Improved Fuzzy C-Means Algorithm and Its Application in Image Segmentation,TP391.41
  9. The Application of Clustering Analysis and Classification Analysis in Rail Transit System,U239.5
  10. Research of State Estimation in Power Network Based on Data Mining Technology,TM769
  11. Research on Grey Relation Analysis, Grey Clustering Analysis and Grey System Forecasting and Application in Medicine,R195
  12. Fuzzy clustering method of power system bad data identification,TM76
  13. Research on Classification of Raw Materials of Pulp Combined with Near-infrared Spectroscopy Technology and Pattern Recognition,TN219
  14. Information Theoretic Feature Weighting and Topic-Driven Co-Clustering for Text Dataset,TP18
  15. The Research of Data Mining Application in Tax System,TP311.52
  16. Short-Term Load Forecasting Based on Fuzzy Clustering Analysis and Least Squares Support Vector Machines,TM715
  17. Research on the Safety Evaluation of Transporting Dangerous on Inland Water on Non-accident in Chongqing,U698.3
  18. Demonstration Analysis of Herdsmen’s Consumption Structure in Inner Mongolia and Its Influence Factors,F327
  19. The Related Study on Modern Logistics Industry of the Development in Taizhou,F259.2
  20. Application Research and Implementation of Dss for Chained-retailing Based on Data Warehouse,F721.7
  21. Simultaneous Confidence Intervals for Multivariate Binomial Proportions,O211.3

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer software > Program design,software engineering > Programming > Database theory and systems
© 2012 www.DissertationTopic.Net  Mobile