Dissertation > Excellent graduate degree dissertation topics show

Research and Implementation of the K-means algorithm of text clustering

Author: GaoJuan
Tutor: YuJinPing
School: Jiangxi University of Technology
Course: Applied Computer Technology
Keywords: Text Mining Clustering K-means Modified K-means Fan syncopation
CLC: TP391.1
Type: Master's thesis
Year: 2010
Downloads: 391
Quote: 2
Read: Download Dissertation

Abstract


The rapid development and wide spread of the network makes it exponentially increases the amount of information received. Text as an important carrier of information, contains many valuable resources, waiting for people to explore and study. However, the text message content and format diversity, complexity, information difficult to find interesting people. With the birth and development of text mining technology, people can quickly and effectively from a large number of text resources found simple, concise, understandable knowledge. Text clustering as an important branch of text mining, and its research is also increasingly attracted widespread attention. K-means algorithm is widely used in its simple and fast in the text clustering, but the traditional K-means algorithm depends heavily on the initial value, need to be given in advance the number of clusters to generate k, and this determination of the parameter k is generally given based on the user's experience and knowledge; addition, the initial cluster centers are selected randomly, and this randomness often leads to the instability of the clustering results. It can be said that the impact of different k values ??and different initial cluster centers on the clustering quality and time efficiency is great. Text mining and clustering analysis to do a more comprehensive study, the work carried out mainly include the following aspects: First, in-depth study of the text mining related theories and key technologies, including the Chinese word segmentation, dimensionality reduction operation, text representation the evaluation of the weights and similarity calculation. Secondly, in-depth study of traditional K-means algorithm, improved K-means algorithm for its shortcomings, in two ways: (1) fan segmentation method to determine the initial number of clusters k. Text feature item on a textual description of the degree of importance of division of the important characteristics of items, select a text set as the sample contains only important feature items, and fan syncopation its preliminary clustering, the initial number of clusters as the K- the value of the k-means algorithm. (2) starting from the own characteristics of the text mining, clustering distribution characteristics to determine the initial reverse center, based on similarity between the central object is small and the central object and other objects of the present cluster similarity between large principle, to find the most effective k text object as initial cluster centers: they lower the degree of similarity between any two, and around each center has thereto similarity higher than a certain threshold number of object. Finally, this paper designs and implements a simple text clustering system based on the K-means algorithm, and based on the validity of experiments on improved K-means algorithm validation results show improved K-means algorithm unstable clustering effect generated by the randomness effectively solve the traditional K-means algorithm, its time complexity is also decreased.

Related Dissertations

  1. Research and Implementation of Mining Implicit User Interest,TP311.13
  2. Establishment and Update of Similar Users’ Cluster in Personalized Information Retrieval,TP391.3
  3. Research on Removal Algorithm of Shadows in Image Segmentation,TP391.41
  4. The Research of the Text Extraction Method Based on Spectral Cut,TP391.41
  5. Gao Zhong-ying academic thought and experience and use of Bufei Decoction treatment of common diseases of the respiratory system drug law,R249.2
  6. Research and Improvement on K-Means Clustering Algorithm,TP311.13
  7. Research on Peer-to-Peer Traffic Identification Algorithm Based on Cluster Analysis,TP393.02
  8. Research of Scheduling Algorithm Based on Hybrid Adaptive Genetic Algorithm in Computing Grid,TP393.09
  9. Evaluation of Photosynthetic Efficiancy of Seedlings of the Hybrid Progenies (F1) in Peach,S662.1
  10. The Research of "Ant Group" Phenomenon in the Harmonious Society,D669.5
  11. The Load Research and Comprehensive Evaluation on the Agricultural Non-Point Source Pollution in Nantong,X592
  12. BF-FCM Clustering Algorithm and Its Application in the Image Segmentation,TP391.41
  13. The Application of Ant Colony Algorithm in Meteorological Satellite Cloud Pictures Segmentation,TP391.41
  14. Research on Clustering Algorithm Based on Mutation Particle Swarm Optimization,TP18
  15. Research on K-means Optimization Clustering Algorithm,TP311.13
  16. Research on Fuzzy C-Mean Clustering Algorithm Based on Particle Swarm Optimization and Shuffled Frog Leaping Algorithm,TP18
  17. Research on Clustering Algorithm Based on Genetic Algorithm and Rough Set Theory,TP18
  18. The Study about the Select Strategies of Sportswear Brand Communication Means,G206
  19. Study on Photosynthetic Characteristics of Peach Based on Heterosis of Assimilation Capacity,S662.1
  20. A Snoring Detector for OSAHS Based on Formant,R766
  21. The Research on Routing Protocol of Agricultural Environmental Monitoring System Based on Wir Eless Sensor Networks,TN915.04

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net  Mobile