Dissertation > Excellent graduate degree dissertation topics show

Parallel K-Means clustering method and its resume data

Author: FengLiNa
Tutor: ZhouLiHua
School: Yunnan University
Course: Applied Computer Technology
Keywords: Data Mining Cluster analysis K-Means Parallel computing MPI Resume data
CLC: TP311.13
Type: Master's thesis
Year: 2010
Downloads: 85
Quote: 1
Read: Download Dissertation

Abstract


With the rapid development of technology in all areas , the amount of data is also a sharp increase . The face of the vast amounts of data , data mining technology in traditional K-Means clustering algorithm challenges , so the study of how to improve the efficiency of traditional K-Means clustering algorithm , and contribute to a better understanding and use of data . Currently, online recruitment has become the primary means of recruitment of the various business units , but every day a lot of resume screening staff inefficient , consuming serious human and material resources . The K-Means clustering algorithm in parallel the resume process data , and help save human resources , at the same time be able to help the employer more quickly to recruit the right talent . Thesis work from the following three aspects : First , in order to improve the efficiency of the traditional K-Means clustering algorithm , the paper presents a parallel K-Means clustering algorithm , the algorithm can reduce traffic and calculate the amount of a certain degree . Papers under the MPI message passing programming environment , Master / Slave programming mode of the proposed algorithm is realized . The paper based on the complexity of the algorithm and the speedup evaluation criteria to evaluate the proposed algorithm , and compared through experiments with the traditional K-Means algorithm . The experimental results show that the proposed algorithm is correct and effective . Secondly, the characteristics of resume data analysis, feature extraction method of resume data , so that the resume data can smaller dimension characteristics described . In order to verify the effectiveness of the mentioned features , the paper uses the traditional K-Means clustering algorithm of resume data clustering . Experimental results show that the extracted features can effectively reflect the resume information , resume data clustering analysis to really be able to quickly tap into the typical resume information . Finally, the proposed parallel K-Means clustering algorithm applied to the resume process data , and the clustering results are compared with the experimental results of the traditional K-Means clustering algorithm . The results show that the parallel K-Means clustering algorithm correctness and effectiveness of the resume data processing . In summary , the paper by the practical problems , the traditional K-Means clustering algorithm from both theoretical analysis and practical application , to improve the efficiency of the implementation of the algorithm , and to expand the field of application of the algorithm .

Related Dissertations

  1. Development of EST-SSR Primers and Application in Analysis of Genetic Realtionships in Tree Peony,S685.11
  2. Pre-hypertension syndrome characteristics,R259
  3. A Study on Healthcare Product Marketing Based on Data Mining Technology,F426.72
  4. Gao Zhong-ying academic thought and experience and use of Bufei Decoction treatment of common diseases of the respiratory system drug law,R249.2
  5. Bing- thick academic thought and clinical experience and empirical studies apply to turtle soups treatment of chronic kidney disease,R249.2
  6. Research and Improvement on K-Means Clustering Algorithm,TP311.13
  7. Comprehensive Quality Assessment of College Students,G645.5
  8. Studieson Effects of Soybean Species on Yuba and Initial Establishment of Quality Evalution System for Yuba,TS214.2
  9. ISSR Analysis of Genetic Diversity on 21 Lotus(Nelumbo Nucifera) Cultivars,S682.32
  10. The Research of "Ant Group" Phenomenon in the Harmonious Society,D669.5
  11. Research on the Soil Environmental Function Zoning,X321
  12. Comparison of Gene Expression Data Cluster Methods and Gene Network Construction for Phytophthora Sojae Genes,S435.651
  13. Study on Heterosis and Genetic Basis of Soybean,S565.1
  14. Evaluation on Forage Quality and Biomass Energy Characters of Inbred Vegetative Lines of Napier Grass,S543.9
  15. The Design and Implementation of Bicluster Data Analyzing Software,TP311.52
  16. BF-FCM Clustering Algorithm and Its Application in the Image Segmentation,TP391.41
  17. Research on Clustering Algorithm Based on Mutation Particle Swarm Optimization,TP18
  18. Research on K-means Optimization Clustering Algorithm,TP311.13
  19. Research on Fuzzy C-Mean Clustering Algorithm Based on Particle Swarm Optimization and Shuffled Frog Leaping Algorithm,TP18
  20. Research on Clustering Algorithm Based on Genetic Algorithm and Rough Set Theory,TP18
  21. The Study about the Select Strategies of Sportswear Brand Communication Means,G206

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer software > Program design,software engineering > Programming > Database theory and systems
© 2012 www.DissertationTopic.Net  Mobile