Dissertation > Excellent graduate degree dissertation topics show
Parallel K-Means clustering method and its resume data
Author: FengLiNa
Tutor: ZhouLiHua
School: Yunnan University
Course: Applied Computer Technology
Keywords: Data Mining Cluster analysis K-Means Parallel computing MPI Resume data
CLC: TP311.13
Type: Master's thesis
Year: 2010
Downloads: 85
Quote: 1
Read: Download Dissertation
Abstract
|
With the rapid development of technology in all areas , the amount of data is also a sharp increase . The face of the vast amounts of data , data mining technology in traditional K-Means clustering algorithm challenges , so the study of how to improve the efficiency of traditional K-Means clustering algorithm , and contribute to a better understanding and use of data . Currently, online recruitment has become the primary means of recruitment of the various business units , but every day a lot of resume screening staff inefficient , consuming serious human and material resources . The K-Means clustering algorithm in parallel the resume process data , and help save human resources , at the same time be able to help the employer more quickly to recruit the right talent . Thesis work from the following three aspects : First , in order to improve the efficiency of the traditional K-Means clustering algorithm , the paper presents a parallel K-Means clustering algorithm , the algorithm can reduce traffic and calculate the amount of a certain degree . Papers under the MPI message passing programming environment , Master / Slave programming mode of the proposed algorithm is realized . The paper based on the complexity of the algorithm and the speedup evaluation criteria to evaluate the proposed algorithm , and compared through experiments with the traditional K-Means algorithm . The experimental results show that the proposed algorithm is correct and effective . Secondly, the characteristics of resume data analysis, feature extraction method of resume data , so that the resume data can smaller dimension characteristics described . In order to verify the effectiveness of the mentioned features , the paper uses the traditional K-Means clustering algorithm of resume data clustering . Experimental results show that the extracted features can effectively reflect the resume information , resume data clustering analysis to really be able to quickly tap into the typical resume information . Finally, the proposed parallel K-Means clustering algorithm applied to the resume process data , and the clustering results are compared with the experimental results of the traditional K-Means clustering algorithm . The results show that the parallel K-Means clustering algorithm correctness and effectiveness of the resume data processing . In summary , the paper by the practical problems , the traditional K-Means clustering algorithm from both theoretical analysis and practical application , to improve the efficiency of the implementation of the algorithm , and to expand the field of application of the algorithm .
|
Related Dissertations
- Development of EST-SSR Primers and Application in Analysis of Genetic Realtionships in Tree Peony,S685.11
- Pre-hypertension syndrome characteristics,R259
- A Study on Healthcare Product Marketing Based on Data Mining Technology,F426.72
- Gao Zhong-ying academic thought and experience and use of Bufei Decoction treatment of common diseases of the respiratory system drug law,R249.2
- Bing- thick academic thought and clinical experience and empirical studies apply to turtle soups treatment of chronic kidney disease,R249.2
- Research and Improvement on K-Means Clustering Algorithm,TP311.13
- Comprehensive Quality Assessment of College Students,G645.5
- Studieson Effects of Soybean Species on Yuba and Initial Establishment of Quality Evalution System for Yuba,TS214.2
- ISSR Analysis of Genetic Diversity on 21 Lotus(Nelumbo Nucifera) Cultivars,S682.32
- The Research of "Ant Group" Phenomenon in the Harmonious Society,D669.5
- Research on the Soil Environmental Function Zoning,X321
- Comparison of Gene Expression Data Cluster Methods and Gene Network Construction for Phytophthora Sojae Genes,S435.651
- Study on Heterosis and Genetic Basis of Soybean,S565.1
- Evaluation on Forage Quality and Biomass Energy Characters of Inbred Vegetative Lines of Napier Grass,S543.9
- The Design and Implementation of Bicluster Data Analyzing Software,TP311.52
- BF-FCM Clustering Algorithm and Its Application in the Image Segmentation,TP391.41
- Research on Clustering Algorithm Based on Mutation Particle Swarm Optimization,TP18
- Research on K-means Optimization Clustering Algorithm,TP311.13
- Research on Fuzzy C-Mean Clustering Algorithm Based on Particle Swarm Optimization and Shuffled Frog Leaping Algorithm,TP18
- Research on Clustering Algorithm Based on Genetic Algorithm and Rough Set Theory,TP18
- The Study about the Select Strategies of Sportswear Brand Communication Means,G206
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer software > Program design,software engineering > Programming > Database theory and systems
© 2012 www.DissertationTopic.Net Mobile
|