Dissertation > Excellent graduate degree dissertation topics show

Research on Clustering Algorithm for Source Code Mining

Author: LiuYangYang
Tutor: ZhangYang
School: Northwest University of Science and Technology
Course: Applied Computer Technology
Keywords: Data Mining Source code mining Reference point CURD algorithm NCURD algorithm Clustering
CLC: TP311.13
Type: Master's thesis
Year: 2009
Downloads: 40
Quote: 0
Read: Download Dissertation

Abstract


Software engineering data has a large amount of data, mostly for discrete attributes and characteristics, in order to software engineering, data mining and consider the efficiency of the implementation must be fast, efficient method. Cluster analysis is an important part of the data mining, has made considerable development in the field of research in recent years. Cluster analysis of software engineering data, which is an important aspect of data mining applications. A variety of clustering algorithms, the characteristics of software engineering data, how to choose an appropriate clustering algorithm is worth exploring. K-means algorithm clustering speed, but the algorithm, such as the number of clusters is difficult to determine in advance of dirty data sensitive defects such as clustering ineffective; the DBSCAN algorithm clustering effect of dirty data is not sensitive clustering time grows exponentially, not suitable for the processing of large databases. CURD clustering algorithm has a fast, efficient, but not directly to the data with discrete attribute clustering. CURD algorithm to improve clustering has discrete attribute data, and applied to the source code mining. By comparison with the K-means algorithm DBSCAN clustering results proved that the improved algorithm is very suitable for the processing of large databases, not only the clustering effect, but high efficiency. The main contents are as follows: (1) by improving the CURD algorithm proposed NCURD (Nominal-CURD) algorithm. NCURD algorithm will be introduced to the technology of the K-mode algorithm in the processing data with the discrete attribute CURD algorithm, thereby having discrete attribute data mining. The experiments show that the algorithm overcome poor clustering results of K-means algorithm and the DBSCAN clustering time into exponential growth defects in the source code mining, has good clustering effect. At the same time, also has similar to K-means linear time complexity, very suitable for the processing of large databases. (2) as a tool to the Eclipse platform, using the JAVA programming language achieve NCURD clustering algorithm, and open source data mining software WEKA source code for the test object, the first elements of the source code to extract, respectively, to obtain the class, method, class members were stored in the database; then, K-means, DBSCAN algorithm NCURD its clustering analysis; clustering results of the last three algorithms. Experimental results prove CURD algorithm improved data mining software engineering with discrete attributes, not only has good clustering effect, but also has a very high efficiency in the implementation.

Related Dissertations

  1. Research and Implementation of Mining Implicit User Interest,TP311.13
  2. Establishment and Update of Similar Users’ Cluster in Personalized Information Retrieval,TP391.3
  3. Research on Removal Algorithm of Shadows in Image Segmentation,TP391.41
  4. The Research of the Text Extraction Method Based on Spectral Cut,TP391.41
  5. A Study on Healthcare Product Marketing Based on Data Mining Technology,F426.72
  6. Gao Zhong-ying academic thought and experience and use of Bufei Decoction treatment of common diseases of the respiratory system drug law,R249.2
  7. Bing- thick academic thought and clinical experience and empirical studies apply to turtle soups treatment of chronic kidney disease,R249.2
  8. Research and Improvement on K-Means Clustering Algorithm,TP311.13
  9. Research on Peer-to-Peer Traffic Identification Algorithm Based on Cluster Analysis,TP393.02
  10. Research of Scheduling Algorithm Based on Hybrid Adaptive Genetic Algorithm in Computing Grid,TP393.09
  11. Evaluation of Photosynthetic Efficiancy of Seedlings of the Hybrid Progenies (F1) in Peach,S662.1
  12. The Load Research and Comprehensive Evaluation on the Agricultural Non-Point Source Pollution in Nantong,X592
  13. The Design and Implementation of Bicluster Data Analyzing Software,TP311.52
  14. BF-FCM Clustering Algorithm and Its Application in the Image Segmentation,TP391.41
  15. The Application of Ant Colony Algorithm in Meteorological Satellite Cloud Pictures Segmentation,TP391.41
  16. Research on Clustering Algorithm Based on Mutation Particle Swarm Optimization,TP18
  17. Research on K-means Optimization Clustering Algorithm,TP311.13
  18. Research on Fuzzy C-Mean Clustering Algorithm Based on Particle Swarm Optimization and Shuffled Frog Leaping Algorithm,TP18
  19. Research on Clustering Algorithm Based on Genetic Algorithm and Rough Set Theory,TP18
  20. Study on Photosynthetic Characteristics of Peach Based on Heterosis of Assimilation Capacity,S662.1
  21. The Research on Routing Protocol of Agricultural Environmental Monitoring System Based on Wir Eless Sensor Networks,TN915.04

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer software > Program design,software engineering > Programming > Database theory and systems
© 2012 www.DissertationTopic.Net  Mobile