Dissertation > Excellent graduate degree dissertation topics show

Research of Clustering Algorithm Based on Relative Density

Author: LiXiaoChang
Tutor: HuangDeCai
School: Zhejiang University of Technology
Course: Applied Computer Technology
Keywords: relative density cluster analysis mixture data incremental clustering single update mode batch update mode
CLC: TP311.13
Type: Master's thesis
Year: 2012
Downloads: 58
Quote: 0
Read: Download Dissertation

Abstract


Cluster analysis is a hot topic in the research field of data mining; many existed proposed clustering algorithms are suitable for different application environments according to their own different characteristics. Among them, the traditional density-based clustering algorithm has been widely used for its good scalability, noise immunity and the ability to discover clusters of arbitrary shape. However it adopts global density parameters, also called absolute density as the metrics of all clusters, so the algorithm has a vital weakness that it’s lack of ability to find clusters of different density levels. Adopting the relative density as the metrics of clusters can be a good solution to this problem, so the relative density–based clustering algorithm has been presented.Meantime, the data which need to be clustered usually is dynamic in practical application, when the data changes, the original clustering pattern dug out should also be updated. Therefore, that how to design the incremental clustering algorithm for saving compute resources and improving clustering efficiency has become an important challenge of current cluster analysis.Firstly, the paper describes the related knowledge of clustering, and introduces the common used concepts of data mining such as the definition of clustering, similarity measurement, density, etc. The paper also conducts the review of traditional clustering algorithms, then gives their classification and performance comparison.Secondly, after analyzing the defects of traditional density-based clustering algorithm, a Relative Density-Based Clustering Algorithm for Mixture Data Sets(M_RDBCA) has been presented. The algorithm overcomes the defects of traditional density-based algorithm while maintaining its advantages which include the ability to find clusters of arbitrary shape and insensitive to the noise. Since the algorithm defines the distance of mixture data in the purpose of measuring similarity of mixed-attribute objects, and introduces the concept of pure neighbor, the algorithm is not only suitable for numerical data, but also can be applied to categorical data and mixture data. The algorithm can distinguish clusters of different density levels because it adopts relative density as clustering metrics, and it proposes a concept of pure core set of objects, so the objects in one cluster can be better integrated. In addition, the paper provides a theory for setting the parameter, so it avoids the problem of DBSCAN that the clustering results is too sensitive to algorithm parameters.At last, we performed a deep and comprehensive research on incremental clustering methods based on M_RDBCA in data warehouse environment. First, the paper described the data model and basic ideas of incremental clustering algorithm. Second, we discussed incremental clustering methods based on M_RDBCA of three different manipulation mode:①Clustering the affected set: after ascertaining the objects which can be affected by the insertion or deletion of an object, the incremental clustering algorithm only be performed on the affected set.②Single update mode: according to the impact on clustering caused by inserting or deleting an object, some clustering operation such as merge, split, or absorb will be conducted. An experiment is carried out for performance comparison, and performance speed-up chart is drew.②Batch update mode: briefly states the idea and method of incremental clustering in the batch mode which analyze the impact on clustering not only focus on existing objects in database but also consider the inserting and deleting objects in the update table.

Related Dissertations

  1. Development of EST-SSR Primers and Application in Analysis of Genetic Realtionships in Tree Peony,S685.11
  2. Pre-hypertension syndrome characteristics,R259
  3. Comprehensive Quality Assessment of College Students,G645.5
  4. Studieson Effects of Soybean Species on Yuba and Initial Establishment of Quality Evalution System for Yuba,TS214.2
  5. ISSR Analysis of Genetic Diversity on 21 Lotus(Nelumbo Nucifera) Cultivars,S682.32
  6. Research on the Soil Environmental Function Zoning,X321
  7. Comparison of Gene Expression Data Cluster Methods and Gene Network Construction for Phytophthora Sojae Genes,S435.651
  8. Study on Heterosis and Genetic Basis of Soybean,S565.1
  9. Evaluation on Forage Quality and Biomass Energy Characters of Inbred Vegetative Lines of Napier Grass,S543.9
  10. Clustering Method Research Based on Divided and Conquered Method,TP311.13
  11. Elements of Thirty Kinds of Proprietary Chinese Medicines and Classification Based on Elements,R286.0
  12. Community-oriented education, personalized learning system and its implementation,TP391.6
  13. Design to E-learning System in Senior Vocational School Base on Moodle,TP311.52
  14. Research on Cultural Industrial Competitiveness of Chong Qing,F224
  15. The Study of Competitiveness of Coastal Ports in China Based on Principal Component Analysis,F552
  16. Joint analysis method based on Chinese industry, consumer preferences study Inn,F719
  17. Studies on the Mesozooplankton and Microzooplankton Community in the Yellow River Estuary and its Adjacent Area in Summer and Autumn,Q958.8
  18. Application Research in Crm of Bank Based on the Improved K-means Cluster Algorithm,TP311.13
  19. Study on the Multi-fiber Combination and the Wearability and Style of the Woven Fabrics,TS941.15
  20. The Research on the Automatic Generation of Concept Weights and Optimal Composition for Test Item in E-learning System,TP391.6
  21. The Evaluation Study of High-tech Industrial Competitiveness in Jiangsu Province,F224

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer software > Program design,software engineering > Programming > Database theory and systems
© 2012 www.DissertationTopic.Net  Mobile