Dissertation > Excellent graduate degree dissertation topics show

Research on High-dimensional Data Clustering Based on Genetic Algorithm

Author: XiongLangHuan
Tutor: SunHaoJun
School: Shantou University
Course: Applied Computer Technology
Keywords: cluster analysis genetic algorithm high-dimensional data feature subspace
CLC: TP311.13
Type: Master's thesis
Year: 2010
Downloads: 128
Quote: 1
Read: Download Dissertation

Abstract


Data Mining is a hot research area in information technology industry, and cluster analysis is one of the most important research topics in this area. Clustering is the process of grouping data into a number of clusters according to a similarity metric, which has a wide range of applications in real word. Nowadays, there are many classical clustering algorithms which work well on low-dimensional data, while those algorithms are often invalid when processing high-dimensional data because of“the curse of dimensionality”. However, the data is always with high dimensions in the real applications. For examples, gene expression data, finance data, multimedia data and web data. The universality of high-dimensional data makes it very important to research on clustering algorithms for high-dimensional data.The direct approach for high-dimensional data clustering is feature transformation which transforms the high dimensional space into low dimensional space. After that, the traditional clustering algorithms can be used to solve the problem. In high dimensional data space, not all of the dimensions are related to clustering. In order to find the most appropriate feature subspace, all the feature subsets need to be tested, while this is a very large cost in computing for high-dimensional data. The traditional search algorithms like greedy algorithms may only find optimal solutions; in this study, genetic algorithms (GA) is used for searching the feature subspace. Genetic algorithms are adaptive heuristic search algorithm premised on the evolutionary ideas of natural selection and natural genetics. In our approach, the searching capability of GA is exploited to search for appropriate feature subsets for clustering, meanwhile, in order to illustrate the characteristics of features (or dimensions) shown in clustering, a fitness function which is based on the degree of features contribute to subspace clustering is proposed. Those researches in this thesis have some theoretical and practical significance.The main contributions of the study are summarized as following: (1) determine the searching space and encoding method. The traditional encoding method using GA for clustering focus on cluster centers, while in this study, the encoding space is made up of feature subspace and cluster centers and attach some conditions to limit the length of encoding string; (2) propose a new fitness function which is based on the degree of features contribute to subspace clustering. As the evaluation function of subspace clustering, it has ability to compare the clustering result of different subspaces, that is evaluate the clustering result and the features included in such subspace at the same time; (3) design and implement a high-dimensional data clustering using genetic algorithms, called GA-HDclustering; (4) the experiments on an artificial data set generated by computer and real-life data sets got from UCI and literature of Brian Tjaden indicate the feasibility and efficiency of GA-HDclustering.

Related Dissertations

  1. Development of the Platform for Compressor Optimization Design and Aerodynamic Optimization Design in the Transonic Compressor,TH45
  2. Development of EST-SSR Primers and Application in Analysis of Genetic Realtionships in Tree Peony,S685.11
  3. The Application of Fuzzy Comprehensive Evaluation Based on Genetic Algorithm in Vocational Evaluation of Classroom Teaching,G712
  4. Comprehensive Quality Assessment of College Students,G645.5
  5. Study on Taste Characteristic of Taste Peptide Enzymatic Production from Oyster Base on A Neural Network Method,TS254.4
  6. Studieson Effects of Soybean Species on Yuba and Initial Establishment of Quality Evalution System for Yuba,TS214.2
  7. ISSR Analysis of Genetic Diversity on 21 Lotus(Nelumbo Nucifera) Cultivars,S682.32
  8. Design and Realization of the Magnetic Antenna in MW and SW Bands Based on Genetic Algorithm,TN820
  9. Research of Scheduling Algorithm Based on Hybrid Adaptive Genetic Algorithm in Computing Grid,TP393.09
  10. Public Transport Optimal Dispatching Based on the Genetic-Newton Algorithm,TP18
  11. BP network optimization based on genetic algorithm optimization of the biodiesel process,TE667
  12. Research on the Soil Environmental Function Zoning,X321
  13. Study on Heterosis and Genetic Basis of Soybean,S565.1
  14. Evaluation on Forage Quality and Biomass Energy Characters of Inbred Vegetative Lines of Napier Grass,S543.9
  15. Research on Clustering Algorithm Based on Genetic Algorithm and Rough Set Theory,TP18
  16. Community-oriented education, personalized learning system and its implementation,TP391.6
  17. Mining resources based on genetic algorithm optimization model of,O224
  18. The magnetorheological damper mechanical properties and Gun Recoil,TB535.1
  19. Optimization Study on Gating System and Molding Process Parameters of Injection Mold Based on Simulation,TQ320.662
  20. Research on the Milling Performance and Parameters Optimization with Large Parts of Heavy Machine,TG54
  21. Design and Realization of Automatic Course Arrangement System Based on Genetic Algorithm,TP311.52

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer software > Program design,software engineering > Programming > Database theory and systems
© 2012 www.DissertationTopic.Net  Mobile