Dissertation > Excellent graduate degree dissertation topics show
Grid Clustering Research Based on Impact Factor of Grid Density
Author: YangBin
Tutor: LiZhong
School: Zhejiang University of Technology
Course: Basic mathematics
Keywords: data mining clustering grid density impact factor of grid density K-means
CLC: TP301.6
Type: Master's thesis
Year: 2014
Downloads: 3
Quote: 0
Read: Download Dissertation
Abstract
|
Data mining is generated by the needing of extracting valuable information fromhuge amounts of data information. As one of the important branch of data mining,clustering analysis can automatically identify classes that composed by similar datapoints that have not been labeled. Among various kinds of clustering algorithms, theclustering algorithm based on density can identify classes that have different densityand shape. However, clustering based on density often needs to set the globalparameters, and more than one other parameter. When the density of classes changesstrongly, the cluster algorithm will encounter difficulties. The clustering algorithmbased on grid calculates the information of mesh instead of data points, so it improvesthe processing speed. But the algorithm increases the speed at the expense of reducingprecision. Mesh size influences the quality of clustering, the smaller the particle size,the more precise clustering, but the computation cost is higher, and the greater theparticle size, the more crude of the quality of clustering.According to these disadvantages of clustering algorithms based on density andgrid, we consider the impact factor of grid density in this paper, propose an improvedclustering algorithm: Impact Factor of Grid Density based Clustering algorithm(IFGDC). The algorithm’s main work includes:(1)clustering grid unit instead of datapoints by partitioning the data space to reduce the complexity of clustering operationeffectively;(2)defining some concepts based on grid adjacent relations to avoid theinconvenience of determining the radius in traditional algorithm based on density;(3)putting forward the concept: impact factor of grid density to determine the core cellfrom the high-density grid mesh;(4)giving a method to extract clustering edgeboundary point and further to improve the clustering accuracy. Finally, we test IFGDCclustering algorithm through some experiments, verify the correctness and validity ofthe algorithm.K-means clustering algorithm is simple and it becomes a classical clusteringalgorithm. However K-means clustering is sensitive to parameters, it relies on theuser’s experience to select cluster number and initial cluster centers; also, the algorithm is susceptible to noise interference, and it’s results depend on the order ofdata input. Considering improving these deficiencies, this paper presents an improvedK-means algorithm based on IFGDC. Firstly, the new algorithm uses the speedadvantage by grid-clustering to preprocess data points to find the general structure anddistribution of the data sets, obtain clustering number k and initial mass center of theclusters. Then we use these two parameters for K-means clustering. Compared withsimply using K-means clustering directly, the experiment results show that theimproved algorithm can improve the quality of the parameter k and initial center ofmass. The new algorithm can reduce the sensitivity of the "noise", the cluster resultsare certain, and do not rely on the order of data entry. So the new algorithm canimprove the clustering effect effectively.The main content is summarized, and the future work is prospected at the end ofthis paper.
|
Related Dissertations
- Research on Graph-Based Algorithm for Tagsnps Selection,Q78
- Research on Mud Density Measuring Method Using Ultrasonic,TE256.7
- Numerical Analysis of Flow Process in a Molten Carbonate Fuel Cell,TM911.4
- Grid-Side Converter Control and Wind Turbine Emulator in Direct Drive Wind Power System,TM46
- The Effect of Current Density and Electric Field on Ni and Its Alloy Electrodeposition for Copper Crystallizer,TQ153.2
- The Optimization of Design for BICM System Based on Irregular LDPC Codes,TN911.2
- BioLab a Bioinformatics Oriented Grid Portal,TP399-C8
- Research and Implementation of Mining Implicit User Interest,TP311.13
- Establishment and Update of Similar Users’ Cluster in Personalized Information Retrieval,TP391.3
- Research on Removal Algorithm of Shadows in Image Segmentation,TP391.41
- The Research of the Text Extraction Method Based on Spectral Cut,TP391.41
- Design and Realize of Family Cleaning Robot Path-Coverage System,TP242
- A Study on Healthcare Product Marketing Based on Data Mining Technology,F426.72
- Gao Zhong-ying academic thought and experience and use of Bufei Decoction treatment of common diseases of the respiratory system drug law,R249.2
- Bing- thick academic thought and clinical experience and empirical studies apply to turtle soups treatment of chronic kidney disease,R249.2
- Research and Improvement on K-Means Clustering Algorithm,TP311.13
- Study on Preparation and Properties of Liquid-Core Capsule and Application in Continuous Inoculation of Milk,TS252.4
- Micro- grid with distributed power control strategy research,TM61
- The Grid-Connected Wind-solar Hybrid Generation System and Maximum Power Point Tracking,TM61
- Effect of Planting Densities and Modes on Lodging Resistance of Stem in Summer Maize,S513
- Research on Peer-to-Peer Traffic Identification Algorithm Based on Cluster Analysis,TP393.02
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > General issues > Theories, methods > Algorithm Theory
© 2012 www.DissertationTopic.Net Mobile
|