Dissertation > Excellent graduate degree dissertation topics show
Cluster analysis algorithm in data mining
Author: WangZuo
Tutor: LiuXiYu;WangHongGuo
School: Shandong Normal University
Course: Computer Software and Theory
Keywords: Data Mining Clustering Analysis data partitioning nearest neighbor first
CLC: TP311.13
Type: Master's thesis
Year: 2006
Downloads: 703
Quote: 9
Read: Download Dissertation
Abstract
|
With the wide usage of information technology, data generated from differentinformation systems become more and more. How to utilize the huge original data toanalyse current situation and predict future of quantities effectively, have alreadybecome a great challenge that the mankind has faced. Therefore the data miningtechnology is arised at the historic moment and can be developed rapidly, which isattributed to the necessary consequence of the conflicting movement between therapid increasing data and the poor information day by day. Data Mining, also called as knowledge discovery of databases (KDD), is aprocessing procedure of extracting credible, novel, effective and understandablepatterns from databases. Data Mining is a relatively young research and applicationarea based on database techniques, which synthesizes multidisciplinary productions,such as logic statistics, machine learning, fuzzy theory and visual computing, in orderto acquire usable information from database .It has achieved increasing attention inthe past years ,and has been applied to finance, insurance, communalfacilities, government, education, telecommunication, software developmentof the bank, transporting, etc.Cluster analysis is an important technology in data mining. Clustering, anunsupervised classifying method, is the process of grouping together similarmulti-dimensional data vectors into a number of clusters or bins. Clustering processesare always carried out in the condition with no pre-known knowledge, so the mostresearch task is to solve that how to get the clustering result in this premises. The mostresearch about clustering is focused on clustering algorithms, the main purpose is toproduce practical algorithms with better performance. Up to now, many clusteringalgorithms have been presented, but these algorithms are only suited special problemsand users. Furthermore, they are imperfect both theoretically and methodologically,even severe fault. Optimizing deeply clustering algorithms will not only help toperfect its theory, but also its popularization and application. This dissertation systematically, deeply, roundly and detailedly studies andanalyses the data mining technique, especially clustering analysis. The main contentsare listed as follows:(1) Description in brief of the data mining technique. This paper introduces thebasic concept, classes, main functions, key technique and typical applications.(2)Research Of data mining Tools .The paper introduces the common tools ofdata mining, Analyses and compares the functions of overseas leading data miningtools(SPSS and DBMiner)based on the real examples and the conclusion is given inthe paper.(3) Description in brief of clustering analysis. The paper analyses the clusteringmethods and the representative clustering algorithms, puts forward the typicalrequests of clustering and compares the common clustering algorithm, so that peoplecan easily find a clustering method that suits a special problem.(4) An improved algorithm of K-Means is proposed. This paper analyses theNearest Neighbors Absorbed First (NNAF) clustering algorithm. This algorithm cancluster quickly with noisy . However, clustering quality will degrade when the clusterdensity and distance between clusters are not even. In this paper aNearest-Neighbors-First clustering algorithm based on data partitioning is proposed.The analysis on theory and experiments show that the new algorithm can wellimproves the the quality of clustering.
|
Related Dissertations
- A Study on Healthcare Product Marketing Based on Data Mining Technology,F426.72
- Gao Zhong-ying academic thought and experience and use of Bufei Decoction treatment of common diseases of the respiratory system drug law,R249.2
- Bing- thick academic thought and clinical experience and empirical studies apply to turtle soups treatment of chronic kidney disease,R249.2
- Evaluation of Photosynthetic Efficiancy of Seedlings of the Hybrid Progenies (F1) in Peach,S662.1
- The Design and Implementation of Bicluster Data Analyzing Software,TP311.52
- Research on Fuzzy C-Mean Clustering Algorithm Based on Particle Swarm Optimization and Shuffled Frog Leaping Algorithm,TP18
- Study on Photosynthetic Characteristics of Peach Based on Heterosis of Assimilation Capacity,S662.1
- Hebei Ecological Garden City Evaluation System and Applied Research,TU986
- Study on Privacy Preserving Classification Data Mining,TP311.13
- The Research and Development of Reservoir Decision Support System,TP311.52
- Research and Implementation of Data Mining Based on Distributed Computing,TP311.13
- Study on Data Mining Application in Short-term Load Forecasting,TM715
- Study and Implementation of Human Resource Statistic Analysis System Based on BI Technology,TP311.52
- The Application of Association Rules on Load Variation Mining,TM715
- Data warehouse and data mining, alloy production control,TP311.13
- The Research of Methods for Well Logging Evaluation of Complex Reservoir in Biyang Depression,P618.13
- Research on Peer-to-Peer Traffic Identification Algorithm Based on Cluster Analysis,TP393.02
- The uncertainty of the data mining algorithm design,TP311.13
- Analysis and Prediction of Corporation Power Consumption Based on Support Vector Machines,TM714
- The Research and Implementation of Discriminant Analysis Based on the Bayes Theory,TP311.13
- Research on the Transformation and Application of the Individual Out-trans-carried Tacit Knowledge,TP311.13
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer software > Program design,software engineering > Programming > Database theory and systems
© 2012 www.DissertationTopic.Net Mobile
|