Dissertation > Excellent graduate degree dissertation topics show
SVM-Based Data Mining Technology Research
Author: LiuSenHua
Tutor: CuiGuangCai
School: Changchun University of Science and Technology
Course: Computer Software and Theory
Keywords: Machine Learning Statistical Learning Theory Data Mining Support Vector Machine Incremental Learning
CLC: TP311.13
Type: Master's thesis
Year: 2009
Downloads: 244
Quote: 2
Read: Download Dissertation
Abstract
|
With the rapid development of information technology all over the world, the amount of information also has improved with the geometric exponential growth, how to utilize the data mining technology to obtain the useful information from the massive and complex data has become the focus in the field of the research on the information technology in recent years. As a new technology Data Mining involves the characteristics of many disciplines, such as database, integrating artificial intelligence, mathematical statistics and so on. Among the aspects of Data Mining, Classification Mining is the most important and the most common data mining task.As a new general learning machine, Support Vector Machine is developed on the basis of statistical learning theory. The development of Support Vector Machine has become the research focus of data classification mining techniques because of its excellent theory (including VC dimension, Structural risk minimization and kernel space theory). In order to solve a complicated classification task, the vectors are mapped from input space to feature space in which a linear separating hyperplane is structured. As a structure risk minimized implement, Vector Machine has the advantages of global optimization, simple structure and high practicability.The article firstly introduce the basic concepts and technology of data mining, discusses and analyzes on the basis of the theory, basic concepts and basic algorithm of Support Vector Machine, as well as all kinds of algorithms on the support vector machine at present.During the applications, data mining technology is often dealt with large amounts of data or incremental data. The article is focus on the fast learning methods of the support vector machine and the incremental learning strategies while dealing with large-scale data. Since the training speed of learning algorithms of the general SVM can be increased, since the processing of historical data is too rough and cannot adapt to the different circumstances that vary with the data distribution of new samples concentration which leads to the injured accuracy, and since it is difficult to choose the penalty Component C of learning algorithms of the general SVM optimization problem C-SVM and it is difficult to adapt to the different circumstances that vary with the data distribution of new samples concentration, we propose a improved Bv-SVM incremental learning method based on self-tuning parameters of the two pre-check border vector. The UCI database experiments show that the method is significantly better than the learning algorithms of the standard C-SVM and simple incremental support vector machine in the training time, which is also better than the learning algorithm of a simple incremental support vector machine on the training accuracy. If the training data is much less, the training accuracy is similar with standard C-SVM, but as the increase of training data, training accuracy is beyond the standard C-SVM gradually, which shows that the method is very suitable for incremental learning of large-scale data sets.
|
Related Dissertations
- Research on Automatic Detection Algorithm for Substructure Distress of Highway Pavement Based on SVM,U418.6
- Research on Transductive Support Vector Machine and Its Application in Image Retrieval,TP391.41
- A Study on Healthcare Product Marketing Based on Data Mining Technology,F426.72
- Gao Zhong-ying academic thought and experience and use of Bufei Decoction treatment of common diseases of the respiratory system drug law,R249.2
- Bing- thick academic thought and clinical experience and empirical studies apply to turtle soups treatment of chronic kidney disease,R249.2
- Research for Infrared Image Target Identification and Tracking Technology,TP391.41
- The Design and Implementation of Bicluster Data Analyzing Software,TP311.52
- Research on Clustering Algorithm Based on Mutation Particle Swarm Optimization,TP18
- Research on Fuzzy C-Mean Clustering Algorithm Based on Particle Swarm Optimization and Shuffled Frog Leaping Algorithm,TP18
- Research on Clustering Algorithm Based on Genetic Algorithm and Rough Set Theory,TP18
- Research of Diagnosing Cucumber Diseases Based on Hyperspectral Imaging,S436.421
- The Research and Application of Data Mart in the Telecommunication Business Analysis,TP311.13
- Research on Intrusion Detection Based on Feature Selection,TP393.08
- Web Usage Mining and the Research of Personalized Recommendation,TP311.13
- Adaptive Learning Environment Behavior Associated with the Level of Learning Research Model,G434
- Research of Fault Diagnosis Method of Analog Circuit Based on Improved Support Vector Machines,TN710
- Prediction Research of the Surrounding Rock Deformation in Engineering by the Support Vector Machine,TV223.1
- Markov Chain Monte Carlo Based Improvements to the Learning Algorithm of Restricted Boltzmann Machines,O211.62
- Study on Fast Detection of Gray Mold on Plants Using Spectral and Multi-Source Spectrum Imaging Technology,S432.4
- A Static Behavior-Based Method to Detect Malware on Android,TP309
- Modulation Classification Algorithms of Digital Communication Signals,TN914.3
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer software > Program design,software engineering > Programming > Database theory and systems
© 2012 www.DissertationTopic.Net Mobile
|