Dissertation > Excellent graduate degree dissertation topics show

Research on Semi-supervised Learning Algorithm Based on Tri-training Algorithm

Author: ChangZuo
Tutor: LiangJiYe
School: Shanxi University
Course: Systems Engineering
Keywords: Semi-supervised learning Tri-training algorithm Seeds sets Pair-wise constraints Active learning
CLC: TP181
Type: Master's thesis
Year: 2012
Downloads: 149
Quote: 0
Read: Download Dissertation

Abstract


Semi-supervised Learning (SSL), a kind of application-driven machine learning method, has become one of hot topics of artificial intelligence and pattern recognition areas. As the main branches of SSL, semi-supervised clustering gives a small amount of supervised information into the search process of optimal clustering, and semi-supervised classification attempts to exploit implied useful information from unlabeled samples in order to assist the classifier’s training. Recently, researchers have continuously introduced all kinds of SSL algorithms, which are applied to some actual fields, natural language processing, image processing, biometric identification, and so on.Tri-training is a representative method based on co-training mechanism. Although it can use classifiers for unlabeled samples’ annotation, this algorithm requires given sufficient labeled samples to guarantee initial classifiers with greater difference, besides it can’t deal with the situation where given supervised information includes pair-wise constraints.Therefore, regarding these described above shortcomings, this paper employs Tri-training as starting point, and concerns how to effectively select and mark unlabeled samples for semi-supervised clustering and classification, when there are different forms of supervised information. The main work is summarized in three aspects as follows:(1) A semi-supervised clustering algorithm based on Tri-training is introduced, as supervised information not only includes labeled samples, but also pair-wise constraints. Firstly, this algorithm selects some unlabeled samples and requires their class label, to enlarge the number of initial labeled samples. Secondly, pair-wise constraints are utilized to optimize enlarged labeled samples, with the purpose of improving its quality. Finally, parameters of K-Means algorithm are initialized by optimized labeled samples, and in the search process, pair-wise constraints are used to modify the clustering results each time. We also apply the proposed method to K-Means, Seeded-K-Means and COP-K-Means algorithm. Experimental results demonstrate that this method can take full advantage of given supervised information and get a better clustering results.(2) An active semi-supervised classification algorithm is proposed, based on Tri-training and few labeled samples, when there are only very few labeled samples in given supervised information. This method selects certain unlabeled samples which are most possibly wrong predicted or most typically represented class attribute, by means of integrating active learning thought. And these unlabeled samples are marked by expert users, to increase the number of initial labeled samples. Comparative experimental results show that when given initial labeled samples are very few in number, and Tri-training is unable to obtain satisfactory results, the proposed method can attain a classification model of better capability.(3) An active semi-supervised classification algorithm is designed, based on Tri-training and pair-wise constraints, when there are pair-wise constraints in supervised information. This method requires informative samples which are marked for expert users, so that there are enough labeled samples. And in the process of classification, pair-wise constraints are used to optimize labeled samples for training its classifier each time, in order to improve the data security. Experimental results illustrate that this method can effectively deal with the case where supervised information includes pair-wise constraints, compared with Tri-training. Furthermore, with the algorithm which isn’t introduced pair-wise constraints optimization mechanism, the proposed method not only improves prediction accuracy, but also less affected by parameters change, and more stable of the performance.Towards different forms of given supervised information, the paper’s research results can provide references about how to conduct SSL effectively, and further extend tri-training’s application prospects in the actual fields.

Related Dissertations

  1. Research and Implementation of Application Traffic Classification & Restoring,TP393.08
  2. Research on Intrusion Prevention Based on Semi-Supervised Fuzzy Clustering,TP393.08
  3. Research on Lie Group Deep Structure Learning Algorithm,TP181
  4. A Novel Labels and Similarity Reconstruction Based on K-means Algorithm Application on Text Clustering,TP181
  5. Research on Fingerprint Segmentation Appraisal and Fingerprint Segmentation with Semi-Supervised Learning,TP391.41
  6. Semi-supervised Learning and Active Learning of Sentiment Classification Coupled with Domain Knowledge,TP181
  7. Study on Tracking Algorithm Based on Semi-supervised Online Learning,TP391.41
  8. Semi-supervised hashing algorithm based Image Retrieval Methods,TP391.41
  9. Domain-specific Chinese term extraction,TP391.1
  10. Research and Implementation of Time Series Classification Based on Semi-supervised Learning,TP181
  11. Researches on Semi-supervised Learning,TP181
  12. Based on semi- supervised learning condition recognition method and copper flash smelting process application,TF811
  13. The Graph-based Semi-supervised Learning with Missing Data,TP181
  14. Research on the Application of Data Mining Technology in the Regulating Matriculation for Postgraduate,TP311.13
  15. Research of Query-by-committee Method of Active Learning,TP181
  16. Comparison and Improvement of Two Methods Based on Semi-Supervised Learning,TP18
  17. Research on Intrusion Detection Based on Semi-Supervised SVM,TP393.08
  18. The LVCSR system based on adaptive methods of semi-supervised learning,TN912.34
  19. Design and Implementation of Hydropower Fault Classifer Based on Support Vector Machine,TV738
  20. A Study on Some Problems of Semi-supervised Learning,TP181
  21. Model Misfit in Classification and Its Solution,TP181

CLC: > Industrial Technology > Automation technology,computer technology > Automated basic theory > Artificial intelligence theory > Automated reasoning,machine learning
© 2012 www.DissertationTopic.Net  Mobile