Dissertation > Excellent graduate degree dissertation topics show
A Study of Text Categorization Algorithms Based on Latent Semantics
Author: LuoLe
Tutor: LiLi
School: Southwestern University
Course: Applied Computer Technology
Keywords: LSA LDA SVM Dimension Reduction Classification Clustering
CLC: TP391.1
Type: Master's thesis
Year: 2013
Downloads: 31
Quote: 0
Read: Download Dissertation
Abstract
|
With the development of information technology, the Internet data and resources increase massively. In order to effectively manage and utilize this massive information, content-based information retrieval and data mining have gradually been focused on. Text classification (TC) technology, which is a sort of supervised machine learning methods, plays an important role in information retrieval and text mining and is applied extensively in the field of natural language processing and understanding, information organization and management and information filtering. Text classification methods based on the machine learning arising in the1990s pay more attention to the classifiers’capabilities of automatically mining and generation and dynamic optimization and surpass the knowledge engineering and expert systems in the classification performance and flexibility.In recent years, however, with the rise of web2.0, social networking, such as Renren, Facebook, Twitter, Sina Weibo, etc. not only become very popular, but also become a way of the modern life. And thus they result in the massive user-generated content data (UGC) filling with the huge cyberspace, of which more than80%is natural language text. Since these texts are special and bear their own characteristics, many traditional text classification techniques cannot achieve good results. There are three salient features in these texts from the social network: high-dimensional, sparse, not normative. In other words, thousands of these short texts likely produce the vector with more than ten thousands of dimensions, which is too time-consuming for the classification algorithm; compared with the long texts, these texts have even less keywords, producing the sparse "document-word" matrix and thus difficult to extract the effective features and to exploit the correlation between the features; in addition, abbreviations and catchwords are used extensively in social network, increasing the synonyms in the texts and making the text classification task more difficult to handle. This study aims at handling the problems in text classification in the social network. On one hand, traditional classification methods neglect semantic information----that is, not taking into account the synonyms and polysemys----and hence result in a lower classification accuracy; on the other hand, in order to ensure the classification accuracy traditional methods often select more than ten thousands of features, which makes the classification algorithm is less efficient. In other words, we cannot boost both of the classification efficiency and accuracy. With the above two issues, we proposed two algorithms:1) Mining the latent semantics of the WSDL texts through the LSI, and then clustering the services using the EM algorithm, and finally classifying the text using the SVM to observe the efficiency of multi-class classification;2) Analysis of the potential topics of texts by LDA, proposing the LDA+SVM classification algorithm to classify the two datasets efficiently.The experimental results show that1) and2) are two ways to both reduce the dimensionality of the data, but also improve the classification accuracy, solving the problem of classification efficiency and accuracy clash perfectly.
|
Related Dissertations
- Soft Sensor of Naphtha Dry Point on Support Vector Machines Regression,TE622.1
- The Research of the Fault Diagnoses Algorithm for the Liquid Rocket Engine Testing Bed Based on PCA-SVM,V433.9
- ISAR Imaging Simulation of Space Targets and Target Recognition Based on ISAR Images,TN957.52
- Research on Autamatic Music Structrue Analysis,TN912.3
- Research on Feature Extraction and Classification of Pulse Waveform for Cholecystitis and Nephrotic Syndrome Diagnosis,TP391.41
- Research and Implementation of Mining Implicit User Interest,TP311.13
- Establishment and Update of Similar Users’ Cluster in Personalized Information Retrieval,TP391.3
- Tourism Comments on the Internet’s Semantic Analysis and Usefulness Research,TP391.1
- Research on Classification Method of Tongue Substance Color and Tongue Coating Color Based on SVM,TP391.41
- The Research on Paper Currency Classification Method Based on Harr-Like Feature and Minimal Ball Including Samples,TP391.41
- Gao Zhong-ying academic thought and experience and use of Bufei Decoction treatment of common diseases of the respiratory system drug law,R249.2
- Evaluation of Photosynthetic Efficiancy of Seedlings of the Hybrid Progenies (F1) in Peach,S662.1
- Research on Nondestructive Detection Technology for External Qualities of Papayas Based-on Vision,S667.9
- Research on Focused Crawler Based on SVM Classification Algorithm,TP391.3
- Research on Predicting Intrinsic Disorder Protein Structure Based on Supervision Manifold Learning Algorithm,Q51
- Methodology for Linkage-map-based Quantitative Trait Loci Synthesis Analysis,S562
- Research on K-means Optimization Clustering Algorithm,TP311.13
- Study on Photosynthetic Characteristics of Peach Based on Heterosis of Assimilation Capacity,S662.1
- Study on the Road Condition Monitoring Based on Vehicular 3D Acceleration Sensor,TP274
- A Preliminary Study on Bare Soil Extraction and Its Moisture Retrieval from High Resolution SAR Data,S152.7
- Moving target trajectory analysis based Intelligent Traffic Monitoring System,TP277
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net Mobile
|