Dissertation > Excellent graduate degree dissertation topics show

Study on Selected Issues of Large Scale Text Classification

Author: LiZiQiang
Tutor: ZhouMingTian
School: University of Electronic Science and Technology
Course: Applied Computer Technology
Keywords: SVM MEB optimal core set coreference analysis syntact phrase
CLC: TP391.1
Type: PhD thesis
Year: 2013
Downloads: 73
Quote: 0
Read: Download Dissertation

Abstract


Text classification is an elemental issue in text information processing, and hasreceived wide attentions. However, with the development of social internetization, largescale text information has emerged violently. This makes text classification to face agreat new challenge. This dissertation explores the problem from the viewpoints of textrepresentation and efficiently SVM training:1. A text is always represented as a vector of weights of the words. Each weight isevaluated from the occuring frequency of a word in a text and the times a word appearsin different texts. The distributing characteristics of these measurements are examinedon real datasets. The results show that, when selecting features, words with middlefrequency are chosen preferentially, or words are divided into three groups, the highfrequency, the middle frequency and the small frequency, respectivly. It is also shownthat the IDF factor should be enhanced by frequency of a word on whole dataset.2. Phrases provide more semantic information than words. But feature selectionalgorithms are traditionally used to choose a subset of phrases to represent a text. Thisdissertation observes that, if phrases are chosen from the viewpoint of the levels of theparse tree, then the recall performance can be improved. This type of phrase can reflectits role and function in the sentence where it emerges. Experimental results show thatthe new text representation can enhance the recall performance.3. The semantic relationships between adjacent words are usually used to reformVSM. Beyond that, the dictionary semantic relationships between nonajacent words,even not arising in any text, are used to do this work. This dissertation explorescoreference in the context to enhance word frequency. In this way, the real frequenciesof features are expressed accurately from the viewpoint of semantics. Experimentalresults show that the new text representation can improve the recall performance.4. To save part of the kernel matrix in a cache is an important accelerating methodfor SVM decomposition optimization. But the behavior of traditional decompositionalgorithms doesn’t show good localization all the time. A type of three-layer workingset selection framework is proposed to localize the iterating of decomposition algorithms. Combined with multiple working set selection strategies, furtheracceleration of traditional decomposition algorithms is achieved.5. For large scale text classification, it is a good strategy to refine the problemsstep by step. Intuitively, the profile of each class of data is the most important subset forclassification task. This dissertation fits each class of data with one hyperplane.Modeled as a MEB problem, the fitting problems can be solved by optimal core setalgorithms. Experimental result shows that a very high efficiency is got when the SVMis trained on this type of small subsets. Furthermore, a very sparse solution is obtained.6. Different from fitting each class of data isolatedly, an improved fitting modelthat considers the separateness between two classes is developed. This model not onlyfits each class of data with one plane, but also leaves the other class of data on one sideof the plane as possible. Experimental result shows that a very sparse solution is madeout efficiently, and it has comparable average generalizing performance with standardSVM.7. Another improved fitting model considering the separateness between twoclasses is explored. It fits each class of data with one plane, and requires the other classof data not only lie on one side of the plane, but also as far as possible from the plane.This new method keeps similar average generalizing performance with standard SVM.It is notable that the new improvement has higher potential training efficiency becausethe fitting cost is almost near half what the training data needs in fact.

Related Dissertations

  1. Soft Sensor of Naphtha Dry Point on Support Vector Machines Regression,TE622.1
  2. The Research of the Fault Diagnoses Algorithm for the Liquid Rocket Engine Testing Bed Based on PCA-SVM,V433.9
  3. ISAR Imaging Simulation of Space Targets and Target Recognition Based on ISAR Images,TN957.52
  4. Research on Autamatic Music Structrue Analysis,TN912.3
  5. Research on Feature Extraction and Classification of Pulse Waveform for Cholecystitis and Nephrotic Syndrome Diagnosis,TP391.41
  6. Research on Classification Method of Tongue Substance Color and Tongue Coating Color Based on SVM,TP391.41
  7. The Research on Paper Currency Classification Method Based on Harr-Like Feature and Minimal Ball Including Samples,TP391.41
  8. Research on Focused Crawler Based on SVM Classification Algorithm,TP391.3
  9. Research on Predicting Intrinsic Disorder Protein Structure Based on Supervision Manifold Learning Algorithm,Q51
  10. Study on the Road Condition Monitoring Based on Vehicular 3D Acceleration Sensor,TP274
  11. Research and Application of Gas Emission Prediction Based on LS-SVM,TD712.5
  12. Research of Support Vector Machine Based Fault Diagnosis System,TH165.3
  13. Computer-aided Hepatopathy Detection Based on MRI Perfusion Images,R445.2
  14. Research of Chinese Ellipsis Identification,TP391.1
  15. Research of Orange Quality Classification Technology Based on Computer Vision,TP391.41
  16. Research on Supplier-Selection Based on Strategic Sourcing,F426.31
  17. Analysis on Synoptic Climatology Characteristics and Forecast Methods of Fog in Hainan,P457
  18. Research of License Plate Recognition Based on Rough Sets and Fuzzy SVM,TP391.41
  19. Study on Visual Target Detection Based on SVM,TP391.41
  20. Vision-based Human Movement Analysis,TP391.41
  21. Research on Semantic Role Labeling for Chinese Nominal Predicates,TP391.1

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net  Mobile