Dissertation > Excellent graduate degree dissertation topics show

POS Tagging System on Hierarchical Classification Labels

Author: PanZuo
Tutor: WuLiDe
School: Fudan University
Course: Applied Computer Technology
Keywords: Hierarchical classification label Part of Speech Tagging Maximum Entropy Wikipedia WordNet
CLC: TP391.1
Type: Master's thesis
Year: 2009
Downloads: 87
Quote: 1
Read: Download Dissertation

Abstract


As the basic operation of natural language processing , POS tagging can provide a lot of useful information about the word and its adjacent components , and therefore is often one of the modules that form a complex application . POS tagging task is text understanding , the text generation an important foundation for the field of natural language processing , after years of development , is considered a relatively mature field . However, with the development of the Internet and the growing number of information irregular phrase text as hierarchical classification label . On the other hand , the speech tagging tools are built on the basis of conventional long sentence , unsatisfactory performance on the phrase text . It is in this context that an in-depth study hierarchical classification label - oriented part of speech tagging algorithm . This article describes a core part-of-speech tagging technology and research methods , including four classic POS tagging model and algorithms . By manually labeled , we successfully analyzed six major differences between the text of the hierarchical classification label with a long sentence , to find the root cause of the traditional part-of-speech tagging tool for poor performance , and pointed out that the two key issues that need to be resolved : the path information and proper nouns . On this basis , this paper presents a part of speech labeling algorithm based on maximum entropy models . To the path information is integrated into the input , we introduce a new marking PATH while adding the characteristics of the three types of use of the path information . Hierarchical classification label terms , we proceed from WordNet and Wikipedia , are constructed out of a dictionary and a database , and then introduced in the form of binary features maximum entropy model . Category Dmoz , these improvements have made ??very significant performance gains , thereby demonstrating the effectiveness of the method . Part-of-speech tagging hierarchical classification label can be applied to automatic Web page classification system . This page classification system is too dependent on human-annotated pages as the training corpus . We design such a system , corpus , the clever use of the part of speech of Classification and Labelling of label information and search engine , automatically construct a number of training corpus . The experimental results show that the performance of this method is fully acceptable .

Related Dissertations

  1. Research on Domain Entity Attribute and Event Extraction Technology,TP391.1
  2. WordNet and the \,G254
  3. Based on Wikipedia disambiguation of named entities,TP391.1
  4. Based on Chinese Wikipedia semantic correlation computation Research and Implementation,TP391.1
  5. Research on the Semantic Annotation of Domain with WordNet,TP391.1
  6. Research on Prediction of Protein-Protein Interactions Based Maximum Entropy Clustering Algorithm,Q51
  7. The Research of Expanding the Semantic Information Function to Search Engine,TP391.3
  8. Based on Maximum Entropy Model Research on extracting Chinese entity relation,TP391.1
  9. Research on Text Clustering Based on Latent Semantic Analysis and Self-organizing Maps,TP391.1
  10. Encyclopedia of Chinese knowledge -based search systems design and implementation,TP391.3
  11. The Automatic Identification of the Semantic Core Words for Frame Elements,TP391.1
  12. Music Named Entity Recognition Technology,TP391.1
  13. Hybrid method based on the complexity of named entity extraction,TP391.1
  14. The Research of Semantic Annotation System for Scientific Literature,TP391.1
  15. Query Expansion Based on User Log Clustering,TP393.09
  16. Automatically Chinese Address Recognition and Normalization,TP391.43
  17. Research on Ontology-based Query Translation in Deep Web,TP311.13
  18. Application Research on Chinese Named Entity Recognition Based on Domain Ontology,TP391.1
  19. Extracting Structured Information from the Chinese Wikipedia and Measuring Relatedness between Words,TP391.1
  20. Analysis of Storm Flood Variation Characteristics in Daqing River Basin,TV122.1
  21. Study the Human Dynamics in Wikipedia,C912

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net  Mobile