Dissertation > Excellent graduate degree dissertation topics show
POS Tagging System on Hierarchical Classification Labels
Author: PanZuo
Tutor: WuLiDe
School: Fudan University
Course: Applied Computer Technology
Keywords: Hierarchical classification label Part of Speech Tagging Maximum Entropy Wikipedia WordNet
CLC: TP391.1
Type: Master's thesis
Year: 2009
Downloads: 87
Quote: 1
Read: Download Dissertation
Abstract
|
As the basic operation of natural language processing , POS tagging can provide a lot of useful information about the word and its adjacent components , and therefore is often one of the modules that form a complex application . POS tagging task is text understanding , the text generation an important foundation for the field of natural language processing , after years of development , is considered a relatively mature field . However, with the development of the Internet and the growing number of information irregular phrase text as hierarchical classification label . On the other hand , the speech tagging tools are built on the basis of conventional long sentence , unsatisfactory performance on the phrase text . It is in this context that an in-depth study hierarchical classification label - oriented part of speech tagging algorithm . This article describes a core part-of-speech tagging technology and research methods , including four classic POS tagging model and algorithms . By manually labeled , we successfully analyzed six major differences between the text of the hierarchical classification label with a long sentence , to find the root cause of the traditional part-of-speech tagging tool for poor performance , and pointed out that the two key issues that need to be resolved : the path information and proper nouns . On this basis , this paper presents a part of speech labeling algorithm based on maximum entropy models . To the path information is integrated into the input , we introduce a new marking PATH while adding the characteristics of the three types of use of the path information . Hierarchical classification label terms , we proceed from WordNet and Wikipedia , are constructed out of a dictionary and a database , and then introduced in the form of binary features maximum entropy model . Category Dmoz , these improvements have made ??very significant performance gains , thereby demonstrating the effectiveness of the method . Part-of-speech tagging hierarchical classification label can be applied to automatic Web page classification system . This page classification system is too dependent on human-annotated pages as the training corpus . We design such a system , corpus , the clever use of the part of speech of Classification and Labelling of label information and search engine , automatically construct a number of training corpus . The experimental results show that the performance of this method is fully acceptable .
|
Related Dissertations
- Research on Domain Entity Attribute and Event Extraction Technology,TP391.1
- WordNet and the \,G254
- Based on Wikipedia disambiguation of named entities,TP391.1
- Based on Chinese Wikipedia semantic correlation computation Research and Implementation,TP391.1
- Research on the Semantic Annotation of Domain with WordNet,TP391.1
- Research on Prediction of Protein-Protein Interactions Based Maximum Entropy Clustering Algorithm,Q51
- The Research of Expanding the Semantic Information Function to Search Engine,TP391.3
- Based on Maximum Entropy Model Research on extracting Chinese entity relation,TP391.1
- Research on Text Clustering Based on Latent Semantic Analysis and Self-organizing Maps,TP391.1
- Encyclopedia of Chinese knowledge -based search systems design and implementation,TP391.3
- The Automatic Identification of the Semantic Core Words for Frame Elements,TP391.1
- Music Named Entity Recognition Technology,TP391.1
- Hybrid method based on the complexity of named entity extraction,TP391.1
- The Research of Semantic Annotation System for Scientific Literature,TP391.1
- Query Expansion Based on User Log Clustering,TP393.09
- Automatically Chinese Address Recognition and Normalization,TP391.43
- Research on Ontology-based Query Translation in Deep Web,TP311.13
- Application Research on Chinese Named Entity Recognition Based on Domain Ontology,TP391.1
- Extracting Structured Information from the Chinese Wikipedia and Measuring Relatedness between Words,TP391.1
- Analysis of Storm Flood Variation Characteristics in Daqing River Basin,TV122.1
- Study the Human Dynamics in Wikipedia,C912
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net Mobile
|