Dissertation > Excellent graduate degree dissertation topics show

HMM-based Chinese Part-of-Speech Tagging and Improvement

Author: ZhuYing
Tutor: MaJianFen
School: Taiyuan University of Technology
Course: Applied Computer Technology
Keywords: speech tagging hidden Markov model second-order hidden Markov model Viterbi algorithm
CLC: TP391.1
Type: Master's thesis
Year: 2011
Downloads: 159
Quote: 0
Read: Download Dissertation

Abstract


Part-of-Speech (POS) tagging is one of the research points on Natural Language Processing which has important significance. It involves a wide range of applications, and it plays an important foundation role in the context of Information Processing. The quality of POS tagging has a direct impact on the accuracy of all Information Processing which based on the results of POS tagging, such as Syntax Analysis, Speech Recognition, Text Classification, Text to Speech, Information Retrieval, machine translation and so on. There are some difficult in the implementation process of POS tagging. Such as, the ambiguity processing of concurrent words, the processing of unknown words and proper noun. Because of the characteristics of Chinese language itself and the restrictions on Chinese Linguistics Research, Chinese POS tagging has more difficulties and complexities.There are many ways of POS tagging, and them can be grouped into two categories as Rule-based methods and Statistical methods in general. HMM-based POS tagging is a typical example of statistical methods. Although the applications of HMM in POS tagging are very mature, but how to improve the tagging accuracies of Concurrent words and Unknown words is still focal points of the study on HMM-based POS tagging. This text bases on the tagged Chinese Corpus named《People’s Daily (Jan.1998)》, establishing the second-order Hidden Markov Modes(HMM2), improving the tagging of Unknown words, by training, testing, and evaluating the model to achieve the Chinese POS tagging. As follows:(1) Because the selection of the corpus plays an important influence in the results of POS tagging, preprocessing the corpus before training and testing. The preprocessing is removing the second dimension and the sign of Proper noun tagging (continue to have the Proper nouns and their tags) to improve the accuracies of the experiments.(2) When the general HMM is carrying out the POS tagging, it is just relied on the tagging of the previous word to estimate the tagging of the current word. Considering based on the linguistic knowledge, this method is not exhaustively extract the semantic information of context. Therefore, put forward the idea that establish the second-order HMM to increase the use of the semantic information of context, thereby increasing the accuracy of the POS tagging results. In the establishment of the second-order HMM, the state transition probability which gets from the training date has been smoothed; as the same time, according to the test in the actual situation, modifying the acquisition method of observation probability, and processing the unknown words in order to further ensure the accuracy of the experiments. (3) In testing, the traditional Viterbi Algorithm can’t meet the improved second-order HMM. So, making the Viterbi Algorithm to be improved and expanded in order to meet the needs of the modified second-order HMM.After the open testing of a ten thousand words on the training corpuses with the annotation of 26 tags and the other annotation of 39 tags, proving the improved second-order HMM in this text has a good effect than the general HMM and HMM2. Finally, this text gives a prospect of the development of POS tagging.

Related Dissertations

  1. Packet Loss Recovering Technology for Speech Transmission over Network,TN912.3
  2. Research on Domain Entity Attribute and Event Extraction Technology,TP391.1
  3. Research and Design of Electronic Equalization Based on Most Likelihood Sequence Estimation,TN911.5
  4. The Study of the Marking of Part of Speech Based on Class a in the Outline of the Graded Vocabulary for HSK,H146
  5. Multi-threaded fusion soccer video semantic analysis and event detection,TP391.41
  6. Chinese Speech Synthesis System Improvement and Implementation,TN912.33
  7. Study on Real-time Rhythmic Information Retrieval of Musical Audio Signal and the System Implementation,TN912.3
  8. Research on Characteristic Analysis and Recognition Algorithm of Heart Sound Signal,R318.04
  9. Extended Hidden Markov Models and Parameter Estimation Based on Genetic Algorithm,O211.62
  10. Research of Multi-Sensory Myoelectric Prosthetic Hand with Hardness and Thermal Conductivity,TP242
  11. The Study and Analysis of Oracle Bone Inscriptions Based on Statistical Natural Language Processing,TP391.1
  12. A Research on Chinese Word Segmentation Based on Phonetic Annotation,TP391.1
  13. Research on DBN-Based Continuous Speech Recognition,TN912.34
  14. Research on the Key Technologies fo Speech Recognition for Robot Communication,TN912.34
  15. The LVCSR system based on adaptive methods of semi-supervised learning,TN912.34
  16. Some Strong Laws for Markov Chain Fields Indexed by a Nonhomogeneous Tree of Module M,O211.62
  17. Research on Automatic Notation of Word for Tibetan Corpus Based on HMM,H214
  18. Research on Information Awareness Technology Oriented to Cognitive Networks,TP393.02
  19. Research and Implementation on Community Discovery from Network Based on Data Mining,TP393.094
  20. Statistical Image Modeling and Image Segmentation in Contourlet Domain,TP391.41
  21. Event Detection Modeling and Optimization in Intelligent Video Surveillance,TP391.41

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net  Mobile