Dissertation > Excellent graduate degree dissertation topics show

Study of Chinese morpheme -based lexical analysis methods

Author: WangQi
Tutor: FuGuoHong
School: Heilongjiang University
Course: Applied Computer Technology
Keywords: Chinese lexical analysis word segmentation part-of-speech tagging sense tagging unknown word
CLC: TP391.1
Type: Master's thesis
Year: 2011
Downloads: 27
Quote: 0
Read: Download Dissertation

Abstract


As the key problem of Chinese information processing, Chinese lexical analysis mainly includes three tasks:word segmentation,part of speech tagging and meaning disambiguation. Although in recent years the Chinese lexical analysis has made great progress, it is facing with huge challenges when dealing with large-scale open text, especially with the problem of unknown words.Therefore, how to excavate and shows the morphological features effectively, in addition, how to solve the problem of unknown words identification and prediction are the major difficulties which Chinese lexical analysis should face, also, it is a research focus.Based on large-scale training corpora, under the machine learning framework, this paper does a research in Chinese morphology analysis methods and focus on the problem of Chinese unknown words recognition and meaning prediction. Specifically, this paper does the research following three aspects:Firstly, taking morpheme as the basic tokens of word-formation, the paper studies the influnces of different tagging sets and different window sizes on Chinese word segmentation. The research takes advantage of Conditional Random Fields model. The experimental results from the data of SIGHAN Bakeoff 2005 show that the introduction of morpheme is good for improving the unknown words recognition ability.Secondly, in view of the unknown words prediction problem in Chinese part of speech tagging, taking morpheme as the base, this paper realized a Chinese part of speech tagging system which is based on maximum entropy model. It discovers and combines the lexical features of words internals to create the system. The experiment based on SIGHAN Bakeoff 2007 speech proves that morpheme-based Chinese part of speech tagging method has the large advantages to predict the POS of unknown words.Last but not least, this paper puts forward a Nave-Bayes model based on central morphemes. At the same time, it studies in meaning prediction problem for Chinese unknown words under the affection of this model. Experiment shows that the Nave-Bayes model based on central morphemes could, to some degree, soleve the meaning prediction problem of Chinese unknown words.

Related Dissertations

  1. Chinese word segmentation based on understanding of system design and implementation,TP391.1
  2. The Research on Full-Text Search and Related Technologies,TP391.3
  3. Research and Implement of Chinese Word Segment Techniques Based on the Conditional Random Field,TP391.1
  4. Research on Approaches of the Subjective Automated Assessment,TP391.1
  5. Opinion Mining Based Product Usability Modeling and Evaluation,F274
  6. The Study of Chinese Word Segmentation on Uygur Students’and Han Students’ Eye Movement,B842.3
  7. The Application of Web Data Mining Technology on Network Education BBS,G434
  8. The Study of College Digital Campus Web Information Filtering,TP393.09
  9. The Study of the Marking of Part of Speech Based on Class a in the Outline of the Graded Vocabulary for HSK,H146
  10. Based WebHarvest the Chinese financial news search engine design and implementation,TP311.52
  11. Chinese XML Compression Technology,TP311.11
  12. Corporate e-mail monitoring system design and implementation,TP393.098
  13. Research on Chinese Word Segmentation of Search Engine,TP391.1
  14. The Study of Modern Chinese New Word Extraction,H08
  15. Research on Text Watermarking Based on Natural Language Processing,TP309.7
  16. Teaching resources search platform - SiseNutch,TP311.52
  17. Chinese semantic annotation consistency test methods Research and Implementation,TP391.1
  18. Study of Kazak Part-of-Speech Tagging Based Upon HMM,TP391.1
  19. Study on the Ambiguity of "Bu X Shi" Character String and the Strategies of the Elimination,H146
  20. Word segmentation Chinese reading for Russian students,H195
  21. Search engine for the field of art education applications,J20-4

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net  Mobile