Dissertation > Excellent graduate degree dissertation topics show

Forum corpus -based automatic identification of new methods unknown words

Author: DuZuo
Tutor: XiongHaiLing
School: Southwestern University
Course: Computer Software and Theory
Keywords: unknown word Chinese word segmentation web spider corpus
CLC: TP391.1
Type: Master's thesis
Year: 2010
Downloads: 75
Quote: 1
Read: Download Dissertation

Abstract


Identification of unknown Chinese words is the bottleneck in the field.This paper presented that download adequate web documents from BBS with web spider in order to construct a corpus which was updated periodicity. Then generate candidate words list by extracting words from the corpus with this new function. Finally, compare this candidate words list and the previous lexicon, so as to recognize the unknown words. Experiments showed that the proposed method was more efficient.Different with English word, Chinese word has its own characteristics. As the composition and use habit of Chinese language, parser Chinese word is a harder problem than the English.At present, the Chinese word segmentation algorithm is mainly in three ways:based on string matching algorithms, based on understanding algorithm and based on statistical algorithms. These three methods, both in the unknown word to varying degrees, there are some problems:based on string matching algorithms can not recognize unknown words fundamentally. Based on understanding algorithm is more difficult and complexity of the time complexity and the space complexity. So it is not widely used. Based on statistical algorithm is more feasible and popular method at present, but there are also some errors in identification.Over all, based on statistical algorithm is a relatively feasible and practical application of a method. This paper studed unknown Chinese word based on statistical algorithms for unknown words identification. First, the Chinese word segmentation, especially in unknown word recognition is descripted. Secondly, the traditional word segmentation algorithms and segmentation system has been analyzed and compared. There are three kinds of traditional Chinese word segmentation algorithms:based on string matching algorithm; based on understanding of the algorithm; based on statistical algorithms. Mechanical matching algorithms can not extract unknown word from a fundamentally reason; understanding algorithm due to algorithm complexity and great difficulty, practical development and application is not widespread; Statistics algorithm in a certain extent, may solve some of unknown words, the algorithm became more popular, but it is still available in more statistical algorithms can not determine the miscarriage of justice and circumstances. This paper presented methods that download adequate web documents from BBS with web spider in order to construct a corpus which was updated periodicity which was contrarily against to the shortage of traditional ways. This step can ensure the timeliness of the corpus. Then generate candidate words list by extracting words from the corpus with this new function MD (the Mutual Information function and Duplicated Combination Frequency are combinated to construct a new statistic MD). This candidate words list and the previous lexicon were compared, so as to recognize the unknown words. Subsequently, according to this thinking program designed to test, set up a test environment. New word recall rate and accuracy of two indicators shows that this design of unknown words automatically recognize the new method is feasible.

Related Dissertations

  1. A Corpus-based Study on That-Clause in Political Documents,H314
  2. A Corpus-Based Study on English Personal Pronouns in Political Speeches and Statements,H314
  3. A Contrastive Study of the Translation of Nominal Concept in Taoism,H059
  4. A Corpus-based Study on Verb-noun Collocation Errors by Chinese English Majors from the Perspective of Conceptual Transfer,H319
  5. A Corpus-based Study on the Chinese Learners’ Use of the Word AS in Light of the Lexical Priming Theory,H314
  6. The Research of Using BNC to Correct the Chinese English Emerging in the English Writing,G633.41
  7. A Research on the Approach of Composition Teaching Reform in Preliminary School,G623.24
  8. The Diagnostic Value of the Measurement of Corpus Callosum Area on MRI Midsagittal Section to Brain Atrophy,R445.2
  9. An Empirical Study of Corpus-Based Lexical Approach to Non-English Majors’ Writing,H319
  10. Investigation on the English Abstracts of TCM Experimental Dissertations,H315
  11. Chinese word segmentation based on understanding of system design and implementation,TP391.1
  12. The Research on Full-Text Search and Related Technologies,TP391.3
  13. A Corpus-based Study on the Vocabulary Errors in the Writings of Chinese EFL Learners and Its Pedagogical Implications,H319
  14. A Corpus-Based Contrastive Study on the Different Usage of Denominal Verbs between ENL Speakers and Chinese EFL Learners,H319
  15. Research and Implement of Chinese Word Segment Techniques Based on the Conditional Random Field,TP391.1
  16. Research on Approaches of the Subjective Automated Assessment,TP391.1
  17. Opinion Mining Based Product Usability Modeling and Evaluation,F274
  18. A Study on Hybridity in Translations of Wuthering Heights,I046
  19. Norm and Creative Treason: a Contrastive Study on Three Chinese Translations of A Tale of Two Cities,I046
  20. Genre Analysis of Prefaces of Academic Monographs in Linguistics,H052
  21. The Application of Web Data Mining Technology on Network Education BBS,G434

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net  Mobile