Dissertation > Excellent graduate degree dissertation topics show

Research and Implementation of Chinese Word Segmentation Algorithm

Author: LinDongSheng
Tutor: LiuXiaoXia
School: Northwestern University
Course: Computer Software and Theory
Keywords: Participle Dictionary Disambiguation Unknown words Rulebase
CLC: TP391.1
Type: Master's thesis
Year: 2011
Downloads: 380
Quote: 4
Read: Download Dissertation

Abstract


Chinese word segmentation is based on a particular specification Chinese word sequence contiguous sequence of words cut into rational process . As a basis for natural language processing tasks , Chinese word segmentation has been widely used in related fields . Therefore, the study Chinese word segmentation algorithm has important theoretical and practical significance. In order to meet the upper application utility requirement for segmentation , this paper based on statistical mechanical word and sub- lexical combination is proposed based on the Chinese dictionary and statistical rules Segmentation. The algorithm is fast segmentation using mechanical pretreatment of lexical text after the initial segmentation using improved bidirectional maximum matching word ambiguity detection method detects Duan , and use statistical models based on binary segmentation Digestion full ambiguity. Secondly, the use of role-based named entity recognition method to identify unknown words . Finally , the introduction of the rule base for segmentation results further amendment. This research work are: 1 ) use of secondary indexes dictionary structure, enhance the dictionary lookup speed , using the Java object serialization technology to achieve the dictionary file loading ( deserialization ) and dictionary object serialization . 2 ) the detection ambiguity , an improved bidirectional maximum matching detection algorithm can detect not only the chain length of odd ambiguities , and to detect all satisfy both chain length and cross section length is an even number of 1 ambiguities . In the ambiguous phrase , using the full syncopation digestion ambiguity. 3 ) In the unknown word recognition , hidden Markov model will solve coding problems Viterbi algorithm for forward role labeling , using role-based pattern matching on the set of the Chinese proper nouns . Use a small correction rule base word fragments amendments. 4 ) At present Chinese word segmentation package is mostly in C language development, and as one of the mainstream development language Java, its Chinese word components is relatively small. Therefore, the word segmentation algorithm based on the study , design and implementation support for the Java language Chinese word segmentation system . Experiments show that the Chinese word segmentation algorithm CPU3.0GHZ, 2GB memory environment, segmentation speed of approximately 21,000 words / sec , word accuracy indicator F-1 value reaches about 95%, basically able to meet most of the upper application requirements.

Related Dissertations

  1. Word Sense Disambiguation Corpus Automatic Acquisition,TP391.1
  2. Context-Dependent Lexical Paraphrasing,TP391.1
  3. Storage, Management and Sharing of Farmland Information Based on Metadata,S126
  4. The Frame Disambiguation of Automatic Identification of Chinese Frame,TP391.1
  5. Digital image forensics technology research,TP391.41
  6. The Supplementation and Revision of the Radical of "Fu" in Chinese Dictionary,H164
  7. Design and Implementation of Data Dictionary in Da Meng DBMS,TP311.13
  8. Chinese XML Compression Technology,TP311.11
  9. Web theme concept extraction process and Visualization,TP393.092
  10. The Disambiguation of Juxtaposition of Noun Phrases in NLU and Its Application in Intelligent Instrument Design,TP391.1
  11. The Study of EPDM Data Dictionary Constructing and Model Management Based on the Data Element,TP315
  12. Technology for Domain-Oriented Automatic Information Extraction from Semi-Structured Web,TP391.1
  13. Word Sense Disambiguation Technology Research Based on HowNet and Bayesian Model,TP391.1
  14. Analysis and Identification of Chinese United Phrase Structure,H146
  15. Research and Application of Information Extraction -based the multidimensional semantics Internet drugs,TP393.09
  16. Research on the Equivalentanalysis Model and Extracting Features of the Heart Sound Signal,TN911.6
  17. Research on the Words in the Kezuozhuiyu,H131
  18. The Use of Be + Past Participle Construction in English Writing of Chinese College Students,H319
  19. Sparse Decomposition Based on Cascade Dictionary and Its Application,TN911.7
  20. Key Technologies Research of Downhole Operation Data Center,TP308
  21. Research on Chinese Word Segmentation of Search Engine,TP391.1

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net  Mobile