Dissertation > Excellent graduate degree dissertation topics show

Research and Implementation of Chinese Auto-segmentation System

Author: DaiJianYing
Tutor: HeZhongShi
School: Chongqing University
Course: Applied Computer Technology
Keywords: natural language processing Chinese auto-segmentation statistical language model time word
CLC: TP391.1
Type: Master's thesis
Year: 2005
Downloads: 541
Quote: 19
Read: Download Dissertation

Abstract


NLP (Natural Language Processing) is an important branch of Artificial Intelligence. Chinese auto-segmentation is the foundation of NLP, and also is a crucial issue in NLP. Chinese auto-segmentation system can be used for automatically recognizing the Chinese words. Up to now, although there are some research efforts in this field, however, there are still some problems in the practical applications, which need to be solved by further research. The main objective of this research is to design and implement a Chinese auto-segmentation system. After an analysis of the main difficulties, in order to reduce the difficulty of segmentation and improve the precision of segmentation, this research has designed and realized a Chinese auto-segmentation system based on a multi-step process strategy. The main work of this paper includes: First, this paper introduces the language model and the algorithms of Chinese auto-segmentation systems, and presents an algorithm to deal with the ambiguity of Chinese time words. Time words refer to expressions indicating both exact time and periods of time. The algorithm gets a 90% accuracy, which shows the effectiveness of the proposed algorithm. Second, this paper gathers, coordinates and establishes natural language resource the study needed, which mainly includes the manual segmentation label corpus’s gathering, processing and settling, and the crude corpus’s gathering and processing, dictionary, and the knowledge warehouse’ building. At the same time, the text’s non-Chinese characters and Chinese figure strand is also studied. The core work of the paper is designing and implementing a Chinese auto-segmentation system based on a multi-step processing strategy. The system includes some modules such as originally segmenting, POS tagging, ambiguity processing, model smoothing and Unknown Word Recognizing. Original segmenting is to find out the potential routes in sentences. Ambiguity processing refers to eliminating ambiguities using Bi-gram or the POS label wholehearted model, and combining ambiguities by SVM. The POS detecting method is used to realize the function of the Unknown Word Recognizing. Model smoothing technique is embodied in the process of the POS label and ambiguity processing. Last, the paper validates the system’s performance by experimentation. The system reaches a precision of 96.94% compared with artificial segmentation. The speed is between 1000 to 1400 words per second. Although the effect and precision is not as good as ICTCLAS, a Chinese auto-segmentation system developed by CAS, some new methods will be helpful for the future study. At the same time, the paper summarizes all the work and puts forward further work.

Related Dissertations

  1. Word Sense Disambiguation Corpus Automatic Acquisition,TP391.1
  2. Research on Semantic Role Labeling for Chinese Nominal Predicates,TP391.1
  3. Study of Structuralization for Electrocardiography Diagnostic Report,R444
  4. Home Academic Information Extraction System,TP393.092
  5. Printers based on natural language HCI Research and implementation,TP11
  6. Based on Chinese Wikipedia semantic correlation computation Research and Implementation,TP391.1
  7. The Research of Topic Based Multi-document Summarization,TP391.1
  8. Research of Protein-Protein Interaction Extraction Based on Rich Feature and Multiple Kernels Learning,Q51
  9. The Research of Web-based Community Medical Intelligent Service System,TP311.52
  10. AraOntoLT: A Framework for Ontology Learning from Arabic Text,TP391.1
  11. Research on the Construction of a Nlp-oriented Chinese Sentence Semantic Knowledge Database,H13
  12. Research and Design of Intelligent Generation of Filtered Rules,TP393.08
  13. The Study and Analysis of Oracle Bone Inscriptions Based on Statistical Natural Language Processing,TP391.1
  14. Web Knowledge Service Oriented of Medical Information Classification Approach,TP391.1
  15. Research in Thesaurus-based Ontology Building Method,TP391.1
  16. Chinese \,H146
  17. Research on Time Words Expressing "Present" in Contemporary Chinese,H146
  18. A Back-off Smoothing Algorithm of Language Model Based on Mutual Information,TP391.1
  19. Application Research on Statistical Language Model of Large Vocabulary Continuous Speech Recognition System,TP391.42
  20. Research on Context Sensitive Personalized Information Retrieval Technology,TP391.3
  21. The timed automata new recognition model and long span area automata research,TP23

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net  Mobile