Dissertation > Excellent graduate degree dissertation topics show

Chinese Word Segmentation Method Based on Dictionary and Statistics of the Words

Author: YueZhongYuan
Tutor: HuYan
School: Wuhan University of Technology
Course: Applied Computer Technology
Keywords: Chinese word segmentation Unkown words word frequency statistic Named Entity
CLC: TP391.1
Type: Master's thesis
Year: 2010
Downloads: 324
Quote: 18
Read: Download Dissertation

Abstract


With the development of the informationize, it is easy to obtain large amount of information. However, to deal with massive information manually is not possible,need to rely on the help of computer. And different from Western languages, Chinese, between words in no obvious signs segmentation. So, to make the computer capable of handling Chinese text, text must do chinese word segmentation first. As the complexity of chinese syntactic network and the the continual emergence of new words, Chinese word segmentation system has not achieved satisfactory results.This paper analyzes the actual use of Chinese word segmentation algorithm.many kinds of dictionary structure.Studied the current problems of Chinese word segmentation.In this paper,we used the combination method based on statistical and dictionary. Achieve Improvements in various aspects. First, we divide the whole text into shorter sentenees according to the Punetuations in this text.in the statistics, through the statistics of results of the fragmentation from Chinese word segmentation. Identified the unknown words in the text appeared more than once. And added to the temporary dictionary. Improved the structure of the dictionary. Put dictionary into a single basic dictionary and extended dictionary. This paper describes a Chinese word segmentation algorithm method based on statistical and dictionary we inerease the number of dictionaries, we add some sPeeial dictionaries which can be used to eliminate the ambiguousnesses and reeognize new words during the segmentation Process besides the basic dictionary.At the same time we reconstruted the data structure of basic dictionary in the memory of computer by using data structure"Hashtable",we choose the first two single charaeters of every word in the basic dictionary as the keywords of the main and sub Hashtables, the remanent words are stored in an array aceording to length.With these data structures, whenever our program meets a word, the program will be able to loeate the word straightly and quickly in dietionary.And adding word frequency information into the dictionary for ambiguity resolution. Extended dictionary include quantifiers dictionary, name dictionary, temporary dictionary, word dictionary Disable and so on.with the correct segmentation of the quantifiers reduce the number of ambiguity.Used improved mechanical method to do second word segmentation.Finally use rules to identification new words appear only once.This method has good ability to identification new words and ambiguity elimination. Basically satisfy the practical application of Chinese information processing requirements.

Related Dissertations

  1. Chinese word segmentation based on understanding of system design and implementation,TP391.1
  2. The Research on Full-Text Search and Related Technologies,TP391.3
  3. Click data and search results based on fragments excavated named entities,TP391.3
  4. Research and Implementation of Website Search Technology Based on Ajax/Lucene,TP393.092
  5. The Study of Ontology-Based Query Expansion,TP391.1
  6. A Design and Application of Personalized Information Retrieve and User Recommendation on Search Engine,TP391.3
  7. Research on Vertical Search Engine Based on SSH and Lucene,TP391.3
  8. Research and Application of Lucene Full-text Retrieval Technology in Patent Information Service Platform,TP391.3
  9. Research and Realization on Document Repository Based on DSpace,TP391.3
  10. Research and Design Topital Crawler for Agricultural Information,TP391.3
  11. Full-text search based on Lucene database research and application,TP391.3
  12. Implementation and Analysis of Tree to String Alignment Template Model in Statistical Machine Translation,TP391.2
  13. The Study of Graininess and Feature Selection for Chinese Named Entity Recognition,TP391.1
  14. The Studies on Chinese Text Categorization Based on Pso and Svm,TP391.1
  15. The Research of Word Index Method Based on Inter-Relevant Successive Trees Model,TP391.3
  16. Research on Named Entity Recognition Based on Rules,TP391.1
  17. Full-text search engine used in the design and implementation of e-mail,TP391.3
  18. Ontology-based protein-protein interaction information text mining method,Q51
  19. The Design and Implementation of Text Data Acquisition System Focused on News Field,TP391.1
  20. Resarch and Improvement of Lucene-Based Search Engine,TP391.3
  21. The Research of Semantic Annotation System for Scientific Literature,TP391.1

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net  Mobile