Dissertation > Excellent graduate degree dissertation topics show

Research of a New Dictionary in the Search Engine

Author: CaiRui
Tutor: ZhaoHeJi
School: Shandong University
Course: Computer Software and Theory
Keywords: Search engines related technologies the Chinese word segmentation verbatim binary search
CLC: TP391.1
Type: Master's thesis
Year: 2010
Downloads: 110
Quote: 0
Read: Download Dissertation

Abstract


Search engine is the system which collects specific information using a certain strategy and program on the Internet, organizes and processes them, submits the following informations, and provide search services for users. At present, as an important area of the Chinese information processing, the search engine plays a more important role in practice. The Chinese word segmentation is the primary key of Search engine technology.In natural language, the word is the smallest independently of meaningful activities. Be different from western language, there is no obvious symbol of segmentation between the word and the word, thus, the Chinese word segmentation is a foundation of Chinese information processing. The Chinese word segmentation widely utilized in the text information retrieval, automatic classification, search engines, automatic based, speech synthesis, automatic proofreading, machine translation (MT), etc. Word segmentation method directly affects the real-time and accurate of the performance of Chinese information processing.The existing Chinese word segmentation algorithms can be classified into 3 groups:the method based on String Matching, the method based on Understanding and the method based on Statistics. The segmentation methods based on the string matching also called mechanical word segmentation or the segmentation method based on the dictionary, and the largest matching method is the commonly used methods. This algorithm needs a new dictionary. The structure and search algorithm of the dictionary for the words segmentation algorithm is quite important. Existing dictionary mechanism has 3 kinds:dictionary based on the word binary search, dictionary based on the TRIE tree and dictionary based on the verbatim binary search. The dictionary structure based on the word binary search has simple data structure, takes up the little space, and is easy to maintain. However since the dictionary structure use the word binary search which need many test matches to get final results, the efficiency is low; the dictionary structure based on the TRIE tree has complex data structure, waste more space, and is hard to maintain. However since the dictionary structure use the verbatim search. The efficiency is high; although the dictionary structure based on the verbatim binary search adopted more efficient verbatim search method, but essentially the method is not perfect. Based on the analysis of three kinds of algorithms, this paper puts forward a new segmentation algorithm-layered binary word, it improve efficiency, control the complexity in a certain level, and find a balance between the efficiency and complexity.This paper describes the common segmentation algorithm, and then introduces the new dictionary mechanism which improves the maximal matching speed.Finally the paper compare the two dictionary mechanism and give the experimental analysis.

Related Dissertations

  1. Research and Implement of Chinese Word Segment Techniques Based on the Conditional Random Field,TP391.1
  2. Research on Approaches of the Subjective Automated Assessment,TP391.1
  3. Based WebHarvest the Chinese financial news search engine design and implementation,TP311.52
  4. Chinese XML Compression Technology,TP311.11
  5. Corporate e-mail monitoring system design and implementation,TP393.098
  6. Legislative Regulation on the PPC of Search Engines,D923.43
  7. Intelligent Analysis and Implement of Compliant Information Based on Filter Technology,TP391.1
  8. Probability and statistics based on dictionaries and Chinese word segmentation algorithm,TP391.1
  9. Research and Implementation of Website Search Technology Based on Ajax/Lucene,TP393.092
  10. Research and analysis for users interested in model - based personalized search engine,TP391.3
  11. Optimization of SOM Algorithm and Application in Chinese Text Clustering,TP391.1
  12. Study on Chinese Text Categorization,TP391.1
  13. The Research and Implemenation of the Chinese Word Segmentation System Combining Omini-Segmentation with Statistic,TP391.1
  14. Research and Implementation of Network Product Search Engine Based on Domain Ontology,TP391.3
  15. Research and Development of Digital Resources Search Engine Technologies on Campus Network,TP391.3
  16. Chinese Spam Filtering Based on SVM,TP393.098
  17. Research and Implementation of an Information Pre-process Platform of Public Opinion,TP393.09
  18. Reverse Backtracking Research of Chinese Segmentation Based on Last Word Dictionary,TP391.1
  19. Research and Implementation of Key Technologies on Web Text Mining,TP391.1
  20. Design and Implementation of WEB automatic text classification,TP391.1

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net  Mobile