Dissertation > Excellent graduate degree dissertation topics show

Research on Wu-Manber Mutiple String Matching Algorithm

Author: MoDeMin
Tutor: LiuYaoJun
School: Taiyuan University of Science and Technology
Course: Computer Software and Theory
Keywords: String matching Exact string matching Wu-Manber Matching algorithm Information Retrieval
CLC: TP301.6
Type: Master's thesis
Year: 2008
Downloads: 83
Quote: 0
Read: Download Dissertation

Abstract


String matching in the field of computer research an ancient, classic and was the subject of extensive research, is one of the key technologies of the field of information retrieval and computer biology. In today's age of the Internet, the ever-changing needs of matching algorithm, the growing demand for real-time processing. The new challenge of the original string matching techniques, the need to improve and optimize the original string matching techniques. This paper is efficient multi-string matching algorithms - Wu-Manber algorithm. Wu-Manber algorithm hash table HASH entry: the lack of a linked list HASH (), raised the processing algorithm based on non-empty common sub-suffix pattern; then synthesize former SHIFT table on the Wu-Manber algorithm improvements, more efficient string matching algorithm; then put forward a Wu-Manber algorithm improvements in large-scale string matching. The research for this article, as follows: Wu-Manber algorithm improvements based non-empty common sub-suffix: suffix pattern common sub insufficient for the Wu-Manber algorithm, this paper proposes a non-empty public sub suffix The mode of processing algorithms. The algorithm in the linked list of the same next non-empty mode of public sub-suffix to come together, to further reduce the average length of the next linked list. Reduce the number of comparisons of the character in the matching process, thereby improving the operating efficiency of the algorithm. This article related given Sogou laboratory, document full-text retrieval experiments, compare and Wu-Manber algorithm, the improved algorithm proposed by predecessors. Experimental results show that the improved algorithm proposed in this paper to effectively reduce the number of times a character comparison in the matching process, thereby improving the speed and efficiency of the matching. Wu-Manber algorithm improvements: a proposed algorithm has been improved in the next linked list are classified common sub suffix contains a non-empty linked list nodes mentioned before; integrated before made the \The new and improved algorithms make full use of the advantages of the two algorithms above, so that the character number of comparisons in the matching process has been further reduced. In the experiment, the new and improved Wu-Manber matching algorithm achieved higher efficiency, increased to 4.6% or more than the original algorithm. Improvement in the large-scale pattern string of 3.Wu-Manber algorithm: a proposed algorithm has been improved, the original algorithm next linked list the junction Same_Subsuffix domain split into two sub-domains, such that the string in the large-scale mode , the number of comparisons of the characters in the search process is further reduced, a further increase in the efficiency of the new algorithm than the original algorithm. The experimental results show that, when the pattern string less, the efficiency of the new algorithm, with some loss compared to the original algorithm. With the increase of the pattern string, the new algorithm has a higher efficiency. Therefore, the new algorithm is faster than the original algorithm has a greater scope of application.

Related Dissertations

  1. Research and Implementation of Exact String Matchiing Algorithms,TP391.41
  2. Research of Question Answering System Based on the Analysis of Lexical and Semantic Meanings,TP391.1
  3. Establishment and Update of Similar Users’ Cluster in Personalized Information Retrieval,TP391.3
  4. Research on Query Expansion Technique of Retrieval System in Biomedical Field,TP391.3
  5. Research and Implementation of Retrieval System on Massive Mail,TP393.098
  6. Research of the Model of Enterprise Competitive Intelligence Collection System Based on Cross-Language Information Retrieval,TP391.3
  7. An Optimized Transmission Algorithm Based on Network Coding in Wireless Network,TN911.22
  8. Research and Implementation of Intelligent Agent-Based Personalized Information Retrieval System,TP391.3
  9. CUDA-based regular expression matching system design and implementation,TP311.52
  10. Windows Kernel Rootkit Detection Technology Research,TP309
  11. Non-negative matrix factorization based on sparse image retrieval,TP391.41
  12. SURF algorithm based on multiple images of three-dimensional model reconstruction method,TP391.41
  13. Research on cross-language text categorization,TP391.1
  14. Classification model based monitoring of e-commerce Prohibited Research and Implementation,TP393.09
  15. Based on the DM6446 electronic image stabilization technology Research and Implementation,TP391.41
  16. Publish subscribe model based design and implementation of middleware,TP311.52
  17. Sort learning loss function studies,TP181
  18. Click-based user clustering research,TP311.13
  19. Semantic Document Retrieval for English to Chinese Cross-Language Question Answering System,TP391.1
  20. The Implementation and Research of the Probabilistic Latent Semantic Analysis Model in the Search Engine’s Business Text Classification System,TP391.1
  21. CPU GPU -based heterogeneous platforms string matching algorithm and implementation,TP301.6

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > General issues > Theories, methods > Algorithm Theory
© 2012 www.DissertationTopic.Net  Mobile