|
In the information society, with the rapid development of computer technology, communication technology, network technology, the network has become an important way for people to obtain information. It is predicted that online media will replace print media in the near future become the main channel of access to information. The face of the Internet continues to explode growth, how to get to quickly find the information needed by the user to become a problem. To solve this problem, the search engine technology came into being. However, the search engine technology is not perfect, there are a large number of duplicate pages return a result set. These duplicate pages mainly because the reprint of the site, it not only increased the difficulty of the users to retrieve the desired information, but also a waste of storage space. Therefore, the detection of a large number of duplicate pages, to avoid duplication of storage, information retrieval is done quickly and accurately is a meaningful work. On the other hand, the Internet e-commerce environment, the digital goods easily be illegally copied and spread, which will undoubtedly hinder the healthy development of e-commerce. Copy detection technology to some extent, it could help solve these problems. Domestic Chinese text copy detection research is not mature, not a perfect solution, many problems to be solved. Text copy detection technology status and development brief review of the text copy detection technology related technologies are studied, and the Chinese text pre-processing, text blocks, and feature extraction strategy, text similarity metrics method, copy the text detection algorithm and other issues are discussed in detail. Copy detection algorithm, algorithm performance, advantages and disadvantages are analyzed and compared. Then, this article focuses on the Chinese text copy detection algorithm, and propose two improved algorithms. Tradition based on N-Gram Chinese text copy detection method to avoid the text of the word, but in the text feature extraction is not perfect. Improved combine text copy detection method based on N-gram, N-Gram method with the sliding window technique, so that extraction of a small amount of text features can more accurately calculate the text similarity, thereby improving the efficiency of the algorithm. Experiments show that the method is effective, the ideal recall and precision rate. This paper presents an improved copy detection method based on the text of the sentence comparison. Methods sentences - document storage structure of multi-layer index during text copy detection, the sentence can directly find the existence of the sentences of all documents. Finally, with improved two text copy detection method for testing artificial label text test corpus, the test results were analyzed and compared with the precision and recall performance evaluation of two replication detection method The test results were evaluated. Experimental results show that the improved two replication detection method can achieve good detection results.
|