Dissertation > Excellent graduate degree dissertation topics show

Study on Chinese Text Similarity Computing Based on Word Segmentation

Author: ShenBin
Tutor: HuaBin
School: Tianjin University of Finance and Economics
Course: Business management
Keywords: Text Similarity Participle MM Law Text feature vectors TF-IDF
CLC: TP391.1
Type: Master's thesis
Year: 2006
Downloads: 1221
Quote: 43
Read: Download Dissertation

Abstract


In Chinese information processing, text similarity calculation widely used in information retrieval , machine translation, question answering systems, text mining and other fields, is a very basic and critical issues , has long been one of the hot and difficult . Computer processing of the Chinese culture with respect to the treatment of the West there is a greater degree of difficulty , embodied in the text of word processing . Segmentation is the Chinese text similarity calculation basis and prerequisite for using an efficient segmentation algorithm can greatly improve the accuracy of text similarity calculation . In this paper, for the common Chinese word segmentation algorithm analysis and comparison on the basis of an improved forward maximum matching segmentation (MM) algorithm and disambiguation strategy, word dictionary is created by , segmentation step and ambiguities on the treatment of new and improved methods to improve the completeness and accuracy of the segmentation . Subsequent analysis and comparison of the existing text similarity calculation method based on vector space model using the TF-IDF method combines previously proposed segmentation algorithm gives Chinese text segmentation and similarity calculation computer system implementation process, and technology text test example , the methods used for the validation. Study of this topic and its results for the Chinese information processing in a variety of fields, especially tech text similarity calculation and comparison , will have a certain reference value and a good prospect .

Related Dissertations

  1. Design and Implementation of the Character Classification System Used in Search Engine,TP391.3
  2. An Interactive and Intelligent Mechanism for Establishing Knowledge Structure of On-line Talents,TP391.6
  3. Sentiment Classification by Combining Lexicon-based and Machine Learning Methods,TP391.1
  4. Research on Similarity Computing Method for Domain Texts,TP391.1
  5. A Improved Text Similarity Model Based on PageRank Value,TP391.1
  6. Research and Application of Data Mining Based on web Literature,TP311.13
  7. Research on the Topical Search Engine Based on Semantic,TP391.3
  8. Research on Paper Similarity Based on Semantic Understanding,TP391.1
  9. Component test reuse Research and implementation,TP311.52
  10. Research on Parameters Correlation and Optimization in Text Similarity Measurement,TP391.1
  11. Chinese Text Similarity Algorithm Research Based on Semantic Similarity,TP391.1
  12. Armed Forces public opinion monitoring system design and implementation,TP393.09
  13. Research of Text Classification Algorithm Based on Biomimetic Pattern Recongnition,TP391.1
  14. Key Techniques of Text Ming on Criminal Cases,TP391.1
  15. Research and Implementation of the Text Cluster Based on Text Similarity Caculation,TP391.1
  16. Study of Chinese Text Similarity Based on Number Difference Gene,TP391.1
  17. Study on Information Retrieval of Quality Internet Public Opinion Monitoring System,TP393.09
  18. Research of Intelligent Remark Generation Based on GA-BP Nerve Network,TP391.1
  19. Chinese full - text similarity computation and application,TP391.1
  20. Chinese text based on vector space model similarity algorithm,TP391.1
  21. The Research on Graph Structure Representation Method Based Chinese Text Clustering,TP391.1

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net  Mobile