Dissertation > Excellent graduate degree dissertation topics show

Study on the word and sentence similarity algorithm in network study platform

Author: HuangNi
Tutor: WuXiaoJun
School: Shaanxi Normal University
Course: Applied Computer Technology
Keywords: distance learning platform segmentation algorithm semantic similarity maximum matching
CLC: TP391.1
Type: Master's thesis
Year: 2012
Downloads: 74
Quote: 0
Read: Download Dissertation

Abstract


With the continuous development of the social information process, information is increasingly in demand and dependence. It has become the focus of the research to approach the access to information from the mass of useful information fast and efficiently. The language is the main carrier of information thus the requirement for language processing is increasing in the modern society. With the application of computers and the Internet popularity, Chinese information processing technology is facing new and greater challenges and opportunities. Chinese Information Retrieval can help users find useful knowledge on their own.As a basis for natural language understanding, the sentence similarity computing is a hot and difficult problem in information processing technology research. The status of the research has a direct impact on the field of intelligent question answering systems, machine translation, and information retrieval. Because of the complexity and evolution of Chinese language, the similarity calculation is still in the application of the initial stage. Since information processing in the national construction is of great importance, it is necessary to research on the similarity computation in-depth. Chinese sentence processing concludes word segmentation, the contents of the sentence similarity algorithm research. The identification of sentence similarity is based on the correct word, so the word quality determines the accuracy of the results of similarity calculation. The paper mainly discussed sentence segmentation technology and the sentence similarity, and it is verified by experiments the effectiveness of the proposed segmentation algorithm and sentence similarity calculation method. Finally, online learning intelligent question answering system is used as the media with sentence similarity calculation in the "question-answer" library.The main work of the paper is as follows:(1)The segmentation algorithm based on the word length is proposed in this paper, which can solve the word length limit of the maximum segmentation algorithm.(2)Semantic similarity algorithm for multi-factor is proposed in this paper. The method takes the similarity of the word form (key words), semantic similarity, sentence similarity of the three aspects into account to calculate the similarity of sentence meaning.(3)A detailed study is made for similarity calculation method based on keywords, sentence and semantic. For the calculation of similarity of word type (key words), vector space model-based approach only consider sided character of the Frequencies of Keywords and other surface information to calculate the similarity. This article will be fully effective in the information such as the order of key words, sentence length to use, consideration of the Frequencies of key words, word spacing, word order and sentence length and other information, so that the word shape similarity calculation results are more accurate. Semantic similarity, sentence affect the right weight for different keywords, the introduction of the weight of each word in sentence is used to calculate the sentence semantic similarity.

Related Dissertations

  1. Road extraction algorithm based on region segmentation of remote sensing image,TP751
  2. Ontology -based Semantic Web service matching and composition method,TP393.09
  3. Segmentation of cDNA Microarray Image Using Fuzzy C-means Algorithm Optimized by Particle Swarm,TP391.41
  4. Research on Relationship Extraction Based on Semantic Pattern Matching in Web Environment,TP391.1
  5. Based on web service research and application of distance education system,TP311.52
  6. Maximum Matching Problem of DNA Algorithms Rearch,TP301.6
  7. Research on Image Segmentation Method Based on Measure of Medium Truth Degree,TP391.41
  8. Research on the Fingerprint of Short Texts,TP391.1
  9. Optimization Techology Study and Implementation of Web Pages Ranking for Meta Search Engine,TP393.09
  10. Research on Automatic Construction Technology of Chinese Verb Frame Base,TP391.1
  11. The Research of Sentiment Analysis Techniques for Short-Texts,TP391.1
  12. Small Molecule Drug Targets’ Research of Congenital Heart Block (CHB),R714.5
  13. Study on Exploration of Reputation Dimensions Based on Text-Clustering and Corpus,TP391.1
  14. The importance of link-based prediction algorithm for dynamic linking,TP393.03
  15. Ontology-based Knowledge Retrieval System Smart Grid,TM76
  16. Based on Chinese Wikipedia semantic correlation computation Research and Implementation,TP391.1
  17. Research on Protein Complex Identification and Visualization in Protein Interaction Networks,TP391.41
  18. The Application Research of Chinese Word Segmentation Algorithm in GIS,TP391.3
  19. Combination Clustering Algorithm Based on Density and Grid and Its Application in Image Segmentation,TP391.41;O157.5
  20. Research and Implementation of K- means clustering of remote sensing images and watershed segmentation algorithm,TP751
  21. Research on the Instant Task-Solve-Oriented Composition of Semantic Web Services,TP393.09

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net  Mobile