Dissertation > Excellent graduate degree dissertation topics show
Research on the Topic-oriented Summarization for Web Documents
Author: DengGuangXi
Tutor: LiPeiFeng
School: Suzhou University
Course: Applied Computer Technology
Keywords: multi-document summarization Web Documents sentence classification dependency relation sentence ordering
CLC: TP391.1
Type: Master's thesis
Year: 2011
Downloads: 11
Quote: 0
Read: Download Dissertation
Abstract
|
With the rapid development of the Internet, information explosion has become a very serious problem currently. It becomes more and more difficult for us to acquire information effectively for the huge set of web pages that contains much duplicate information. So how to provide concise information for users to meet their requirement is an urgent problem. To solve above problem, we make a study on the method of topic-oriented summarization for Web documents. The research can be summarized as follows:1. We did research on classifying those sentences in Web documents into each aspect of the special topic. We proposed a method which based on dependency relation to compute sentence similarity. A clustering approach based on that method also was applied to cluster those sentences into some clusters and the features (words and grammar) from each cluster were extracted to form patterns. These patterns can be used to identify those topic related sentences. Further more, a classifier was used to identify topic related sentences which adjusted the weight of word according to its level in a parse tree, selected some effective dependency relations as features and used syntactic tree as structure features based on verb-driven tree pruning.2. In the part of sentence selection, this dissertation proposed a novel sentence selection method. Firstly each candidate sentence was labeled according to those features of Web documents. Then the score of a candidate sentence was adjusted based on sentence similarity, document links, sentence anchors and the relative position between sentences. Finally, a MMI algorithm is used to choose sentence according to scores of candidate sentences.3. A novel context-based approach to sentence ordering for multi-document summarization was proposed. Our method first detected whether two summary sentences should be adjacent according to the similarity between one summary sentence and the context of the other summary sentence, and then computes the reliability based on the similarity and relative position. The first sentence is decided based on features of the sentence. The sentence that has the maximum adjacency reliability with the previous sentence will be selected as next one.Evaluation shows our approaches achieve high performance and the latter is better than the state-of-the-art ones.
|
Related Dissertations
- Research on Lexical Chains and PageRank Based Multi-document Summarization,TP391.1
- Research on Key Technologies of Chinese Multi-Document Summarization,TP391.1
- Chinese Query-Focused Multi-document Summarization Based on Cloud Model,TP391.1
- Graph model of knowledge discovery,TP181
- Multi-document Summarization Based on Basic Element,TP391.1
- Research on Event-based News Story Analysis Technology,TP311.5
- Research on the Decision Support System Model Based on Web Mining,TP393.09
- Multimedia information automatically Summary and its technologies,TP393.09
- Omit the performance of the end of the sentence spoken Japanese inspection,H36
- Research on Basic Algorithms of Digital Image Processing and Implementation with FPGA,TP391.41
- Research on Facial Feature Extraction and Matching Algorithms for Image Retrieval,TP391.41
- Research of High Speed Image Pre-processing System Based on FPGA,TP391.41
- Research on Algorithms of 2D Face Template Protection,TP391.41
- Research of Visualization Technology in the Virtual Test of Missile,TP391.9
- The Research and Implemention of Image Retrieval Based on User Interested Feature,TP391.41
- Research of Image Mosaic Technology,TP391.41
- Research and Implementation of Exact String Matchiing Algorithms,TP391.41
- Research of Question Answering System Based on the Analysis of Lexical and Semantic Meanings,TP391.1
- Research on the Classification Based on the Reconstruction of Solder Joint,TP391.41
- Tongue Feature Extraction and Research of Fusion Classification,TP391.41
- Research on Structure Transition Technology for SMT,TP391.2
- The Fatigue State Recognition of the Driver Based on Eye Detection,TP391.41
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net Mobile
|