Dissertation > Excellent graduate degree dissertation topics show
Chinese and English Automatic Summarization Based on Topic Modeling
Author: ZhangMingHui
Tutor: ZhouGuoDong;WangHongLing
School: Suzhou University
Course: Applied Computer Technology
Keywords: topic modeling Multi-document summarization latent Dirichlet allocation natural language processing
CLC: TP391.1
Type: Master's thesis
Year: 2011
Downloads: 40
Quote: 0
Read: Download Dissertation
Abstract
|
With the rapid development of the computer technology and the Internet, various in formation is increasing explosively; people’s demand for precisely locating information give a strong impetus to the research in the natural language processing technology (NLP). Meanwhile, as the progressively research on cross-document information fusion technologies, multi-document summarization become a hot research subject, multi-document summarization can be used in question answering, search engines, topic detection and other applications.In this paper, we analysis the existing methods on automatic multi-document summarization deeply, and apply a topic model to the sentence silence detection. In addition, we use a dynamic model to control redundancy. At last, we implement an automatic multi-document summarization system based on those methods. Experimental results on TAC2008 and TAC2009 corpus show that the system has a good Rouge performance.This paper mainly analysis the most two key multi-document summarization technologies:Sentence salience determination and redundancy control. In terms of sentence salience determination, we propose a sentence topic feature based on topic modeling. The results show that the topics feature plays a significant role in the MDS. And the combination of topic feature and other traditional features can also improve the system performance. In terms of redundancy control, we use dynamic modeling to control redundancies; we also design the update dynamic modeling for the update summarization task based on this. After using the update dynamic modeling, the summary can effectively avoid history redundancies. The results of TAC2008 corpus show that after combined the two strategies (Sentence salience determination and redundancy control) we can achieve a better system performance. Especially in the update summarization task, our result is better than the best result in the entry system. Finally, this paper also gives the evaluation of Chinese corpus before and after joining topic model and dynamic model. The result shows that topic modeling and dynamic model have equally effective on the Chinese corpus. However, the result of Chinese MDS is obviously worse than the one of English MDS, and the reason may be that the Chinese corpus needs more preprocessing which can affect the performance of the whole system.
|
Related Dissertations
- Word Sense Disambiguation Corpus Automatic Acquisition,TP391.1
- Combination of user and modeling of geographic information Images Themes,TP391.41
- Printers based on natural language HCI Research and implementation,TP11
- The Time Sketch of Country Style,J905
- Research of Word Sense Disambiguation Based on Word-sense Category Extending,TP391.1
- Research on Query Expansion & Key Technologies Based on Semantic Analysis,TP391.1
- Question and Answer Recommendation in Question Answering Communities,TP391.1
- Scenario Abstraction From Children’s Story for Interactive Animation,TP391.1
- Research on a Multi-Strategy Approach to Answer Extraction in Chinese Question Answering,TP391.3
- Research on the Topic-oriented Summarization for Web Documents,TP391.1
- Chinese-English machine translation system in a model of semantic disambiguation,TP391.2
- Design and Implementation of the list class automatic question answering system,TP311.52
- Probabilistic grammar based on binary combinations disambiguation model design and implementation,TP391.1
- Mechanical product design verb attribute structure Research and Application,TH122
- Chinese word segmentation system and its implementation,TP391.1
- Chinese new word recognition,TP391.1
- Based on statistical and semantic analysis of multi- document summarization,TP391.1
- A Fingerprint Engine for Author Profiling,TP391.3
- Research of Chinese and Japanese Question Classification,TP391.1
- Research of Automatic Summarization Based on Named Entity,TP391.1
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net Mobile
|