Dissertation > Excellent graduate degree dissertation topics show

Research on System of Multi-field Information Extraction Based on Semantic Role and Concept Graphs

Author: YangXuanXuan
Tutor: ZhangLei
School: Northwestern University
Course: Applied Computer Technology
Keywords: information extraction semantic roles similarity calculating of conceptual graphs Semantic Computation HowNet
CLC: TP391.1
Type: Master's thesis
Year: 2010
Downloads: 136
Quote: 3
Read: Download Dissertation

Abstract


Information extraction is a hot research in the field of natural language processing. In abroad, the research work began in the last century and has achieved many results. About the research of Chinese information extraction is still at the initial stage. Other existing methods of extraction are basically based on statistics. This method is lack of semantic information to support. Using statistic method to extract information is low efficient and imprecise, because the approach ignores the semantic links among the words in sentence. Besides, this A simple probability model can not understand the semantics about the sentence at all, so the extraction results is in low quality, and could not satisfy the intelligent demand. In order to solve the deficiency of traditional extraction method, this article that based on the previous studies presents a new approach using semantic role and the semantic resources in the "Hownet" to extract information. This new model is a multi-field information extraction system which based on semantic information.In this paper, our research model is based on semantic roles and the conceptual graph. The relevant processes are as follows:First, labeling the semantic role of sentence. Later, going to the module of pretreatment and filtering out empty word and interjection. Secondly, according to the semantic information, using relevant Algorithm which paper presented to generate conceptual graphs. Thirdly, the relevant module calculates the similarities of conceptual graphs so as to identify areas of the scene. We have adopted a automatic classification method to generate extraction templates, this main idea come from prototype of Bootstrapping. At last, we should construct the rules of extraction. In this module, we take advantage of semantic role to generate the rules in order to improve the accuracy of extraction. Among them, the field division of the scene, template generation and extraction rules are the main content of this article study.In the final part of paper, we give the evaluation results of system. The results contain two aspects:vertical comparison horizontal comparison. In the section of vertical comparison, we extract the information of same domain but in different technology. In the horizontal comparison, we did the two experiments. One is the scene division, another is the cross-cutting information extraction. We use three different approaches to extract the Multidisciplinary message. Experimental results show that our method of extraction is feasible and effective. Besides, this approach improves the accuracy of information extraction system and the recall rate.

Related Dissertations

  1. Research on Domain Entity Attribute and Event Extraction Technology,TP391.1
  2. Based on high-resolution remote sensing data mining houses information extraction,TP751
  3. Research of Text Clustering on Food Complaint Documents Based on Ontology,TP391.1
  4. Related Studied on Information Extraction and Information Recommendation Based on Web Data Mining,TP393.09
  5. Research on Technology of Deep Web Oriented Data Extraction and Semantic Annotation,TP311.13
  6. Research on Temporal Information Recognition and Normalization,TP391.1
  7. Crawler and Incremental Update Strategy Research in Deep Web,TP391.3
  8. Statistic-based Automatic Keypharse Extraction and Summarization from Multi-document,TP391.1
  9. Research and Implementation of Topic-Based Event Fusion,TP393.09
  10. Research on Cross-lingual Word Similarity Computation,TP391.1
  11. Chinese BBS Information Extraction and Classification,TP393.094
  12. The Research of Paper in Quality Evaluation Based on Sciencepaper Online,TP393.092
  13. Research on Web Information Extraction Tool,TP393.092
  14. Research and Application on the Technology of Web Information Extraction Based on the HTML,TP393.09
  15. The Design and Implementation of Web Information Extraction System Based on Dom,TP393.09
  16. Design and Implementation of Ontology Based Extraction and Analysis Platform of Logistics Information,TP393.09
  17. Query Results Dealing Technology of Deep Web Based on Ontology and Conceptual Schema,TP393.09
  18. The User Attention Analysis and Individualized Study Based on "Chinese Scienecpaper Online",TP18
  19. The Research and Application of the Information Extraction Algorithm in FTIR Multi-spectral Microscopic Images,TP18
  20. Medical Advertisments Monitoring System Based on Web Content Mining,TP393.09
  21. Research on Auto-Construction Technology for University Teacher Social Network,TP181

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net  Mobile