Dissertation > Excellent graduate degree dissertation topics show

Research on Named Entity Disambiguation in Deep Web

Author: ZhangRui
Tutor: ShenDeRong
School: Northeastern University
Course: Computer Software and Theory
Keywords: Deep Web named entity disambiguation XML Schema similarity affinitypropagation clustering algorithm
CLC: TP311.13
Type: Master's thesis
Year: 2010
Downloads: 29
Quote: 0
Read: Download Dissertation

Abstract


Due to the heterogeneity and autonomy of the Web databases, it is a challenging task to integrate the results extracted from various Web databases. The entity records in Deep Web are generally redundant or with multiple forms, which can not meet the data quality well (including data consistency, accuracy and completeness). Therefore, named entity disambiguation is an essential part of data integration during the data cleaning processing. In this thesis, we put forward a named entity disambiguation model used in Deep Web, which consists of data preprocessing module and similar entity clustering module.First, we apply XML schema to describe the global mode of entity records in data integration, which can achieve a unified format of entity records attribute information and solve the mode conflict between heterogeneous data sources; meanwhile, by defining different data types, we can obtain normalized data. Second, we propose the cluster-based named entity disambiguation algorithm, which consists of three parts:entity attribute similarity calculation, entity record similarity calculation and similar entity clustering. Take the entity records with global mode as input, we define constraint rules according to the domain knowledge, and choose the suitable similarity algorithm to calculate the entity attribute similarity; and then we calculate the entity record similarity according to the feature weights; finally, based on the entity similarity matrix, we employ the Affinity Propagation clustering algorithm to disambiguate the named entity records. The outputs are several distinct clusters, and each cluster represents a single entity of the real world.Thanks to the named entity disambiguation algorithm, we can obtain consistent and accurate entity record information, which is helpful to enhance the data quality during the multiple data sources integration in Deep Web and improve the user experience. The experiment results show our named entity disambiguation model is feasible and efficient.

Related Dissertations

  1. Syntactic Features Based Pronoun Resolution,TP391.1
  2. Research of Multiple Emails Automatic Summarization,TP391.1
  3. Research of IRC Botnet Detection Based on Behavior,TP393.08
  4. Research on Auto-Evaluation Method of Programming Based on Similarity,TP312.1
  5. Comprehensive Evaluation of Flue-cured Tobacco Quality in Pingdingshan and Comparative Analysis with American Tobacco,S572
  6. The Impact of Tourism on Typical Vegetation in Luya Mountain Nature Reserve, Shanxi Province,S759.9
  7. Ontology -based Semantic Web service matching and composition method,TP393.09
  8. Research on Method of Deep Web Schema Matching Based on Query Interface,TP311.13
  9. Study on Data Extraction and Semantic Annotation for Specific Field Deep Web,TP311.13
  10. The Research of Incremental Deep Web Acquisition Technology,TP393.09
  11. Based on the Deep Web of Books Information Integration and Inquires the System,TP311.52
  12. Outlier Detection and Semi-supervised Clustering Algorithm Based on Shared Nearest Neighbors,TP311.13
  13. Measure of Fuzzy Information and Its Application,O159
  14. Study and Implement on Spherical Harmonc Based 3D Models Retrieval System,TP391.41
  15. The Research on Deep Web Interfaces Integration and Query Results Ranking,TP274
  16. The Effects of Adult Attachment and I-sharing on Interpersonal Attraction,B849
  17. Research of Bi-directional Transformation between XML Data and Relational Data with Semantic Constraints,TP311.13
  18. Sentence Similarity Computing Research and Application of Intelligent Question Answering System,TP391.1
  19. The Research on Key Technologies for Web Information Personalization Collection and Management,TP393.09
  20. Synthesis and Design of Microwave Filters,TN713
  21. Finding Web Services Based on Clustering Probabilistic Semantic Approach,TP393.09

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer software > Program design,software engineering > Programming > Database theory and systems
© 2012 www.DissertationTopic.Net  Mobile