Dissertation > Excellent graduate degree dissertation topics show

The Study on Automatical Domain-Specific Knowledge Extraction from Websites Based on Bootstrapping

Author: KangZuo
Tutor: MaJun
School: Shandong University
Course: Computer System Architecture
Keywords: Domain knowledge extraction semi-structured websites pattern learning ontology mapping bootstrapping
CLC: TP391.1
Type: Master's thesis
Year: 2012
Downloads: 77
Quote: 0
Read: Download Dissertation

Abstract


With the striking expansion of Internet and the speedy development of online applications, the scale of web information is becoming larger and larger. Web becomes an important knowledge repository; it is high desirable for people to obtain wanted information efficiently. The web contains a large amount of semi-structured domain knowledge on movies, books, restaurants and so on, which is closely related to our life. Nowadays, people can achieve information retrieval from web by search engine easily, but the results are not so credible. Meanwhile, domain-specific knowledge encoded in semi-structured pages is often come from underlying databases of commercial providers, it is difficult for search engines based on keyword-matching to crawl and index such knowledge. How to automatically extract and organize such domain-specific knowledge has become a research hotspot in information extraction field.According to analyzing the current web information extraction methods, this paper employs tag path technique to represent HTML page instead of DOM tree. Such representation technique reduces the scale of tags in tag tree dramatically and improves the performance of our algorithm. This paper proposes a novel automatically extracting domain-specific knowledge algorithm based on bootstrapping from semi-structured websites:Domain-specific Knowledge Extraction from Websites, DKEW. DKEW utilizes ontology to unify the labeling of domain-specific semi-structured data extracted, which is helpful to organize and query knowledge. DKEW firstly clusters the target pages based on tag path technique to filter noisy pages and keep detail pages which contains more semi-structured information. To extract information from detail pages, we propose a novel pattern based on tag path representation. For detailed cluster, DKEW uses a machine learning method to learn the pattern with the help of the known seed. Then DKEW automatically extracts domain knowledge using the learned pattern and maps the extracted knowledge to predefined ontology in the form of table. New mapped knowledge with high reliability will be used to expand domain seed and ontology for the next iteration process. Finally, DKEW utilizes a bootstrapping solution to iterate the whole processes and integrates them as an automatic information extraction algorithm. DKEW only requires a tiny human effort to initialize seed by annotating a few pages from Wikipedia in the specific domain. This paper crawls large-scale data from several popular domains by our self-defined web crawler to verify DKEW. Experimental results show that DKEW is better than RoadRunner, which is an automatical web information extraction method, in terms of effectiveness and efficiency. Moreover, our approach employs automatic mapping technique instead of manual labeling of RoadRunner, which saves a lot of labor and time. Experimental results also verify the generalization of DEKW in large-scale domains.

Related Dissertations

  1. WordNet and the \,G254
  2. Research on Mapping RDF/RDFS to Relational Database Schema,TP311.13
  3. Printers based on natural language HCI Research and implementation,TP11
  4. Semi-automatic ontology based on relational database construction method,TP311.13
  5. Research on Mapping Relational Database Schema to Ontology,TP311.13
  6. Research on Medical Case-Based Reasoning Based on Ontology,R197.3
  7. Research on Graph Partitioning-based Large-ontologies Partitioning and Mapping,O157.5
  8. Immersed Power Transformer Fault Diagnosis,TM407
  9. Research on Laplace Matrix-based Partitioning and Mapping of Large Ontology,TP391.1
  10. Research on Key Technology of LTE Intrusion Detection Systems,TP393.08
  11. Body in the integration of medical information and research,TP391.1
  12. Research on Large-Scale Ontology Partition and Mapping Based on Fragment,TP391.1
  13. Study of Ontology Mapping Based on BP Neural Networks,TP311.52
  14. Study on Sentiment Polarity Analysis of Subjective Text,TP391.1
  15. Study on Ontology Learning and Its Application to Semantic Retrieval,TP391.1
  16. Pricing Switch-Option-Embedded Notes in a LIBOR Market Model,F224
  17. A Supervised Way in Word Sense Disambiguation,TP391.1
  18. The Identification of Vehicle Based on Bootstrapping,TP391.1
  19. Research on Sentiment Classification of Product Review Texts Based on Maximal Consistent Block,TP391.1
  20. The Research of Ontology Mapping Method Based on Semantic Similarity,TP391.1
  21. Ontology Mapping and Its Application in Inter-Smartspaces,TP18

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net  Mobile