Dissertation > Excellent graduate degree dissertation topics show
Semantic Search for Wikipedia
Author: LiuQiaoLing
Tutor: YuYong
School: Shanghai Jiaotong University
Course: Applied Computer Technology
Keywords: Semantic Search Wikipedia Large-scale knowledge base Mixed query capabilities
CLC: TP391.3
Type: Master's thesis
Year: 2009
Downloads: 131
Quote: 3
Read: Download Dissertation
Abstract
|
Wikipedia, as one of the most famous Web 2.0 applications, collaborative editing of a large number of volunteers, becoming the world's largest free online encyclopedia. Wikipedia content and articles of organization makes it a convenient online knowledge portal, which is a key factor in making it a great success. Wikipedia, almost every article of an entity described in detail, rich in text messages. In addition, the article also contains a wealth of structured information, such as the class system and the information box data. Recently, more and more these structured information extracted from Wikipedia and expressed as semantic data. However, existing search engine does not make full use of Wikipedia structured and semantic data, thus greatly limiting the use for the large-scale knowledge base. With the rapid development of Wikipedia, text and semantic information is also increasing. Thus, how to better, faster, and more intelligent access to the information becomes an important issue. This requires an effective search service based on text and semantic information. Semantic search is the key technology to solve this problem. Therefore, this research will be applied to semantic search technology Wikipedia build on Wikipedia semantic search system, which allows users to have more powerful search capabilities to facilitate full use of the large-scale knowledge base. Build on Wikipedia semantic search system has three difficulties: how to more fully in semantic data mining Wikipedia to build a large-scale knowledge base; how to provide as powerful as possible mixed query capabilities, supports a mixture of text messages and semantic information query and ensure its efficiency; how user-friendly search interface to support the powerful mixed query capabilities while still maintaining the ease-of-use features. Taking into account existing semantic data based on Wikipedia boxes and text extraction method has lower coverage article, this article further tap the knowledge implicit in the Wikipedia classification system proposed Wikipedia-based parent-child category to automatically extract The method of the tuple data. Wikipedia-based semantic data and rich text information, this paper proposes an extension of the new method of information retrieval engine index structure and function to provide efficient mixing queries. First, we define a tree mixed query, but also has the ability to structured query and keyword search. Then we extend the inverted index technology widely used in the information retrieval engine to support the relationship between the index and the mixed query processing algorithms broken down into the basic operation of inverted index, and to achieve consider the relationship between sorting mechanism. Using this method can be indexed on the mass of semantic data and text information, and can efficiently handle the user's hybrid query. Design-friendly search interface, this paper extends the widely used multi-angle structured data browse and search interface, allowing users to start from the familiar keyword search Back to search results at the same time automatically results in the category and relationship prompts guide the user to iterate the process and complete the complex mixture of query. Experiments show that the method based on the Wikipedia category automatically extracted semantic data beyond the previous work on the coverage of the article, and further expansion of the triple data. In this paper, the realization of semantic search system under the premise to ensure efficient and easy-to-use, greatly enhanced the ability to query the user on Wikipedia, improved search results, allowing users to more fully take advantage of a wealth of knowledge in Wikipedia.
|
Related Dissertations
- Based on Wikipedia disambiguation of named entities,TP391.1
- Based on Chinese Wikipedia semantic correlation computation Research and Implementation,TP391.1
- Research on Ontology-Based ITSM Knowledge Representation and Retrival,TP393.09
- Mining Semantic Knowledge from Chinese Wikipedia,TP391.1
- Research on Ontology-based Semantic Information System,G354
- Query Expansion Based on User Log Clustering,TP393.09
- Distributed intelligent semantic search method based on cloud computing,TP391.1
- Research on Semi-Automatic Construction of Ancient Agronomy Ontology and Its Semantic Retrieval,S-09
- Research on Selection Mechanisms for Web Services in Peer-to-Peer Networks,F49
- P2P environment content-based image retrieval technology,TP391.3
- Image description and retrieval based on the MPEG-7,TN919.81
- Research of Semantic Retrieval Based on the Ontology,TP391.3
- The Web2.0 era wikis Status and Development of,G206
- Research and Implementation of Semantic Retrieval Model Based on Jena,G354
- Research on Text Classification Based on a Keyword,TP391.1
- Wiki publishing model study,G230.7
- POS Tagging System on Hierarchical Classification Labels,TP391.1
- Based on semantic P2P search research and simulation to achieve,TP393.02
- Study the Human Dynamics in Wikipedia,C912
- Research on Semantic Search in Retrieval System of ERP,TP391.3
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Retrieval machine
© 2012 www.DissertationTopic.Net Mobile
|