Dissertation > Excellent graduate degree dissertation topics show

Research on Big Data Mining Algorithm Based on Top-k Subgraph Pattern Matching

Author: HeJian
Tutor: LvFeng
School: Wuhan University of Technology
Course: Communication and Information System
Keywords: Big Data Mining Top-k Subgraph Pattern Matching Top-k GPM JoiningAlgorithm Spanning Trees’ Cost Forecast
CLC: TP311.13
Type: Master's thesis
Year: 2013
Downloads: 18
Quote: 0
Read: Download Dissertation

Abstract


With Big Data Time coming ahead of time in the Internet industry, a series of emerging and effective information extraction technology specially for big data mining has begun to attract the general concern of the researchers to research. The amount of data information generated every year keeps on the state of the explosive growth all the time, so referred to data are not also presented the larger trends in quantitative, in addition, accompanied by a qualitative change; Even as we know, traditional data mining methods are based largely on the relational database to be developed, so this way can’t apply to that situation such as data types being more diverse, data structure relations being more complex. But graph mining techniques can always be applied to modeling and mining operation with big data when in need of conjunction with graph theory method such as graph query, graph traversal, graph isomorphic and so on. As well as it mainly applied to library management system (such as books’information retrieval) and social networks (such as the relationship matching between the characters) and Bioinformatics Engineering (such as PPI and Genetic Engineering) which both contain a lot of data.In this paper, the Top-k Subgraph Pattern Matching (GPM) algorithm belongs to a typical graphical big data mining technique, which is based on the isomorphism principle of graph theory and applied to single-source data graph with specific incidental tag attributes. The purpose of this kind of mining algorithm is to get a large number of matching results from the mass source data graph on conditions of satisfying the query graph labeling conditions and structure conditions (based on path).And it mainly apples to those data types for the Resource Description Framework (RDF) data.The author’s main researching work as follows.(1) Introduced some basic concept and related algorithms of Big Data Mining’s graph pattern matching, mainly introduced Top-k GPM Join Algorithm which has some reference and comparison value. (2) This paper put forward a kind of efficient and universal Top-k Subgraph Pattern Matching Plan. About its usability, it take consideration of the applicability of the algorithm for the querying graph with cycles or without cycles to get the accurate Top-k Matching Results. In the matching process for the querying graph with cycles, it take the optimal spanning tree, which is chosen by Spanning Trees’Cost Forecast Matching Plan, as the querying subtree to develop into Top-k matching results by expanding edges.(3) In the phase of algorithm’s performance testing, the RDF graph source data adopted comes from real DBLP data. My matching plan of this algorithm not only got all accurate Top-k matching results, but also verified that it is viable to choose the optimal spanning tree to match by Spanning Trees’Cost Forecast Matching Plan. And then this algorithm in this paper has taken time and space performance into overall consideration by comparing with Top-k GPM Joining Algorithm. The results of performance testing showed that its time performance have been improved greatly on the cost of consuming more extra memory space.

Related Dissertations

  1. Research on F&B Index Structure Supporting XML Query,TP311.13
  2. Query Processing and Optimization in Massive Multi-Database Integration,TP311.13
  3. Research on Parallel Frequent Graph Pattern Mining,TP311.13
  4. The Design and Implementation of DICOM Middle Software and Access Control Model in Formation Integration Platform,TP311.13
  5. Research on K-means Optimization Clustering Algorithm,TP311.13
  6. Public Security 110 Command Decisions Business Systems,TP311.13
  7. Clustering Method Research Based on Divided and Conquered Method,TP311.13
  8. Incomplete information on the completeness of the system and its knowledge acquisition,TP311.13
  9. Evolutionary Clustering Algorithm and Its Application,TP311.13
  10. Research on Method of Deep Web Schema Matching Based on Query Interface,TP311.13
  11. Application of Data Mining in the Analysis of Higher Vocational Colleges’ Achievement,TP311.13
  12. Design and Implementation of Course Assessment and Analysis of Decision System Based on Data Mining,TP311.13
  13. Study on the Key Technologies of Domain Database-Oriented Question and Answering System,TP311.13
  14. Research and Application about Ditributed Transaction Process Protocol,TP311.13
  15. Web Usage Mining and the Research of Personalized Recommendation,TP311.13
  16. An Improvement of Cluster on Phylogenetic Profiling Method,TP311.13
  17. Application of the Technique of Data Mining in Educational Administration,TP311.13
  18. An Algorithm on Clustering and Anomaly Detection for Multiple Data Streams,TP311.13
  19. The Design and Implementation of Distributed Database Based on Oracle of Mobile CRBT,TP311.13
  20. Government Decision Support System of Industrial Economics,TP311.13
  21. The Research on Feature Selection for Data Stream,TP311.13

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer software > Program design,software engineering > Programming > Database theory and systems
© 2012 www.DissertationTopic.Net  Mobile