Dissertation > Excellent graduate degree dissertation topics show
The Hot-topic Discovery Based on Density Clustering of Feature Words and Similarity Calculation
Author: HanChenJing
Tutor: LuoHuiQiong
School: University of Electronic Science and Technology
Course: Computer System Architecture
Keywords: Hot-topic Discovery Problem Title Keywords Extracting KeywordsWeight Similarity Calculation Density Clustering Algorithm
CLC: TP391.1
Type: Master's thesis
Year: 2013
Downloads: 29
Quote: 0
Read: Download Dissertation
Abstract
|
With the rapid development of the Internet, Network media has become animportant source of information for people and network public opinion information’sguiding role has become more and more important. At the same time, Internet publicopinion has become a kind of important form of the social intelligence. Althoughnumerous and varied network information plays an active role to the development ofthe society, the Internet public opinion happened widely and spread fast. So thereexists the problem of information security. The harmful information, such asreactionary, obscene, superstition and so on, spread over the Internet, which seriouslyharm the national security and social stability. How to find the Internet hot-topicaccurately and timely has become a hot topic at home and abroad in logistics.Optimizing and improving the related technology of the Internet public hot-topicproblem can grasp the Internet public hot-topic better and improve its efficiency andaccuracy. Internet public hotspot problem mainly considers the feature extraction andclustering algorithm technology. High quality and high speed text clusteringtechnology will distinguish the large numbers of text information into somesignificative clusters. Many researchers have been paying much attention to theclustering algorithm over the past years, such as ARHP, PDDP, K-means, PAM,DBSCAN and OPTICS. These algorithms can cluster texts well. However, thesealgorithms have their limitations, when considering to keywords extraction andsimilarity calculation. However, this paper represents the Internet hotspot problem as abased on keywords extraction and density clustering and similarity calculation problemand optimizes the keywords extraction and similarity calculation at the same time.Under the method of the Internet hotspot, according to the existing problemabout keywords extraction method and optimizing the similarity calculation. Then,new improved algorithm, based on title keywords extraction and improved similaritycalculation and related clustering algorithm, has been designed to realize the Internethotspot discovery. These algorithms are focusing on improving the accuracy aboutInternet hot-topic discovery. The main studies in this paper are as follows: (1) The high or low quality of keywords is closely related to the major points ofthe articles content. Only when fully understanding the content and the exact meaningof the words, which will extract keywords effectively. In order to get high quality ofkeywords, present a method which extracts keywords from the title.(2) When discovering the hot-topic, whether use classification algorithm orclustering algorithm, it needs to analyze the similarity between two vectors forreflecting the text’s real similarity better. According to the giving similaritycomputational formula, present a method which considering the keywords weight intothe similarity formula.(3) Combination of the above-mentioned methods, this paper base on densityclustering bring forward the algorithm for hot-topic discovery based on title keywordsand density clustering and similarity calculation.Finally, via compare experiment data testing and above-mentioned algorithmanalysis, the algorithm shows preferable performance.
|
Related Dissertations
- Methods study recommended expert assessment of the project,TP391.3
- Uncertain Information Vehicle Routing Problem Based on Case-based Reasoning,G358
- Research of Multiple Emails Automatic Summarization,TP391.1
- An Improved Discrete Particle Swarm Optimization for Ontology Matching,TP391.1
- The Analysis of Cellular Signal Transduction Network and the Similarity Evaluation on the HPLC Fingerprint of Breviscapine Injection,R286
- Research on Multi-strategy Ontology Mapping Based on Fragments Separated,TP391.1
- The Field Specific Automatic Question Answering System Based on LDA Model,TP391.1
- Research on the Method of Auto-discovery and Verification of Topic-Websites,TP391.1
- The Research on the Method for Webvideos Clustering Based on Multi-Modal Strategy,TP391.1
- Research and Implementation of a SNS Question Answering System,TP391.1
- The Design and Implementation of Project Information Service Platform,TP391.1
- Design and Application on the Scalable Multi-dimensional Recommendation Engine Framework,TP391.3
- Key Technologies Research on Personalized Information Service Based on Ontology and Context-aware,TP391.1
- Research and Implementation of Personalized News Recommendation System Based on Collaborative Filtering,TP391.3
- The Query Recommendation Algorithm Research Based on the Search Logs,TP391.3
- Personalized Recommender Systems in E-Commerce,TP391.3
- Research on Movie Similarity Calculation Using Multi-feature Method,TP391.41
- The Research of Similarity Comparising Method of Plant Shapes,TP391.41
- Research on Semantic Analysis and Mapping Method of Semi-structured Data,TP311.13
- Research on Evaluation Assistance System of S&T Project Based on Web Text Mining,TP311.13
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net Mobile
|