Dissertation > Excellent graduate degree dissertation topics show
Research on Text Clustering Based on Hownet
Author: ZhangLong
Tutor: ChaiXin
School: Hebei University of Technology
Course: Computer technology
Keywords: text clustering vector space model hownet textual similarity
CLC: TP391.3
Type: Master's thesis
Year: 2012
Downloads: 21
Quote: 0
Read: Download Dissertation
Abstract
|
K-Means algorithm is a classical algorithm of data mining technology, and it has the advantage of brief form and low time and space cost. It is also used widely in text mining. The paper researches on the key technology and algorithm in text clustering and puts forward a new method of calculating the similarity of texts based on hownet and improves the K-Means algorithm.The main work of the paper is to explore the effect of three text similarity calculating methods on K-Means algorithm. Using the classical vector space model based text similarity calculating method, hownet based text similarity calculating method and position information involved text similarity calculating method, the paper completes K-Means algorithm. To define the hownet based text similarity calculating method, the paper put forward a new way of generating vector space. It use the words of one text to generate a vector for the text,thus, the dimension of the vector equals to the number of words in the only text but not the number of words in all the text set. In this method, the high dimension and sparsity is reduced. The paper also talks something about the relation between the space and Euclid space. To define the position information involved text similarity calculating method, The paper also put forward that the similarity of two words should be decided by the words meaning similarity and position similarity. The paper also explore the method that how to correct the similarity of two words.
|
Related Dissertations
- Research and Implementation of Mining Implicit User Interest,TP311.13
- Research on Focused Crawler Based on SVM Classification Algorithm,TP391.3
- Evolutionary Clustering Algorithm and Its Application,TP311.13
- Research on Cross-document Coreference of Chinese Person Name,TP391.1
- Research of Text Clustering on Food Complaint Documents Based on Ontology,TP391.1
- Research on Automatic Marking of Subjective Questions Based on the Computation of Text Similarity,TP391.1
- Tracking Events for Food Complaint Documents Based on Ontology,TP391.1
- Design and Implementation of the Character Classification System Used in Search Engine,TP391.3
- Research and Implement of the Theme Crawler for Automotive Industry,TP391.3
- Research and Application on Short Message Text Clustering,TP391.1
- Research for Event Extraction Method in Specific Domain Based on Tree Conditional Random Field,TP391.1
- Research on Interest Model of User Behavior,TP274
- Term Weight-Based Chinese Text Classification Algorithm,TP391.1
- Online Education News Text Categorization System Design and Implementation,TP391.1
- Forum message text for effective data extraction studies,TP391.1
- SMBSDD: An improved search mechanism in unstructured P2P networks,TP393.02
- Web-based social network search method to judge the identity of names,TP393.09
- Web news hot discovery system design and implementation,TP393.09
- Research on multi-role social networks,TP393.09
- Research of WEB Mining Applying in the Monitoring System of the Network Policing Information,TP393.09
- Research and Realization of Web 2.0-Based Knowledge Management System,TP311.52
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Retrieval machine
© 2012 www.DissertationTopic.Net Mobile
|