Dissertation > Excellent graduate degree dissertation topics show
Weight calculation method research and application of statistical and rule - based feature weight
Author: YangYuZhen
Tutor: LiuPeiYu
School: Shandong Normal University
Course: Computer Software and Theory
Keywords: Network Information Filtering Information gain Text granularity Statistics and rules Syntactic rules Basic phrases
CLC: TP391.1
Type: Master's thesis
Year: 2011
Downloads: 106
Quote: 0
Read: Download Dissertation
Abstract
|
With the rapid development of information technology, the popularity of the Internet, wireless transmission of high-speed mobile broadband mobile communication brought cloud computing \to a great wealth of information era. However, in the face of a broad array of electronic information, people but seems at a loss how to get the information that people need the current problems to be solved. Information filtering came into being, the feature item weight calculation method is the basis of information filtering, but also a prerequisite for protection of information filtering effect. This paper focuses on the selection of the feature item granularity feature selection method, feature weight calculation methods and characteristics between the organization and other related content, there is a certain theoretical value and practical significance. The research topics include:. Characteristic feature selection and weight calculation concept defined summarizes the existing feature selection function and the weights of the evaluation function calculation method, and to distinguish and contrast the two, given the definition on the concept and meaning, as successor The researchers provide a theoretical basis. Improved information gain algorithm analyzes the advantages and disadvantages of the traditional information gain algorithm, and difficult to deal with the imbalance corpus this defect for traditional information gain algorithm, information gain algorithm improves the distribution of feature items and class dispersion and class dispersion as judgments conditions, a combination of improved information gain algorithm with information gain algorithm, so that does not overcome the traditional information gain algorithm is difficult to deal with the imbalance corpus of defects, and retain the advantages of the traditional information gain algorithm. Finally, the feasibility of this approach is proved by experiments. In the VSM introduced the idea of ??syntactic rules, increase the semantic description of the feature items, lexical systems are not perfect is the main reason for the Semantic missing feature items, and also the major bottleneck restricting the development of the Chinese information processing. This article through text indexing and vector space model feature item granularity selection of serious study. For the current lack of lexical system, the syntactic rules are introduced to the VSM, the syntactic rules structure the lexical merger rules, which identify basic phrases in the text, and these basic phrases instead of words in the VSM to expand the semantic description of the feature item . And experiments to prove the effectiveness of this approach. Constructed feature relationship tree, enhanced feature between the VSM VSM between each feature items of independent mutual interference among the feature item, that produces with ambiguous and classification noise, therefore, this paper The relationship between the VSM feature items reorganize combing introduced tree model, VSM feature relationship tree is constructed, and the introduction of this idea to the adverse sentiment classification and achieved good results. Construct a statistical and rules of the right weight calculation method for the currently entitled weight calculation method only use frequency times measurable characteristic items of the document contribution degree, easily cause characterized curve flat difficult embodied entries distinguish degrees of phenomenon, this article consolidated considering the feature items of the distribution, location, frequency times and syntax role information, Attribute Weighting evaluation function is constructed. And by document topic sentences extracted text classification experiments comprehensive evaluation of the algorithm. Experimental results show that the re-calculation method based on the statistics of joint rights not only increase the distinction of degree, and to overcome the flat curve of the phenomenon of feature items, and the algorithm is calculated based on the frequency of strong stability and earlier convergence . 6. Designed and implemented a network information filtering system based on statistics and rules in line with the advanced, reliable, easy-to-use design goals, from the user perspective, designed and implemented a network information filtering system based on statistical rules, the The system can automatically through the native classification filter automatically set according to specific user, the Web page's URL filtering, keyword filtering, and content filtering, which content filtering based primarily on this article improved feature selection algorithm for feature choose to reduce the dimension of the vector space, and the use of the text proposed in this paper based on statistical and rule indicates that the model represents a document vector space, greatly improving the accuracy of the information filtering.
|
Related Dissertations
- A Study on Feature Selection for Cancer Detection Based on Biological Expression Data,R730.4
- Rotating Machinery Failure Diagnosis Based on Support Vector Machine,TH17
- Fuzzy Clustering and Its Cause Analysis of Traffic Accident black spots Application,TP311.13
- Network Information Security Audit Systerm Based on Content,TP393.08
- Decision tree classifier in the \,TP311.13
- Intelligent Intrusion Detection Model Based on Information Gain Bayesian Network,TP393.08
- The Algorithm of Sensor Management Based on Information Theory,TP212
- Airborne Multi- Sensor Management Algorithm Study and Research,V243
- Joint ventures Union Influence of Chinese Enterprise Performance,F273.7;F224
- The use of decision trees to improve the ability of firewall filtering,TP393.08
- Design and Implementation of the campus network based on the network information filtering system,TP393.09
- Research on Network Information Filtering Model Based on Genetic Taboo Algorithm,TP393.09
- Class-based core word text classification technology research,TP391.1
- Decision Tree Algorithm based supermarket customer segmentation applied research,TP311.13
- Support vector machines for feature reduction studies,TP181
- A Weighted Decision Tree-Based Fast Intrusion Detection Model,TP393.08
- Study of Caihu-Guizhi-Ganjiang Decoction of Literature and Applications,R289.5
- Research on Correlation Analysis of the Staff Quality and Performance,F224
- Chinese Text Classification Algorithm,TP391.1
- Research & Practice of a Method on Customer Credit Evaluation in Mobile Communication,F626
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net Mobile
|