|
Implicit in data mining , incomplete , noisy , fuzzy , random data extracted , not known in advance , but potentially useful information and knowledge . With the rapid development of computer network technology , database technology , database management system widely used to produce large-scale data , and the explosive growth . However, people can not understand and effective use of these data implied valuable knowledge , resulting in an awkward position in the \The use of a variety of data mining tools that people have a powerful means to get rid of this embarrassing situation , because the data mining system tools can help people to intelligently and automatically find the potential knowledge left in the huge amounts of data , for analysis decision - making purposes . The research is based on a database and support for data mining standard process , and integrate a variety of data mining algorithms , the platform of the second generation data mining system applicable in many fields . Reads as follows: ( 1) data mining theory , data mining system tools as the starting point , step-by-step analysis of the study data mining platform should have the functionality and features of its design and lay a theoretical foundation . ( 2) data preprocessing is an integral part of the data mining system platform , and carried out a detailed analysis of the need for data pre-processing , content , the main method . ( 3) classification algorithm for data mining , clustering algorithm , correlation analysis, predictive analysis elaborated , part of the data mining classic algorithms : ID3, C4.5, Naive Bayes classifier , K-means and K- modes the shortest distance, the longest distance method of DBSCAN with distance - based outlier analysis , Apriori, linear regression analysis . In the paper , a detailed description of the ID3, the basic idea of the Naive Bayes classifier , K-means , Apriori , linear regression analysis has been improved and other related content , and detailed instructions to the instance , such as the implementation process . (4) the design of common data mining system platform , will be divided into data processing , data mining, and visualization of three modules described in detail .
|