|
With the advent of information society in the late 20th century, there has been a flood of data streams, time series, graphics, social networks, spatio-temporal data, text data, Web data, found useful knowledge from these massive data is a hot issue of data mining. Which repeated measurements by different time value or events constitute a time-series data, in large numbers in the financial, engineering experiments, meteorological, medical, transportation and other fields. Concern in the field of data mining, time series are mainly concentrated in the trend analysis, similarity search, association rules and other issues. In this paper, stock data, for example, in the Summary of domestic and international data mining research and development profile, the expression of the time series, rule discovery, the similarity search research and analysis, segmented symbolic time series similarity search algorithm and time series data association rule discovery has made some achievements, the main research contents and results are as follows: 1) To introduce the basic theory of data mining analysis, including data mining analysis process, the traditional time series analysis and data mining analysis, frequent pattern discovery, the strong association rules, and they were in-depth, systematic study and analysis. 2) The time series of piecewise linear, symbolic representation of commonly used piecewise linear representation method based on time series study compared the use of special points extracted based on the change in slope of the segmentation algorithm to compress the data, good to solve the time-series data of high dimension, a large amount of data; proposed to consider the time factor, segment relative slopes of the eight-mode-based symbolic representation, valid representation of the relationship between the ups and downs of the stock price and time. And stock data validation algorithm to achieve better results. 3) similarity search seriously study the time sequence similarity search method, paragraph breaks and minimum first difference cycle chain code-based fast similarity search algorithm, compared with the previous search algorithm more efficient and effective The solution to the Timeline translation, scaling, rotation similar judgment. Stock data validation, the industry analysis of the stock time series data of great help. 4) the association rule discovery research sequence in the symbolic representation of the correlation algorithms, explore, Apriori and FP-tree combined with the new algorithm, found a single vessel and multivessel frequent pattern of stock data, and generate association rules .
|