Dissertation > Excellent graduate degree dissertation topics show
Research of High Dimensional Retrival Algorithms Based on Adaptive Cluster Distance Bounding
Author: GuoJuan
Tutor: HuangHong; WangBen
School: Zhejiang University of Technology
Course: Applied Computer Technology
Keywords: high-dimensional indexing structure adaptive cluster distance bounding KNNquery approximate similarity search dimension reduction
CLC: TP311.13
Type: Master's thesis
Year: 2012
Downloads: 26
Quote: 0
Read: Download Dissertation
Abstract
|
With the fast development of the high dimensional databases, the size of the high dimensional database is becoming increasing. Now, reasearchers have proposed a lot of high dimensional index structures in order to accelerate the query efficiency. But most of these structures have two defects as follows:first, most of them are based on an unreasonable assumption that data are uniformly distributed and have no connections among dimensions. Second, the query efficiency based on these structures are becoming lower with the increase of the dimension size and a research has improved that the "curse of dimensionality" occurs when the dimensionaliry is more than10. A popular and effective technique to overcome the curse of dimensionality is the vector approximation file, but the acceleration performance of it is limited in a small range, moreover, this structure is also based on the assumption that the data are uniformaly distributed, with independent attributes. All these cannot satisfy the practical use.The adptive cluster distance bounding (ACDB) method for high-dimensional indexing is based on the assumption that the data are non-uniformly distributed and have some connections among dimensions for real datasets. Comparied with other high dimensional index structures, the method can solve practical problems better. Moreover, the indexing method is using the Voronoid clustering, which allows for more general convex polygon structures that are able to efficiently bound the cluster surface. In2011, some researchers proposed the KNN algorithm based on ACDB. This algorithm has reduced the I/O times by filtering out irrelevant clusters though a lower value of the query-cluster distance. Reasearchers have proved that this algorithm based on ACDB is better at reducing the I/O numbers and the query performance than the sequencial query, such as VA-File, iDistance, LDC and so on. But with the increasement of the dimensionality and data size, the query efficiency is becoming lower and lower, so the KNN algorithm based on ACDB is still needed to be optimized. Hence, our main work and results are as follows: First, we investigated the-state-of-art algorithms about the high-dimensional indexing in the field of ACDB, and found the defects in the KNN algorithm based on ACDB.Second, with extra storage requirements and pretreatment, we proposed an IV-KNN search algorithm by reducing unnecessary distance calculations and CPU costs according to the triangle inequality formula. Experiment results based on real data sets proved that the efficiency of IV-KNN search algorithm is much better than V-KNN.Finally, to analyze the current approximate similarity retrival methods, we propose a RDFV-KNN algorithm by both applying the PCA and LLE dimension reduction technology. From experiments of real datasets, the efficiency and accuracy of RDIV-KNN algorithm has been improved straightforward.
|
Related Dissertations
- Study of Data Reduction Technique Based on Manifold Learning,TP311.13
- A Novel Approach for Projection by Optimize Local Area in Face Recognition Applications,TP391.41
- Chinese Keyword Extraction Method Based on Word Span and Its Application in Text Classification,TP391.1
- Face Image Retrieval and Recognition Based on Data Dimension Reduction,TP391.41
- The Study of Text Classification and Retrieval for Chinese Patent,TP391.1
- Feature Dimensional Reduction for Fast Face Recognition,TP391.41
- An Unsupervised Feature Selection Method Based on the Degree of Feature Cooperation,TP18
- Chinese text orientation Classification System,TP391.1
- Research on Hyperspectrum Dimension Reduction Algorithm Based on TDPCA and SPIHT Compression,TP391.41
- Research of Clustering and Characteristic Dimension-reduction Based on Immune Genetic Algorithm,TP311.13
- The Research of Locally Linear Embedding Algorithm Based on Hadoop Platform,TP311.13
- Sparse Representation Fordimensionality Reduction,TP391.41
- Text Orientation Analysis Based on Emotion Chunks Combining with Machine Learning,TP181
- Research on Similarity Search in Hybrid Peer-to-Peer Systems,TP393.02
- A number of studies of high-dimensional data discriminant analysis problem,TP391.41
- Research on Dimensionality Reduction and Clustering Algorithm of Commercial Data Streams,TP311.13
- Study on Algorithms for Fast Outlier Detection,TP311.13
- Research on the Algorithm for Text Clustering Based on Feature Words,TP301.6
- Study on Nonlinear Dynamics of the Low-Pressure Rotor-Bearing System of 200MW Turbine Set,TK263
- The Analysis of the DRT’s Computation Complexity and the Research on Compound DRT Based on DCT,O158
- Numerical Methods and Analysis for the Schr(o|¨)dinger-Poisson Equations and Their Applications,O241.8
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer software > Program design,software engineering > Programming > Database theory and systems
© 2012 www.DissertationTopic.Net Mobile
|