Dissertation > Excellent graduate degree dissertation topics show

An Improvement of Cluster on Phylogenetic Profiling Method

Author: LiDongZuo
Tutor: MaZhiQiang
School: Northeast Normal University
Course:
Keywords: Phylogenetic profile Weights Hierarchical clustering K-means clustering Biology Distance K-means initial sample
CLC: TP311.13
Type: Master's thesis
Year: 2011
Downloads: 10
Quote: 0
Read: Download Dissertation

Abstract


With efficient, automated sequencing technology, bioinformatics central topic, has moved from sequencing of genes transferred to the analysis of genes that have been sequenced, mainly on the study of gene function and notes. Since the method itself homology defects and accuracy problem, people began to pay attention nonhomologous method. Nonhomologous sequence method is mainly through classified attributes, and thus functional predictions. Phylogenetic spectrometry method in numerous nonhomologous most widely used applications. Phylogenetic spectrum method proposed in 1999 by Pellegrini, followed by many scholars reference group selected from the gene, constructed phylogenetic spectrum, spectral similarity analysis of these three aspects of its improvement. This article on this solid foundation, first constructed based on the weight of the phylogenetic profile, then alternate use hierarchical clustering method and K-means clustering method for similarity analysis. In the spectral similarity analysis phase, we propose two improvements: First, propose a new distance for hierarchical clustering clustering phase. The second is from hierarchical clustering method to extract more information for K-means clustering method to provide initial information, the full utilization of the results of hierarchical clustering, K-means clustering method makes the results more accurate. Currently in the clustering algorithm, the main application is the Euclidean distance. Because we are dealing with a sample of mostly belong to Euclidean space, so using Euclidean distance clustering can get good results. The distance used herein, is a non-Euclidean space distance. Compared to the Euclidean distance, it reinforces the known information on the sample from the impact. It is not only the distance between the sample considered, but also consider the sample and the reference sample distance. Using this new distance, allows us to give priority to a similar sample with a known reference. K-means clustering method flaw is that the sensitivity of selected initial conditions: the initial number of clusters K and the initial cluster target selection, will finally have a significant impact on clustering results. Currently the K-means algorithm mainly in the selection of the initial information. Predecessors using hierarchical clustering combined with K-means clustering method used, the purpose is to use hierarchical method of K-means clustering method provides the initial number of clusters K. In this paper, based on the results from the hierarchical clustering method to extract more useful information, K-means clustering method is given an initial clustering objectives. Finally, with Escherichia coli K12 genome as a test sample, test verification for these improvements. The experiment showed that, compared with the previous result, the new algorithm with higher accuracy.

Related Dissertations

  1. Research on Clustering Algorithm Based on Mutation Particle Swarm Optimization,TP18
  2. Research on Fuzzy C-Mean Clustering Algorithm Based on Particle Swarm Optimization and Shuffled Frog Leaping Algorithm,TP18
  3. Research on Clustering Algorithm Based on Genetic Algorithm and Rough Set Theory,TP18
  4. Texture Synthesis Algorithm Based on Samples,TP391.41
  5. The Design and Implementation of Teaching Quality Assessment Based on AHP for Enterprise Practice,F272
  6. Study of Routing and Wavelength Assignment Problem in WDM All Optical Networks,TN929.1
  7. Based on the User’s Perceptual Needs’ Handle GPS Design,TN967.1
  8. Loading Optimization Method Based on Two Level Clustering,U492.22
  9. Extraction Algorithm based on remote sensing images of the road,TP751
  10. The Improvement on the Fuzzy C-means Algorithm,TP311.13
  11. A New Algorithm of Image Segmentation Based on Pixel-level Multi-feature,TP391.41
  12. Arouting Protocol of Wireless Mesh Network Based on Load Aware,TN929.5
  13. The Maximum Priority Index Method with Flexible Content Balance Constraints in Computerized Adaptive Testing,O224
  14. Optimization and empirical analysis of the input-output model of the enterprise information,F270.7;F275
  15. Gene Chip Image Processing Base on Shape Adaptive DCT,TP391.41
  16. Research on the Crack Defects Detection Algorithm of Aircraft Skin,TP274
  17. Study on the Selection of Financial Centers Based on the Methods of Spatial Statistical Analysis and Combined Evaluation,F832.7
  18. Technology Research on Multiple Radar System Multi-Target Tracking,TN957.52
  19. Improvement Research and Implementation of License Plate Recognition,TP391.41
  20. Research and Application of Segmentation Algorithm of Billets in Complex Production Line Scene,TP391.41
  21. Researching of Evaluation in Mini Vehicle Body Shop’s Reforming Program Based on TOPSIS Method,U468.4

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer software > Program design,software engineering > Programming > Database theory and systems
© 2012 www.DissertationTopic.Net  Mobile