Dissertation > Excellent graduate degree dissertation topics show
Research on Missing Value Imputation for Microarray Gene Expression Data
Author: LiYanFang
Tutor: GuanYuDong
School: Harbin Institute of Technology
Course: Information and Communication Engineering
Keywords: missing value imputation global structure similarity structure domain knowledge
CLC: Q78
Type: Master's thesis
Year: 2011
Downloads: 60
Quote: 0
Read: Download Dissertation
Abstract
|
DNA microarray technology is a kind of biological chip technologies. It is mature and widly used. It can measure the mRNA levels of thousands of genes under certain experiments simultaneously. However, microarray gene expression data generally suffers from missing value problem due to a variety of experimental reasons. Actually, public microarray datasets have missing values in various degree, which can adversely affect downstream analysis. Simply deleting the genes with missing values can result in a loss of much usefull information, and to repeate the experiments is obviously very expensive and time consuming.Missing value imputation is researched from several aspects as follows.Firstly, Bayesian principal component analysis imputation method is studied based on the global correlation information in the data set. It consists of three elementary processes, which are principal component regression, Bayesian estimation and a repetitive algorithm. Missing values and model parameters update each other continuously and ultimately achieve convergence to estimate missing values.Secondly, according to the principle of gene co-expression, this dissertation uses local similarity structure in the data set to study K-nearest neighbors imputation method and local least squares imputation method. Both of them have a similar problem that the estimation accuracy declines in the case of high missing rate. Improved methods are proposed, which expand the range of alternative genes by pre-filling eligible genes and estimating missing values according to the missing rate of genes. Experiments show that the proposed algorithms can improve the effectiveness significantly.Additionally, the common theme for algorithms in this category is the integration of domain knowledge or external information into the imputation process. For example, histone acetylation may alter chromatin structure and provide binding surfaces for transcription factors. Histone acetylation and gene expression datasets are combined to select neighbor genes to estimate missing values.Finally, validation of imputation results is an important step in assessing the performance of imputation algorithms. This thesis focuses on internal validation containing indices derived from statistical calculation and clustering methods. What’s more, these indices are also applied to datasets containing genes differentially expressed. In a word, these work is to research the accuracy and application range of the various missing value imputation algorithms.
|
Related Dissertations
- Modeling and retrieval of multimedia personalization,TP391.3
- Human motion sequence data semantic analysis method,TP391.1
- Study on Data Annotation of Deep Web Data Integration System,TP393.09
- Research on Subject and Predicate Relation Identification and Theme Relevance Computation Technology,TP391.1
- Mathematics domain knowledge to learn the knowledge characterization,O1-4
- Researches on Information Extraction Mechanism of Physics Disciplinary Domain Knowledge Study,G633.7
- The Application of Support Vector Machine on RF Power Device Modeling,TN626
- Research on Teaching-Domain Knowledge Repository Based on Ontology,TP391.1
- Domain knowledge uncertainty reasoning,TP391.1
- A Personalized Recommendation Algorithm Based on Knowledge Category Tree,TP182
- Research Knowledge Manage and Construction Based on Ontology,TP182
- MAS-based distance education system in the field of knowledge,TP399-C1
- Research on Framework of Ontology Based Knowledge Discovery System,TP311.52
- Research on Domain-Oriented Ir System Architechture and Related Techniques,TP391.3
- Course Knowledge Based on the Test and Counseling System Research and Application,TP391.6
- Research on Web Text Extraction and Categorization from Hierarchy Technologies,TP391.1
- Research on Test Question Classification and Similar Test Questions Detection Based on Domain Knowledge,TP311.52
- Requirements Specification Meta-model and Its Generation Tool,TP311.52
- Change Detection of Harbor Ship Target Based on Structrue Information,TP75
- Domain knowledge based semi-supervised clustering algorithm,TP311.13
- Research on Ship Domain Knowledge Acquisition Based on Ontology,TP182
CLC: > Biological Sciences > Molecular Biology > Genetic engineering (genetic engineering)
© 2012 www.DissertationTopic.Net Mobile
|