Dissertation > Excellent graduate degree dissertation topics show

Study on Data Quality Assessment Techniques for Telecom Data Mining

Author: WangXiaoHua
Tutor: ZuoJian;SuHongYe
School: Zhejiang University
Course: Control Theory and Control Engineering
Keywords: Telecom Data Mining Insolvency Data Quality Assessment Missing Value Imbalanced Data Outlier Degree
CLC: TP311.13
Type: Master's thesis
Year: 2010
Downloads: 291
Quote: 4
Read: Download Dissertation

Abstract


In recent years, as an effective method of knowledge extraction, data mining has been widely used in the field of telecommunication, for example, telephone fee fraud detection, customer subdivision, customer churn prediction, cross selling, etc. However, telecom data is usually of poor quality, which can not meet the requirements of data mining, and that is why there are so few successful cases of telecom data mining. Poor data quality has already been the bottleneck of data mining’s application to the filed of telecommunication, so it is necessary to do the data quality assessment to measure the availability of data mining, and thus to avoid the unnecessary waste of time and energy. For data quality assessment, there are many research achievements on that subject, but most of them focus on the frame theory, which are seldom related to specific business backgrounds and applications. Till now, there is still no specialized literature discussing about specified mining subject oriented data quality assessment. Aiming at insolvency mining, which is one of the most commonly used subjects of telecom data mining, based on deep research of how missing values and outliers affect the classification result, this thesis conducts a research on data mining oriented data quality assessing techniques. The main research work is listed as follows.1. For Missing Evaluation, the concept of Class Distribution is proposed to measure the relationship between an input attribute and the target attribute, and based on that, a Class-distribution-based Attribute Weighting Algorithm (CAWA), which can discriminate the importance of different input attributes, is presented. Based on CAWA, an Attribute-weight-based Missing Evaluation Algorithm (AMEA) is presented to realize Missing Evaluation. The experiment results show that this missing evaluation algorithm can reasonably measure the affects to mining results caused by missing values.2. For Outlier Evaluation, aiming at the characteristics of the telecom data, especially the imbalanceness of insolvency data, the affects to classification results caused by outliers in imbalanced datasets are analyzed, and the concept of Outlier Degree (OD) is proposed based on the combination of Hyper-graph Outlier Test (HOT) Algorithm. Based on that, an Imbalanced Outlier Evaluation Algorithm (IOEA) is presented to realize Outlier Evaluation. The experiment results show that this outlier evaluation algorithm can reasonably measure the affects to mining results caused by outliers.3. Based on Missing Evaluation and Outlier Evaluation, combined with telecom insolvency data mining’s own characteristics, a comparably complete data quality assessing system is presented. This system is composed of Missing Evaluation Sub-system and Outlier Evaluation Sub-system. Based on the experiments, combined with the experience of telecom experts, a reference value of assessing point vector is given. The results of the experiment show that this reference value can provide meaningful guidance to the mining feasibility analysis.

Related Dissertations

  1. A Study on Healthcare Product Marketing Based on Data Mining Technology,F426.72
  2. Gao Zhong-ying academic thought and experience and use of Bufei Decoction treatment of common diseases of the respiratory system drug law,R249.2
  3. Bing- thick academic thought and clinical experience and empirical studies apply to turtle soups treatment of chronic kidney disease,R249.2
  4. The Design and Implementation of Bicluster Data Analyzing Software,TP311.52
  5. Research on Clustering Algorithm Based on Mutation Particle Swarm Optimization,TP18
  6. Research on Fuzzy C-Mean Clustering Algorithm Based on Particle Swarm Optimization and Shuffled Frog Leaping Algorithm,TP18
  7. Research on Clustering Algorithm Based on Genetic Algorithm and Rough Set Theory,TP18
  8. Based on data mining research tax audit case selection,F812.42
  9. Community-oriented education, personalized learning system and its implementation,TP391.6
  10. Association rule mining based Intrusion Detection System Research and Implementation,TP393.08
  11. Data warehouse technology in the banking customer management systems research and implementation,TP315
  12. Investments Projects Evaluation Supporting System Based on Optimizing Model of Industrial Parameters,F283
  13. The Research and Application of Data Mart in the Telecommunication Business Analysis,TP311.13
  14. Traffic Classification Using Stream Data Mining Algorithm,TP393.06
  15. Design to E-learning System in Senior Vocational School Base on Moodle,TP311.52
  16. The Application of Association Rules Algorithm in Higher Vocational Colleges’ Endorsement of Impoverished Students,G717
  17. Research on Application of Data Mining Technology in Degree of Satisfaction Analysis of Television Customers,TP311.13
  18. Web Usage Mining and the Research of Personalized Recommendation,TP311.13
  19. Data Mining of Application in the School Management and Training Students,TP311.13
  20. Research on Employment Monitoring System of University Graduate,G647.38
  21. The New Privacy Preserving Algorithms for Linear Programming and SVMS,O221.1

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer software > Program design,software engineering > Programming > Database theory and systems
© 2012 www.DissertationTopic.Net  Mobile