Dissertation > Excellent graduate degree dissertation topics show

Research and Implementation of Automatic Text Classification Based on Hadoop

Author: ZhangYongYong
Tutor: ShaXueJun
School: Harbin Institute of Technology
Course: Electronics and Communication Engineering
Keywords: text categorization Hadoop MapReduce HDFS feature wordselection
CLC: TP391.1
Type: Master's thesis
Year: 2013
Downloads: 27
Quote: 0
Read: Download Dissertation

Abstract


With the popularization of Internet and the rapid development of Internettechnology, the data on the internet increase exponentially. In these mass data, themajority is semi-structured and structured data, which is similar to Web text.Therefore, it is an urgent problem how to find and locate the effective informationof the user needs on the internet. In order to improve the efficiency and accuracy ofuser search, it is very important to classify the semi-structured and structured webtext data effectively. The classification technique have been developed to a certainstage, but in face of the massive data, due to the speed of development of computerhardware will never catch up the speed of the customers’ demand, the current speedof computer hardware can’t satisfy the demand of the massive data processing andquick response. The purpose of this paper is that a automatic text categorizationsystem with the massive data is designed and implemented.Based on the exponential growth of the massive data and the urgent needs ofclassification, the rise of cloud computing and the development of classificationtechnology, this paper put forward a design and implementation of automatic textcategorization based on a Hadoop cluster. First of all, the framework of Hadoopsystem, including MapReduce and Hadoop distributed file system(HDFS), isanalyzed. And the Hadoop cluster system is builded. Then, the current matureclassification technology and algorithm, the pretreatment and vectorization of textare studied. And the text classification model is made. Finally, the pretreatment oftext, including tokenization, stemming and removal of stop words at forth, is dealtwith linux shell. And the process of feature word selection, text vectorization,training and testing is dealt according to MapReduce programming model. Thisautomatic text categorization system reduces the hardware requirements, saves thecost, can satisfy the mass data high concurrent processing requirements andimproves the speed and efficiency of the data processing.

Related Dissertations

  1. Research and Application of Map/Reduce Based Distributed Log Analyzer,TP311.52
  2. Design and Implementation of Online Shopping Prototype System Based on Hadoop,TP311.52
  3. Text Categorization Based on Rough Set Theory,TP18
  4. Design of the Mobile Learning System Based on Hadoop,G434
  5. HADOOP architecture based on the social security project web log analysis system,TP311.52
  6. 3D Mannequins Generating Engine Based on eMTM with MapReduce,TP391.41
  7. The Research of Software Service Platform Based on Cloud Computing,TP311.52
  8. Research on Fast Queryalgorithm of Massive Data,TP311.13
  9. An Intrusion Detection System for High-Speed Networks,TP393.08
  10. Incremental Learning Method Based on Cloud Computing,TP311.13
  11. Cloudqueue: An Internet-Scale Messaging Infrastructure Based on Hadoop,TP311.52
  12. The Research of Text Classification Based on Hadoop,TP391.1
  13. Research and Implementation on a Distributed Service Registry Based on HADOOP Platform,TP393.09
  14. Resarch of Task-level Data Processing Based on Multicore CPU and Test of Its Performance on Cluster Platform,TP274
  15. Hadoop-based video transcoding system design and implementation,TN919.81
  16. Hadoop data center deployment and tracking systems research,TP308
  17. Fault Tolerance for MapReduce in the Cloud Environment,TP302.8
  18. Cloud-based mobile data storage backup system,TP309.3
  19. A scalable prototype design and implementation of MapReduce,TP311.52
  20. The Research of Distributed Text-based Data Filtering Technology and System Implementation Based on MapReduce,TP391.1
  21. Design and Implementation of Massive Web Log Analysis System Based on Hadoop/Hive,TP393.09

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net  Mobile