Dissertation > Excellent graduate degree dissertation topics show

A Method Based on the Statistics of Punctuation for Extracting Text from Chinese Type Web Pages

Author: WangZuo
Tutor: XiaoChun
School: Xiangtan University
Course: Computer Science and Technology
Keywords: Information Extraction DOM tree Statistical characteristics k-means Clustering
CLC: TP391.1
Type: Master's thesis
Year: 2012
Downloads: 2
Quote: 0
Read: Download Dissertation

Abstract


As the rapid development of Internet, The web has become a huge source ofinformation, Information Extraction as a necessary means of the user to obtainaccurate information. How quickly and efficiently obtain the information people ingenuine need to become an urgent need to address the important issues from the webdocuments. So the web document information extraction technology is developed.Traditional information extraction techniques to extract design information asplain text. Usually the main use of natural language understanding technology.However, due to the plain text does not have any information that can be used inaddition to text outside. So the development of information extraction technology isrelatively slow. So how to improve the efficiency of web information extraction hasbeen a hot research; For Chinese type of web text information extraction hasimportant practical significance has also become the focus of information extraction.This article in view of the above problems, through in-depth systematic study of webdocuments, information extraction, to obtain the following two researchachievements.1. Distribution characteristics of the punctuation marks in a web documentobtained by experimental observation, found that punctuation can be used as the basisfor web documents, information extraction.2. For the Chinese type of web page, Text information extraction method isproposed based on the punctuation statistical characteristics, This method firstconverted into web documents with the wish to corresponding DOM tree, And thenby the punctuation of the demographic characteristics of constructed leaf node featurevector, weights to calculate the distance between each leaf node in the leaf nodes bythe eigenvectors, select the weight of the maximum and minimum of two leaf nodesas a sample leaf node. Treat them as the initial conditions of k-means algorithm, useof extraction k-means all the leaf nodes into two clusters, finally, select the right valueof the largest leaf node where the cluster, each leaf in the cluster contains all the textinformation as a web document. Experiments show that, the method can accuratelyextract the body of the page information, also has high versatility.

Related Dissertations

  1. Research on Domain Entity Attribute and Event Extraction Technology,TP391.1
  2. Research and Implementation of Mining Implicit User Interest,TP311.13
  3. Research on Temporal Information Recognition and Normalization,TP391.1
  4. Establishment and Update of Similar Users’ Cluster in Personalized Information Retrieval,TP391.3
  5. Research on Removal Algorithm of Shadows in Image Segmentation,TP391.41
  6. The Research of the Text Extraction Method Based on Spectral Cut,TP391.41
  7. Gao Zhong-ying academic thought and experience and use of Bufei Decoction treatment of common diseases of the respiratory system drug law,R249.2
  8. Research and Improvement on K-Means Clustering Algorithm,TP311.13
  9. Research on Peer-to-Peer Traffic Identification Algorithm Based on Cluster Analysis,TP393.02
  10. Research of Scheduling Algorithm Based on Hybrid Adaptive Genetic Algorithm in Computing Grid,TP393.09
  11. Evaluation of Photosynthetic Efficiancy of Seedlings of the Hybrid Progenies (F1) in Peach,S662.1
  12. The Research of "Ant Group" Phenomenon in the Harmonious Society,D669.5
  13. The Load Research and Comprehensive Evaluation on the Agricultural Non-Point Source Pollution in Nantong,X592
  14. Study on Growth Monitoring Technique Based on Pixel Un-Mixing Method and HJ Remote Sensing Images in Paddy Rice,S511
  15. BF-FCM Clustering Algorithm and Its Application in the Image Segmentation,TP391.41
  16. The Application of Ant Colony Algorithm in Meteorological Satellite Cloud Pictures Segmentation,TP391.41
  17. Research on Clustering Algorithm Based on Mutation Particle Swarm Optimization,TP18
  18. Research on K-means Optimization Clustering Algorithm,TP311.13
  19. Research on Fuzzy C-Mean Clustering Algorithm Based on Particle Swarm Optimization and Shuffled Frog Leaping Algorithm,TP18
  20. Research on Clustering Algorithm Based on Genetic Algorithm and Rough Set Theory,TP18
  21. The Study about the Select Strategies of Sportswear Brand Communication Means,G206

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net  Mobile