|
With the rapid development of Internet, the Web has become a huge, shared, distributed information resources collection, how to automatically obtain from the vastness of the Web information resources of interest to the user's data records or useful information, has become depth issue. Currently, most Web data are presented to non-structural or semi-structured form, the lack of a description of the data itself, does not contain a clear semantic model is also not very clear, which makes the application can not directly resolve and take advantage of the mass on the Web Information. In order to avoid the embarrassment of \commodity prices, tracking partners or competitors dynamic, integrated enterprise external information, etc.), a variety of Web data extraction technology came into being, and increasingly play a powerful role. Web data extraction technology has obvious advantages and broad prospects, data extraction technology, artificial intelligence, information retrieval, natural language understanding technology in network information processing applications, the research focus of the field today. Thesis page for a list generated by the template, how to detect its generic template, and extracted data embedded in the template list page automatically extract data. First, a brief description of the semi-structured data, Web data extraction, the concept of the list page; followed by in-depth study of the history of the development of Web data extraction technology, art, and pointed out that the existing technology in the comparison of various technologies advantages and disadvantages of the possible future direction of development; then, details of the previous stage of academic achievements - to automatically extract Web data tree alignment algorithm, it is proposed in this paper based on the tag tree list page data extraction system research foundation and core part. Was implemented this algorithm to improve the alignment of the tree before and after the relevant processes, formed a complete system of Web data extraction; final details that the proposed list of tag-based tree page data extraction system design, achieved with experiment, the system configured tag tree, excavation of the main data area, the identification data is recorded, and generates a recording mode, and gradually narrow the size of the target area, in order to extract the user's interest data record or useful information. The experiments show that the system is able to deal effectively with the list page, extract data information, be able to adapt to a wide range of practical needs, and has the depth to promote the application of value.
|