|
The rapid development of Internet Web has become the main way for people to obtain information , but with the surge in the number of pages full of irrelevant information and make people want to quickly and accurately obtain the information they need to become very difficult. Meanwhile, as \Internet , these pages are search engine \Addition , Internet information on each site independently of each other , collected very difficult , in this case , the usual search engines play a negligible role , and the Web information extraction technology is very necessary. Through the existing information extraction technology summary and analysis , combined with \. And on this basis , to achieve an extraction rule strong applicability , high degree of automation of Web information extraction system . System implementation is divided into sample learning, information extraction and database storage in three stages to complete. Learning phase is the system in which the key points and difficult to achieve , at this stage , the path through the leaf node access algorithms, data acquisition algorithm region , semantic acquisition algorithm and path optimization algorithm design and implementation , while using powerful and flexible XSLT, generate robust and adaptable extraction rules file. In the information extraction phase , the system uses the URL pattern matching and DOM similarity algorithm, Web pages and extraction rules to be extracted files automatically match . Meanwhile, in order to solve the balance of automation and accuracy problems , the system provides simple GUI interface , support for manual guidance training , tested the system for the \
|