Dissertation > Excellent graduate degree dissertation topics show

Research of Disambiguation of Internet People Information Technology

Author: MaErLei
Tutor: LiuYuanChao
School: Harbin Institute of Technology
Course: Computer Science and Technology
Keywords: Disambiguation Social Network Area classification Social attributes Features Library
CLC: TP393.4
Type: Master's thesis
Year: 2010
Downloads: 28
Quote: 0
Read: Download Dissertation

Abstract


With the development of Internet and its relative technology, the WWW has become the largest information area. For the enterprise or the individual, web becomes the main information source gradually. However, because of too many web sites and the information overflow resulting from this, it is more and more difficult to obtain useful information. When searching for person information, you will gain huge information and too much duplication, and the accuracy is not high. So, the person information extraction system is built to allow users faster and more convenient to obtain the required information, and the result simple, refined and beautiful.Because different people may be in different areas, so in this paper, according to this feather, the document information can be divided into seven categories, such as cultural, administrative, military, science, education, sports, health, economic. According to this method, we can avoid the same information processing between people of different areas and can improve the efficiency of the system. In addition, the pre-classification of this method can achieve high recall rate and we can ensure that the information of people in different areas will not crossover, so as to reduce the error rate that in the subsequent processing, the information of people in different areas my be classified into one group.In this paper, we implement the disambiguation processing by combining social networks and context information. If we only use the social networks or the context information, our disambiguation processing can not perform very well, because there will be only one people’s name in the entire social networks or the social network is very small if we only use the social network and if we only use the context information, the context information of the document can not characterize characters very well, so we use this two methods to improve accuracy and the recall rate of the system. Using social network, we can achieve high accuracy, but the recall rate will be low, then use context information, we can overcome the disadvantage, and achieve good performance in both accuracy and recall rate.Character information processing system is a system that first run pre-classification according to the character information which collected through retrieving the character’s name and using web crawler web crawling, then cluster using social network and context information, and finally display the network information according to different character entity in the interface of system.

Related Dissertations

  1. Word Sense Disambiguation Corpus Automatic Acquisition,TP391.1
  2. Scholar Resume Automatic Generation Based on Text Mining,TP391.1
  3. Disambiguation of Biomedical Abbreviations,R-5
  4. The Frame Disambiguation of Automatic Identification of Chinese Frame,TP391.1
  5. Research on Cluster-based Person Name Disambiguation,TP391.1
  6. Study of Safety Management in Subway Construction Based on Social Network Analysis,U215.8
  7. Academic Network Repetitions disambiguation algorithm,TP301.6
  8. Research of Semantic Tag Recommendation Model and Algorithm Based on SCOT,TP391.3
  9. Based on Wikipedia disambiguation of named entities,TP391.1
  10. Figure based on association rules and emotional tendencies sort of sentence,TP391.1
  11. The Study of Text Classification and Retrieval for Chinese Patent,TP391.1
  12. Research on WordNet Based Chinese-english Cross Language Text Similarity Measurement,TP391.1
  13. A Study of Gene Name Normalization and Functional Prediction Based on Semantic Resource,TP391.1
  14. Chinese named entity recognition and disambiguation of,TP391.1
  15. The Disambiguation of Juxtaposition of Noun Phrases in NLU and Its Application in Intelligent Instrument Design,TP391.1
  16. A Study of the Influence of Linguistic Features on the Senses of Modal Verb Must by Naive Bayesian Model,H313
  17. Chinese word sense disambiguation research,TP391.1
  18. Chinese Personal Name Disambiguation in Web People Search,TP391.1
  19. Word Sense Disambiguation of English Modal Verb Will by Support Vector Machines,H313
  20. Improved Sentence Similarity Algorithm Research and Its Application in Question Answering System,TP391.1

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Computer network > Internet
© 2012 www.DissertationTopic.Net  Mobile