Dissertation > Excellent graduate degree dissertation topics show
Research and Construction of Protein Named Entity Recognition System
Author: WangWenXun
Tutor: ZhuXiaoYan
School: Tsinghua University
Course: Computer Science and Technology
Keywords: Information Extraction Protein Names Boundary Fixation Word Shape Features Integrated Strategies
CLC: TP391.4
Type: Master's thesis
Year: 2005
Downloads: 74
Quote: 1
Read: Download Dissertation
Abstract
|
In order to understand biological process, we must clarify howbiomedical substances, namely proteins, interact with each other.Unfortunately, an enormous number of information related with biomedicalare buried in millions of scientific literatures, and accumulating. As achallenging task in bioinformatics, rediscovering them, an automatic efficientprocessing with intelligent information extraction is needed.Since last year (2004), our laboratory starts bioinformatics researchproject on protein-protein interaction (SPIES). SPIES, a mature system forprotein-protein interactions, has only captured interaction between proteinsthrough automatic-generated patterns. Hence, a support system from proteinnamed entity recognition is urgently needed.We present Named Entity recognition system for Protein names(Ne4Pro), an automatic system that extract protein names from the biologicalliterature and link them to the associated entries in sequence database.The extraction system in Ne4Pro is divided into several tasks. The firsttask is the named entity identification, which is to identify name andnon-name part from biomedical literature. The second task is the named entityboundary fixation, which we provide attractive expansion and shrinkagemethod to capture the right boundaries of named entities. Then, we takeidentified named entities into the last task semantic classification, where nameentities are classified into protein and non-protein name class. Each of thetasks is carefully created based on expert knowledge on the nomenclature ofprotein names as Ne4Pro, a novel system that integrates methods fromdictionary-based, rule-based and machine learning-based.The main contributions of this thesis are: first of all, the last task of oursystem, a novel semantic classification, which integrates knowledge-based,dictionary-based, and machine learning-based method, allow us to achievehigh performance improvement over an independent baseline semanticclassification. This task shows that knowledge-based model and curation ofdictionary is important in order to reduce ambiguity of classifier. Secondcontribution is a novel boundary fixation method task, considered to be moreaccurate than just detecting longest named entities. Third contribution is anew word shape feature, which proposed to overcome the limitation ofdictionary through imitating their word shapes. In addition, a unigram/bi-gramthat is of statistical information-rich with rule based class smoothing methodis also introduced. Last but not least is the design of each task which plays animportant role of the system performance.We use GENIA corpus 3.02 to conduct 10-fold-cross validationexperiments. To achieve desirable performance in our system, we used SVMas a machine learning approach which has shown the best performance invarious Biomedical Natural Language Processing task. Unlike most previousdeveloped systems, we are not using longest named entity annotations as ourperformance evaluation due to we want to find more precise named entitiesthan just to find longest named entities. Our experiment shows that proposedboundary fixation task capable to improve the performance of system by 6.7%in precision and 9.3% in recall, while our semantic classification taskperforms better performance improvement over baseline semanticclassification, and other similar systems.
|
Related Dissertations
- Research on Domain Entity Attribute and Event Extraction Technology,TP391.1
- Active faults based radar image information extraction method applied research and demonstration,P542.3
- Based on high-resolution remote sensing data mining houses information extraction,TP751
- Template independent web information extraction,TP393.092
- The Research for Named Entity Recognition and Relation Extraction in Text,TP391.1
- Research on Temporal Information Recognition and Normalization,TP391.1
- The Research and Application of the Information Extraction Algorithm in FTIR Multi-spectral Microscopic Images,TP18
- Study on Extraction of Broad-levaed Forest Information Based on Medium and High Spatial Resolution Remote Sensing Image,S771.8
- Study of the Scale Issues in the Information Extraction of Forest Vegetation Based on Remote Sensing,S771.8
- Dynamic Origin-Destination Matrix Estimation Based on Information Extraction,U491.111
- The laser detection automotive active anti-collision technology and systems research,TN247
- City names addresses coding Matching,P208
- Segmentation of Social Media Web Page and Extraction of Topic Frequent Cluster,TP391.1
- Design and Implementation of EDPF NT-plus Operator Station Simulation Software,TP391.9
- Research of Chinese Meta Search Engine Based on Clustering,TP391.3
- The Research of Text-based Image Retrieval Technology in Uyghur Kazak Kirgiz Search Engine,TP391.3
- Mobile Web Site Development Based on DOM Tree Information Extraction,TP393.092
- The Relation and Analysis of Relation-network Analysis Technology Based on MongoDB,TP311.13
- The Images Interpretation and Information Extraction of the Burst and Potential Debris Flows in Remote Sensing Way,P642.23
- Study on Remote Sensing Geological Mineral Information Exrraction for Tibet Xiongcun Porphyry Coper-gold Deposit,P627
- Algorithms of Snow Cover Extraction from Multi-source Remote Sensing Data Based on Adaboost and Its Application,P426.63
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Pattern Recognition and devices
© 2012 www.DissertationTopic.Net Mobile
|