Dissertation > Excellent graduate degree dissertation topics show
Automatic extraction of domain concepts
Author: YaoXianMing
Tutor: GuoJianYi
School: Kunming University of Science and Technology
Course: Computer Software and Theory
Keywords: Domain term extraction Word co-occurrence Get connotation domain concepts Domain concepts epitaxy ( instance ) learning
CLC: TP391.1
Type: Master's thesis
Year: 2010
Downloads: 112
Quote: 0
Read: Download Dissertation
Abstract
|
Ontology in Semantic Web, information extraction, knowledge management and other areas more obvious role, how to automatically or semi-automatically construct ontology become an urgent task. Ontology domain concepts is the most basic of the most important modeling primitive ontology relations, instances, functions, axioms are built on the basis of the concept, its domain ontology plays an important role. How efficiently from large text fields automatically get the concept of researchers presented a huge challenge. In this paper, the problem domain concepts extracted from the field terminology extraction, meaning access to domain concepts, domain concepts epitaxy (instance) carried out three aspects of learning related research and experimental work. 1, domain terminology extraction. Firstly, using the rules to get the text from the field of terminology may be a string, and then use the mutual information within the string tightly investigated the extent to which the field in the mix to get a meaningful high-frequency words, and finally use the word co-occurrence analysis of these frequency words in the correlation between the terms for the field of terminology. The innovation is the introduction of word co-occurrence analysis of the correlation between the field of terminology, examples of areas such as high-frequency words will be filtered out, enabling the automatic extraction of domain terms. 2, the domain concepts connotation get. In this paper Text (HOWNET) This semantic dictionary, with the meaning of the original meaning of the concept to describe the concept of meaning in order to achieve automatic acquisition; For some unknown word problem, we adopted a strategy to achieve its connotation split automatically obtain . Meanwhile, for the term repetitive questions, we use K-means clustering technology to achieve the term de-emphasis, compute the similarity between the terms of its meaning by reference to the distance between the original. 3, the domain concepts epitaxy (instances) to learn. For the current instance learning rule-based and statistical methods shortcomings, this article will introduce machine learning methods to the instance learning, and try out based on support vector machine (SVM) instance learning method, the experimental results demonstrate the effectiveness of the method . In this paper, the concept of extraction when the field research status and existing problems, put forward the concept of extraction of the basic ideas and terminology extraction according to the field, the field of conceptual content acquisition, domain concepts epitaxy (instance) the concept of learning steps taken to carry out the field related experimental results verify the effectiveness of these methods.
|
Related Dissertations
- Research on Terms Co-occurrence Based Models and Algorithms for Text Mining,TP391.1
- Research on the Coexistence Relation of Words and Its Application in Text Classification,TP391.1
- Automatic Identification of Chinese Organization Name,TP391.12
- Research and Implementation of Method for Web Noise Elimination and Feature Selection,TP393.092
- Research of Chinese Short-Text Classification,TP391.1
- Scientific Paper Discrimination Method Research Based-on Word Co-Occurrence Network and Support Vector Machine,G322
- Scientific Paper Discrimination Method Research Based on Laplacian Spectrum Analysis,G311
- Research on Keywords Acquisition Based on Semantic Distance from Web Pages,TP391.1
- The Description of Text’s Feature Based on Semanteme Concept,TP391.1
- Research on Domain-Specific Term Extraction Based on Semi-Supervised Learning,TP181
- The Research on Keywords Extraction from Chinese News Web Pages Based on Clustering,TP391.1
- Keyword Extraction from News Web Pages,TP391.1
- A Text Categorization Method Based on Features Clustering,TP391.1
- Tthe Research on Modeling User Interest Based Keywords Dependency Profile,O242.1
- Research on the Alignment of Video Features and the Descriptive Vocabulary,TP391.41
- BBS Spam Filtering Model Based on Word Co-ocurrence,TP393.094
- Research on Basic Algorithms of Digital Image Processing and Implementation with FPGA,TP391.41
- Research of Visualization Technology in the Virtual Test of Missile,TP391.9
- Tongue Feature Extraction and Research of Fusion Classification,TP391.41
- Research on Temporal Information Recognition and Normalization,TP391.1
- FPGA/DSP Image Co-processor Technology and Ethernet Data Transmission,TP391.41
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net Mobile
|