Dissertation > Excellent graduate degree dissertation topics show
Research and Implement of Chinese Word Segment Techniques Based on the Conditional Random Field
Author: LuZuo
Tutor: TanYunMeng
School: Huazhong University of Science and Technology
Course: Communication and Information System
Keywords: Chinese word segmentation Conditional Random Fields Machine Learning
CLC: TP391.1
Type: Master's thesis
Year: 2011
Downloads: 65
Quote: 0
Read: Download Dissertation
Abstract
|
The growing mass of information in the Internet has brought great challenges to information processing , especially in the field of Chinese information processing , the most important one is the Chinese word segmentation technology requirements . The main purpose of the Chinese word segmentation is the Chinese statement resolves to the combination with the correct meaning of the word , so that the computer can understand the meaning of the sentence , information indexing, automatic summarization , text classification , automatic clustering , text correction, handwriting input , intelligent response the word first step in Chinese information processing , and other fields , so the strengthening of the Chinese word segmentation technology has a very important significance . In this paper, the conditions with the Airport statistical model for Chinese word tokenization successfully converted to a combination of Chinese characters word formation problem is to introduce the concept of the character lexeme kanji lexeme marked so that it can in accordance with the word the bit combination rules merge words , the whole process to be achieved through the training of machine learning , not only to improve the accuracy of the segmentation can make segmentation work to get rid of the reliance on the dictionary . The paper uses JAVA language based on conditional random field segmentation system , the first true the corpus characteristics of information statistics , scalable maintenance feature information library , and then use the the viterbi algorithm of the lexeme sequence the optimal solution in achieving viterbi introduced in the course of the algorithm rules limit in order to remove an invalid combination, thereby reducing the amount of computation of the viterbi algorithm . In the experimental stage , using real corpus for training and testing , and to prepare the efficient performance evaluation algorithm to analyze the results of the segmentation . The research method by the kanji word bit combination rules to identify unknown words , which can do further research on the unknown word identification problems .
|
Related Dissertations
- Based on Data Distribution Characteristics of Text Classification,TP391.1
- Integration of Spatial Information Bag of Feature in Image Annotation,TP391.41
- Semi-supervised Learning and Active Learning of Sentiment Classification Coupled with Domain Knowledge,TP181
- Research on Approaches of the Subjective Automated Assessment,TP391.1
- A Static Behavior-Based Method to Detect Malware on Android,TP309
- Based WebHarvest the Chinese financial news search engine design and implementation,TP311.52
- Chinese XML Compression Technology,TP311.11
- Distortion effects on image quality evaluation and classification,TP391.41
- Content-based malicious code detection Research and Implementation,TP393.092
- Home Academic Information Extraction System,TP393.092
- P2P Traffic Identification Method,TP393.06
- SVM Based on SIFT and scene classification,TP391.41
- Random Forests Feature Selection,TP311.13
- Online Learning and Its Implementation in Intelligent Traffic and Finantial Market,U495;F830
- The Dynamic Distributed network intrusion patterns,TP393.08
- Research and Implementation of Website Search Technology Based on Ajax/Lucene,TP393.092
- Fast Human Detection Based on Feature of Gradients and Cascade Classification,TP391.41
- Chinese Segmentation Algorithm Research Based on Special Identifier,TP391.1
- Research on Chinese Word Segmentation of Search Engine,TP391.1
- A Design and Application of Personalized Information Retrieve and User Recommendation on Search Engine,TP391.3
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net Mobile
|