Dissertation > Excellent graduate degree dissertation topics show

Cascade Consistency Check of Segmentation of the Chinese Corpus

Author: LiuBo
Tutor: ZhengJiaHeng
School: Shanxi University
Course: Computer Software and Theory
Keywords: Corpus Consistency Verify Inconsistency of Segmentation Word Segmentation
CLC: TP391.1
Type: Master's thesis
Year: 2008
Downloads: 130
Quote: 1
Read: Download Dissertation

Abstract


In the field of Chinese natural language processing,it pays attention to research and obtain knowledge automatically based on large-scale true text corpus,and it becomes primary task that building the large-scale and high-quality corpus.But because of the setting-up of the corpus now,it is need to check-up artificially,the unavoidable carelessness and mistake will cause inconsistent to the same word segmentation results under the same language environment.These inconsistent phenomena have not merely influenced the corpus precision,and take mistake to next process step that utilize corpus resource.So,while processing the corpus,we must check and collate the segmentation consistency to guarantee the corpus quality.So it is important standard that evaluate segmentation corpus quality.Aimed at the problems in the large-scale corpus,we submit a cascaded method to solve Segmentation inconsistent based on the research of inconsistent strings in SXU and MSR corpus,which deals with the segmentation consistency using the database of the rules,then,we make use of the statistic model to collate the corpus.The experience has got the anticipated aim,which proves that the combined method can solve the segmentation inconsistent availably.The main task is as follows:1.Based on the statistic and analysis of the inconsistency of segment for Chinese corpus that has 4 million Chinese characters,we defined the main types of structure for the segment inconsistencies,confirm the research object of the thesis and put the aspects which influenced the precision score as foundation of rules database;2.A rules and examples method was put forward.We extract initial rules and a lot of examples,which apply to collate results.We enhance segmentation corpus quality by using rules self-learning measure;3.We propose a statistical method which can largely enhance segmentation corpus quality;we use vector model expresses abstracted inconsistency strings and word environment,and using synonymous database when computing similarity.Computing similarity and classify measures were used to getting probability value of inconsistency strings.At last,strings classified by mensurable measure.4.Based on the above thought and the method,experimental models are designed respectively that method on the basis of rule,method on the basis of statistical,and method that combined two kinds of method together.A test was done on the corpus,the result prove the methods put forward in the thesis is effective.In the process,we found it is hard to fine one method to resolve this question.So we use a combined method to class.Using classing method,the inconsistent segmentation that have the same structure and similar functions were gathered together,can be treat with the uniform segmentation format. SXU corpus has a good achievement in SIGHAN2007 bakeoff.To carrying on open tests of the combined corpus segmentation consistency check-up system,the precision of consistency checking is 84.50%,the recall is 70.39%, it is obvious that the corpus quality can really improved after adopting the system.

Related Dissertations

  1. Business English professional translators based teaching model Corpus,H319
  2. Research and Implement of Chinese Word Segment Techniques Based on the Conditional Random Field,TP391.1
  3. English Academic title phrase characteristics of,H313
  4. The Research of MRI Measurements on Various Lines and Cross-Sectional Area of Fetal Corpus Callosum,R714.5
  5. A Corpus-based Comparative Study of College English Coursebooks,H319
  6. Research on the Construction of a Nlp-oriented Chinese Sentence Semantic Knowledge Database,H13
  7. The Study and Analysis of Oracle Bone Inscriptions Based on Statistical Natural Language Processing,TP391.1
  8. The Research and Implemenation of the Chinese Word Segmentation System Combining Omini-Segmentation with Statistic,TP391.1
  9. Design and Implementation of Indexing Mechanism for Image Information Based on the Berkeley DB,TP391.3
  10. Corpus -based English language restate comparative study marks,H314
  11. Corpus-based Chinese Non-English Majors Have usage study,H319
  12. Sentence- aligned parallel corpus and its application in machine translation,H315.9
  13. The Lexical Phrases and Spoken Language Proficiency of English Majors,H319
  14. Design and Implementation of Data Management System of the Harmful Internet Video Information,TP393.06
  15. The Study and Analysis of Oracle Configuration Based on Support Vector Machine Analysis,TP391.1
  16. A Corpus-based Study of Image Schema Reflected in English HIGH/LOW and Chinese GAO(高)/DI(低),H05
  17. A Corpus-Based Study of Poetry Translation Criteria,I046
  18. \,H212
  19. A Corpus-based Account of the Lexical Use by High School English Learners,G633.41
  20. Corpus-based comparative study of synonyms semantic prosody,H319
  21. Probability and statistics based on dictionaries and Chinese word segmentation algorithm,TP391.1

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net  Mobile