Dissertation > Excellent graduate degree dissertation topics show

Study on Technology of Corpus Selection for Statistical Machine Translation

Author: LuZuo
Tutor: ZhuJingBo
School: Northeastern University
Course: Computer Software and Theory
Keywords: statistical machine translation corpus selection quality evaluation of sentencepairs corpus’s coverage the Pranking algorithm
CLC: TP391.2
Type: Master's thesis
Year: 2013
Downloads: 3
Quote: 0
Read: Download Dissertation

Abstract


Bilingual training corpus, as an indispensable source of knowledge for statistical machine translation, directly affect the translation performance of systems. In general, the increase in the size of bilingual training data leads to higher translation accuracy. However, as the corpus size increases, the training and decoding cost of machine translation systems also increase. In addition, the noises in the bilingual data can have a negative effect on translation accuracy.In this paper, we discuss the problem of data selection from two different perspectives, data quality and coverage, respectively. The goal is to reduce the size of training data, meanwhile retain the accuracy of translation systems. This way, the cost of training and decoding of machine translation systems can be decreased.We propose multiple features for evaluating the quality of parallel sentence pairs, including the fluency of sentences as well as the likelihood of the sentences in a sentence pair being the translation of each other. We incorporate the features in a linear model and learn feature weights on a labeled dataset by using the Pranking algorithm. Experimental results show that the proposed approach can effectively distinguish between translation sentence pairs of diverse quality and reach an accuracy of83.56%.We propose a model for training data selection targeted at statistical machine translation. In this, we consider both the quality and the coverage of bilingual sentence pairs. Experimental results on the CWMT and NIST datasets show that when a machine translation system selectively uses20%of the whole training dataset, its accuracy can reach97%of the accuracy achieved by using the whole training dataset. When the selected subset is increased to30%, the resulting translation accuracy is comparable with or even higher than the accuracy achieved with the whole training data.We propose to improve the accuracy of a machine translation system by integrating the results of quality evaluation into the training process of the system. Experimental results show that this approach can improve translation accuracy, although improvements achieved are marginal.This thesis is mainly focused on corpus selection for statistical machine translation, including corpus scoring (quality evaluation of parallel data) and data selection for building machine translation systems. In the future, we will go further to study other approaches to parallel corpus processing and will also consider the problem of translation model adaptation.

Related Dissertations

  1. The Research of Decoding Algorithm for Statistical Machine Tranlation,TP391.2
  2. For specific areas of research and application of statistical machine translation,TP391.2
  3. The Research on the Technology of Statistical-Based Chinese-English Machine Translation,TP391.2
  4. Research on Technologies of Chinese and English Verb Subcategorization,H042
  5. The Research of Phrase Extraction Technology for Tibetan and Chinese Statistical Machine Translation,TP391.2
  6. Mining Bilingual Parallel Corpora from Web Automatically and Its Application in Statistical Machine Translation,TP391.2
  7. Research on Automatic Katakana Translation Technology,TP391.2
  8. The Research on English-Chinese Name Entity Translation,TP391.2
  9. Translation Knowledge Acquisition in Corpus-based Machine Translation,TP391.2
  10. Study on Technology of Corpus Processing and Its Quality Evaluation for Statistical Machine Translation,TP391.2
  11. Research and Implement of Statistical Machine Translation System on Handheld Device,TP391.2
  12. Research on Domain Adaptation in Statistical Machine Translation Based on Clustering,TP391.2
  13. Dependency parsing Chinese - Naxi Machine Translation research based on statistics,TP391.2
  14. A Focused Crawler Based on Statistical Machine Translation and Topic Propagation,TP391.3
  15. Research on System Combination in Machine Translation,TP391.2
  16. Consensus Decoding Approaches to Statistical Machine Translation,TP391.2
  17. Research on Discriminative Training Methods for Statistical Machine Translation,TP391.2
  18. Study on Several Key Problems in the Training Process of Phrase-based Statistical Machine Translation,TP391.2
  19. The Research on Bilingual Syntactic Phrase-based Statistical Machine Translation,TP391.2
  20. Research on Word Alignment in Statistical Machine Translation,TP391.2

CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Translator
© 2012 www.DissertationTopic.Net  Mobile