Dissertation > Excellent graduate degree dissertation topics show
The Research of Tibetan Text Categorization Base on N-Gram Information
Author: ZhouDeng
Tutor: DaiYuGang
School: Northwest University for Nationalities
Course: Applied Computer Technology
Keywords: text classification N - Grammodel corpus Naive Bayes Multinomial
CLC: TP391.1
Type: Master's thesis
Year: 2010
Downloads: 47
Quote: 0
Read: Download Dissertation
Abstract
|
With the coming of information age, information resources manifest an explosive growth throughout internet. For the sake of not being submerged by mass of information, good measures should be taken to category and manage them. Text categorization plays an important role on how to use that information. It can effectively organize and manage information so that the efficiency of information retrieval will be improved.Research on text classification and its related technologies are done in the paper, and this paper proposes an approach for Tibetan text categorization without word segmentation. Compared with traditional text classification models, the approach based on character-level N ? Gram language modeling avoids word segmentation so that much computing resources during pre-processing will be saved.This paper firstly introduces the recent research situation of text categorization all over the world; secondly the commonly used text express models are analyzed and N-Gram model for Tibetan text is further studied, while the selection of parameter N and every function of Tibetan text categorization system are discussed. Thirdly, core function classifier of the system is presented in details in Chapter Four. The paper also proposes a corpus Naive Bayes Multinomial, which combines tightly with N ? Gram and makes good classifying effect. Finally, because of a large number of overlapped high-degree bigrams and biased high-degree bigrams in bigram feature set we put forward a novel feature reduction method,δ? OR,which can raise theδ? degreeoverlapped bigrams to corresponding trigrams. The experiment shows thatδ? OR method can not only achieves feature reduction as well deletes redundant information but also improves abilities to describe features and to classify features. In some degree of reduction, the categorization effect can be reasonably made better.
|
Related Dissertations
- Research on Text Classification Based on Biomimetic Pattern Recongnition,TP391.1
- Tourism Comments on the Internet’s Semantic Analysis and Usefulness Research,TP391.1
- Business English professional translators based teaching model Corpus,H319
- ESP genre research -based teacher training model,H319
- Based on Data Distribution Characteristics of Text Classification,TP391.1
- A Study of the Acquisition of Chinese Progressive Complex Sentence Based on the Interlanguage Corpus,H195
- A Corpus-Based Intertextuality Analysis of the Reportage on Shanghai World Expo,H052
- A Corpus-based Study on That-Clause in Political Documents,H314
- A Corpus-driven Comparative Study of Uses of TAKE,H319
- A Corpus-Based Study on English Personal Pronouns in Political Speeches and Statements,H314
- Research on Ontology-based Scientific Papers of Chinese Classification,TP391.1
- Chinese Keyword Extraction Method Based on Word Span and Its Application in Text Classification,TP391.1
- A Contrastive Study of the Translation of Nominal Concept in Taoism,H059
- A Corpus-based Study on Verb-noun Collocation Errors by Chinese English Majors from the Perspective of Conceptual Transfer,H319
- A Corpus-based Study on the Chinese Learners’ Use of the Word AS in Light of the Lexical Priming Theory,H314
- The Research of Using BNC to Correct the Chinese English Emerging in the English Writing,G633.41
- A Research on the Approach of Composition Teaching Reform in Preliminary School,G623.24
- Research on Improved K Neighbor Support Vector Machine Algorithm Faced Text Classification,TP391.1
- The Diagnostic Value of the Measurement of Corpus Callosum Area on MRI Midsagittal Section to Brain Atrophy,R445.2
- An Empirical Study of Corpus-Based Lexical Approach to Non-English Majors’ Writing,H319
- The Research on Cllective Multi-Label Classification,TP391.1
CLC: > Industrial Technology > Automation technology,computer technology > Computing technology,computer technology > Computer applications > Information processing (information processing) > Text Processing
© 2012 www.DissertationTopic.Net Mobile
|