Dissertation > Excellent graduate degree dissertation topics show
Gaussian mixture model based speaker recognition research
Author: JiangWei
Tutor: FanMingZuo
School: University of Electronic Science and Technology
Course: Information and Communication Engineering
Keywords: speaker recognition speech model LPC FCC GMM
CLC: TN912.34
Type: Master's thesis
Year: 2008
Downloads: 490
Quote: 1
Read: Download Dissertation
Abstract
|
Speaker recognition is the processing of automatically recognizing which is speaking by using speaker specific information included in speech signal speaker recognition. In general, it can be classified into text-independent speaker recognition and text-dependent speaker recognition according to recognition condition. This thesis focuses on research of text-independent speaker recognition technology based on Gaussian mixture models (GMM).Firstly, this thesis introduces vocal tract model from acoustic theory of speech production; hereby, introduces all-pole model of speech signal; studies voice production from sonant, unvoiced fricatives, unvoiced plosives using all-pole model of speech signal.Secondly, we study feature extraction, introduce linear prediction algorithm which is used in computing the parameters of all-pole model of speech signal and gives several prediction-derived parameter include reflection coefficient, linear prediction cepstrum coefficient and log area rations coefficient; Also we introduce Mel-Warped cepstrum coefficient and sub-cepstrum coefficient. We study the performance of these features on speakers recognition, the cepstrum coefficient can represent the feature of speaker more accurately, so the cepstrum coefficient has better performance on speaker recognition.In succession, we study the theory of GMM. The individual Gaussian components of a GMM are shown to represent some general speaker-dependent spectral shapes that are effective for modeling speaker identity, these spectral shapes represent a speech classes, for example, phonemes. This thesis discusses estimation of GMM’s parameters, initialization and classification decision. This thesis also makes some improvement on GMM, including making covariance diagonal matrix which can improve computational performance, variance limiting which can avoid model singularities and avoid decrease of recognition performance.We design and implement an automatic speaker recognition system for a complete experimental evaluation of GMM.Finally, a complete experimental evaluation of GMM is conducted with 36 speaker database. The experiments examine model initialization, model order selection and large population performance. Some observations and conclusions are: Identification performance of GMM is insensitive to the method of model initialization; There appears to be a minimum model order needed to adequately model speakers and achieve good identification performance (Sixteen for this 12 speaker database); The GMM maintains high identification performance with increasing population size if the training speech data and test speech data is enough (length of training speech data is large than 90 seconds, length of test speech data is large than 5 seconds).
|
Related Dissertations
- Study on the Chimisorption Removal of FCC Gasoline Sulfur Compound,TE624.55
- Discussion on a New Test of Conventional Asymptotics in GMM,O212.1
- Telephone-based channel voiceprint recognition algorithm,TN912.34
- Complex channel speaker recognition technology,TN912.34
- Research on Speaker Recognition System under the Visual C++ 6.0,TN912.34
- Research on Synthesis Algorithm for Chinese Singing Voice Based on Parametric Modification,TN912.33
- Financial Development and Economic Growth: simultaneous equations econometric model based on research,F832;F124
- Study on Industrial Application of OCT-MD Technology for Catalytic Gasoline Selective Hydrodesulfurization,TE624.55
- Study on Oxidation-extraction Desulfurization of FCC Diesel Oil Oxidized with Q3PW(Mo)12O40,TE624.55
- Mixing characteristics and Gaussian mixture model - based speaker recognition,TN912.34
- Design and Implementation of Speaker Recognition System Based on Windows CE,TN912.34
- Study of Extraction and Optimization Characteristic Parameters in Speaker Recognition,TN912.34
- Audio Architecture Technology Research,TN912.3
- The Research Based on the Text Irrelevant Speaker Distinguishes,TN912.34
- The Empirical Analysis on the Relationship between the Development of State-Owned Banks and Economic Growth,F124;F224
- Emperical Studies on Interest Rate Rules in China,F822.0
- Asset Prices’ Influence on Monetary Policy,F822.0
- Study on Extractive Distillation Phase Equilibrium of FCC Gasoline Desulfurization Process,TE624.5
- Study on the Extractive Desulfurization of the FCC Gasoline,TE624.5
- Information Extraction and Quantitative Analysis of Positive Signals in Tomographic Chip,TP391.41
- Research and Design of Dynamic Human Detection System under Static Background,TP391.41
CLC: > Industrial Technology > Radio electronics, telecommunications technology > Communicate > Electro-acoustic technology and speech signal processing > Speech Signal Processing > Speech Recognition and equipment
© 2012 www.DissertationTopic.Net Mobile
|