Dissertation > Excellent graduate degree dissertation topics show

Research on Time-Varying Robustness in Speaker Recognition

Author: WangLinLin
Tutor: ZhengFang
School: Tsinghua University
Course: Computer Science and Technology
Keywords: speaker recognition time-varying issue time-varying robustness
CLC: TN912.34
Type: PhD thesis
Year: 2013
Downloads: 17
Quote: 0
Read: Download Dissertation

Abstract


The focus of this dissertation is the time-varying issue in speaker recognition andthe time-varying robustness is explored. Major efforts and contributions are:1. A proper longitudinal voiceprint database that specially focuses on thetime-varying issue. After analyzing existing speech databases with the time-varyingattribute, we designed to create a fixed-text read speech database with16recordingsessions within a time span of3years. Since the time-varying effect was the only focus,other factors, such as recording equipment, software, conditions and environment werekept as constant as possible throughout all recording sessions. Gradient time intervalswere used, with the length of intervals increasing gradually.2. Performance evaluation index for a time-varying speaker recognition system.For a time-varying speaker verification task, there are generally a series of EERs,corresponding to each recording session. Then when comparing the performance of twosystems, we are indeed comparing two arrays of EERs. Therefore, it is natural to usemean and standard deviation of each array of EERs to evaluate the overall performanceof a system. The mean value serves as an indicator of the averages performance ofsessions, while the standard deviation value serves as an indicator of the time-varingrobustness across sessions. Specifically in this paper, the product of those two values isused to evaluate the overall time-varying speaker verification performance.3. Time-varying robust feature extraction algorithms with discriminationsensitivity of frequency bands calculated through F-ratio. The concept of overalldiscrimination sensitivity of frequency bands regarding the time-varying speakerecognition task was proposed. Efforts were made to identify frequency bands thatrevealed high discrimination sensitivity for speaker-specific information, while lowdiscrimination sensitivity for time-varying session-specific information. F-ratio wasemployed as an intermediary criterion to calculate the overall discrimination sensitivitybased on the log-energy spectrum. Thus according to the overall discriminationsensitivity, tme-varying robust feature extraction algorithms were presented duringfeature extraction of cepstral coefficients with different emphasis on different frequencybands from two aspects: pre-filtering frequency-warping and post-filtering filter-bankoutputs weighting. Experimental results showed that the two algorithms outperformedthe baseline MFCC by26.90%and5.45%, respectively. 4. Performance-driven feature extraction algorithm based on frequency warping.This algorithm evaluated the overall discrimination sensitivity of frequency bands froma performance-driven point of view instead of the F-ratio criterion. Specifically, theoverall discrimination sensitivity of a designated frequency band is determined by theoverall performance of a time-varying speaker recognition system, which made use offrequency-warping approach to soly emphasize the designated frequency band, leavingother unchanged. Finally, frequency warping was performed and experimental resultsshowed that it yielded a better result than MFCC, with a gain of32.47in overallperformance.5. Discriminative feature extraction algorithm based on filter-bank outputsweighting. This was also a performance-driven approach, yet it was designed for thefilter-bank outputs weighting method. After resigning an initial series of weights forfilter-bank outputs, speaker modeling and utterance scoring were performed; thenaccording to the performance feedback, the series of weights were adjusted by theproposed MCE*MSV criterion. After several iterations of such a process, the bestseries of weights were found automatically. The MCE*MSV criterion wasproposed to minimize the target optimization function of the error rates ofrecording sessions and their standard deviation. The best series of weights wereapplied to filter-bank outputs and experimental results showed that it workedbetter than MFCC by34.08%.

Related Dissertations

  1. Rejection research strategy based SVM speaker,TN912.34
  2. Speaker Recognition Based on SOPC controller,TN912.34
  3. Auditory System Property about Speech Signal Process,TN912.3
  4. Complex channel speaker recognition technology,TN912.34
  5. Research on Speaker Recognition System under the Visual C++ 6.0,TN912.34
  6. Mixing characteristics and Gaussian mixture model - based speaker recognition,TN912.34
  7. Design and Implementation of Speaker Recognition System Based on Windows CE,TN912.34
  8. Study of Extraction and Optimization Characteristic Parameters in Speaker Recognition,TN912.34
  9. The Research Based on the Text Irrelevant Speaker Distinguishes,TN912.34
  10. Experimentation Research of Speaker Recognition System Based on VQ and DTW,TN912.34
  11. Research and Implementation on Real-time Speaker Recognition Algorithmin in Multiplexing Parallel Model,TN912.34
  12. The Development of Speaker Recognition System Based on Support Vector Machine,TN912.34
  13. Research on Text-independence of Open-set Speaker Recognition,TN912.34
  14. Speaker Recognition Research in Noisy Environment,TN912.34
  15. EMD-based speaker recognition,TN912.34
  16. Rapid Speaker Recognition Based on GMM-UBM,TN912.34
  17. Research on Real-Time Audio Decoding and Robust Speaker Recognition System in Network Environment,TN912.34
  18. Research of Text Dependent Speaker Recogniton on Embedded System,TN912.34
  19. Speaker Recognition on the Base of Time-Varying Characteristics of Speech Signal,TN912.34
  20. Research on Speaker Recognition Technology Based on Wavelet Analysis,TN912.34
  21. Research of A Feature Extraction in Speaker Recognition,TN912.34

CLC: > Industrial Technology > Radio electronics, telecommunications technology > Communicate > Electro-acoustic technology and speech signal processing > Speech Signal Processing > Speech Recognition and equipment
© 2012 www.DissertationTopic.Net  Mobile