Dissertation > Excellent graduate degree dissertation topics show

Research on Synthesis Algorithm for Chinese Singing Voice Based on Parametric Modification

Author: LiJinZuo
Tutor: YangHongWu
School: Northwest Normal University
Course: Circuits and Systems
Keywords: Text to Speech Melody control model GMM MIDI STRAIGHT algorithm
CLC: TN912.33
Type: Master's thesis
Year: 2011
Downloads: 50
Quote: 0
Read: Download Dissertation

Abstract


The speech synthesis technology is an important research content in the field of Human-computer interaction. Singing voice synthesis is a branch of the speech synthesis, and it has also become a hot topic in recent years. In order to generate a Mandarin song, the Text To Speech (TTS) technology and voice modification technology are combined together in this paper. Firstly, the differences between speech and singing voice signal are analyzed. Then, the melody control model and spectrum model which based on Gaussian Mixture Model are established. Lastly, a singing voice synthesis system is achieved with variable timbre.The results of our research can reveal inherent differences between the speech and singing voice, and have an important theoretical significance for the speech synthesis. In addition, our system can be applied to entertainment areas and even some music creation fields. It makes the singing voice synthesis more entertaining and interesting. The main research results and innovation are as follows:Firstly, the differences between the speech and the singing voice are compared by analyzing their acoustic parameters. The speech and singing voice are all produced by the same organs, while the perceived effect of them is different. The differences are analyzed in the paper. Speech main transfer semantic content, and singing voice main transfer emotions. The singing voice has rich harmonic and its HHG still has high energy, and the trend of the fundamental frequency and other harmonic are very smooth. Compared with singing voice, the speech’s trends are related with speech’s tones.Secondly, a new singing melody control model is constructed. According the difference between them in fundamental frequency (F0), the F0 control model is mainly based on vibrato. The melody control model is adopted to achieve the lyrics to song conversion by modifying the acoustic parameters of speech signal. The MOS test shows that the average MOS score of synthetic song before timbre conversion is above 3.29.Thirdly, a singing’s spectrum modification model is proposed in this paper, which is based on GMM. Constructing the STRAIGHT spectrum model by using GMM, and the singing’s timbre can be converted. The ABX test demonstrates that the accuracy can be up to 100% in the case of k = 0 or 1, and it can be higher than 62.5% in the case of 0 < k< 1. The professional singer’s timbre can be added proportionally. The experiments also show the mean of GMM has greater impact on a singer’s timbre than weight ratio and covariance.Finally, a new singing voice synthesis system based on STRAIGHT is constructed in this paper. The system consists of text analysis module, concatenative synthesis model, melody control model, and timbre control model. Obtaining the pronunciation information of the lyrics through the textual analysis process, and getting the synthesized speech by using the concatenative model. The melody control model is used to achieve singing voice synthesis, and the timbre control module is used to convert the singing’s timbre.

Related Dissertations

  1. Compensation Methods of Different Speech Coding for Speaker Recognition,TN912.34
  2. Discussion on a New Test of Conventional Asymptotics in GMM,O212.1
  3. Telephone-based channel voiceprint recognition algorithm,TN912.34
  4. MIDI-based instrument control systems and automatic identification method notes,TN912.34
  5. Windows CE-based monitoring room management system design and development,TP311.52
  6. Research of Embeded Speech Synthesis Technology,TN912.33
  7. Financial Development and Economic Growth: simultaneous equations econometric model based on research,F832;F124
  8. The Interactive Music Play Table,TP391.41
  9. Mixing characteristics and Gaussian mixture model - based speaker recognition,TN912.34
  10. Audio Architecture Technology Research,TN912.3
  11. The Empirical Analysis on the Relationship between the Development of State-Owned Banks and Economic Growth,F124;F224
  12. Emperical Studies on Interest Rate Rules in China,F822.0
  13. Asset Prices’ Influence on Monetary Policy,F822.0
  14. On the Computer Music and Polyphony Music Composition,J619
  15. Information Extraction and Quantitative Analysis of Positive Signals in Tomographic Chip,TP391.41
  16. Research and Design of Dynamic Human Detection System under Static Background,TP391.41
  17. Research on Broadcast News Audio Structure Analysis,TN912.3
  18. The Design and Implementation of a Multi-structured Call Center,TN99
  19. Speaker Verification Based on Factor Analysis,TN912.34
  20. The Research of Animal Behavior Recognition System Based on Sound Features,TN912.34

CLC: > Industrial Technology > Radio electronics, telecommunications technology > Communicate > Electro-acoustic technology and speech signal processing > Speech Signal Processing > Speech synthesis
© 2012 www.DissertationTopic.Net  Mobile