Dissertation > Excellent graduate degree dissertation topics show

The Study of Speech Codec Based on Perceptual Quality

Author: YangJie
Tutor: YuShengSheng
School: Huazhong University of Science and Technology
Course: Computer System Architecture
Keywords: Speech coding Perceptual weighting Subspace After filtering Embedded platform
CLC: TN912.3
Type: PhD thesis
Year: 2010
Downloads: 181
Quote: 1
Read: Download Dissertation

Abstract


Voice communication is an efficient way to communicate. In the actual environment, the speech signal will be influenced by the external environment, and network transmission would be further introduction of interference. With the rapid development of mobile communication systems, various communications organization launched and improve the standard variable rate speech coding and stereo speech coding. How to get better sound quality, it comes to speech coding, speech enhancement, decoded voice signal post-processing. The sound quality of the voice through the human auditory perception test, the human auditory system is a complex psychological, physiological and physical conversion process. Noisy Speech for pure voice, and to improve the clarity and intelligibility of speech is the focus of the study, the main research work involved is as follows: The perception of the noise will be greatly reduced under the residual noise in the auditory masking. Subspace enhancement method first by removing the noise subspace, then by the inverse transform of the signal subspace of the added weight coefficient for pure voice. Masking characteristics of the human ear is related band, and key and key band is based on the characteristics of the human ear cochlea reasoning subspace denoising If you need to use the masking properties of the frequency domain to the conversion feature domain. Calculated from the correlation matrix, obtained by eigenvalue decomposition of the matrix of the eigenvalues ??and eigenvectors. Through the power spectrum density calculating a masking threshold obtained after the inverse transform, and then use the feature filter gain to calculate the linear prediction matrix, enhanced voice. Voice pitch and harmonic components have a significant cyclical. In contrast, more important than the trough auditory perception crest. The adaptive variable rate coding (Adaptive Multi Rate, AMR) to search for the linear prediction coefficients in accordance with perception smallest weight standards excitation signal. Accordance with the theory of auditory masking noise perceived formant at relatively insensitive to the human ear. Input speech before the open loop pitch search through a perceptual weighting filter. The adjustment of the coefficients to control the response of the filter. According to the energy value can roughly determine whether the frame is a voice frame. And weaken noise by setting the threshold, use the comb filter to strengthen the voice in the cycle components, able to obtain better denoising effect. When the network transmission quality is poor, the channel can not be error control, the decoded segment will receive the error signal of the frame. General error concealment method based on the signal around the correct frame interpolation or replace reconstruction error frame. Adaptive codebook correlation will affect the speed of recovery of the error frame. When the higher energy limited its contribution excitation signal can reduce the inter-frame correlation. When consecutive error frames received, the pitch value is not a simple self-growth, but fluctuate within a certain range, the cumulative deviation can be avoided. Gain coefficient dressing can also be in the wrong end of the frame, as soon as possible to resume normal decoding. The characteristics of embedded technology is very suitable for the the terminal market trends. Taking into account the limited resources of embedded devices, the need to reduce the complexity of the application. ARM-based mobile phone platform, porting and optimization of adaptive variable rate coding. Algebraic codebook search is an important and complex part of the standard of the voice. , Contains pulses corresponding to the different track. Different location, search for the smallest standard in accordance with the mean square error between the weighted input speech and weighted reconstructed speech. The entire search is nested, which brought a large amount of computation. Can take advantage of a more efficient search algorithm to obtain the corresponding pulse position. Thus saving coding time, to improve the coding efficiency. Also optimize the characteristics of the ARM instruction and platform. Significant savings under the premise of ensuring the sound quality, the encoding complexity. The study of transplantation of multimedia applications on embedded devices is beneficial.

Related Dissertations

  1. Signal Detection Circuit Design for Slow-Light Optic Fiber Gyroscope,V241.5
  2. The Research of Navigation Multi-Sensor Data Fusion Technology Based on Micro Unmanned Platform,V249.32
  3. Attitude Determination and Finite Time Control Algorithms for a Satellite,V448.222
  4. Establishment and Update of Similar Users’ Cluster in Personalized Information Retrieval,TP391.3
  5. 3D Visualization Research of Medicial Ultrasound Image,TP391.41
  6. Active faults based radar image information extraction method applied research and demonstration,P542.3
  7. Research on Personalized Recommendation Algorithm Based on Natural Forgetting,TP311.52
  8. Combined DWT Dynamic Data Reconciliation Research and Application,TP274
  9. Design and Implementation of Embedded Multi-parameter Intelligent Evironment Monitoring Systems,TP274
  10. The Problem of Reducing Subspaces for a Class of Reproducing Kernel Spaces,O177
  11. Preparation and Properties of Filtering Ceramics for High Temperature Exhaust Gas,TQ174.6
  12. Study on Wireless Sensor Networks with Communication Constraints,TN929.5
  13. Coherent Source DOA Estimation Based on a Single Snapshot of the Uniform Circular Array,TN911.23
  14. Linear Operator Broadcast Channels,TN911.22
  15. Design and Realization of Channel Estimation for MIMO-OFDM Systems Based on Subspace,TN919.3
  16. Study on Characteristic Products of Solid-insulation Aging of Power Transformer,TM855
  17. Lightning detection system hardware design,TN202
  18. Online Education News Text Categorization System Design and Implementation,TP391.1
  19. Integrated Access Device Configuration - Management Subsystem Design and Implementation,TP311.52
  20. Band Entropy Method and Its Application to Fault Diagnosis of Rolling Bearings,TH165.3
  21. Research on Filtering for Local Strongly Coupled Systems,TN713

CLC: > Industrial Technology > Radio electronics, telecommunications technology > Communicate > Electro-acoustic technology and speech signal processing > Speech Signal Processing
© 2012 www.DissertationTopic.Net  Mobile