|
To solve common Chinese speech synthesis system , and synthesized speech expressive synthetic speech intelligibility and naturalness of pronunciation and prosodic structure , in particular, to make it accurate , vivid semantic performance capabilities , the need text analysis module can output the richer linguistic information , and use this information to synthesize more accurate , vivid voice . To this end, the paper focuses on : speech database and its the rhythm marked , and synthesis system front-end module , expand research . This paper, on the one hand the speech synthesis processing of the text object by the statement level rise to the chapter level , to build a database chapter news broadcast ; the other hand, the speech synthesis of the front end - text processing module for a certain degree of improvement . The main tasks are as follows : 1 . Selected news broadcast corpus Research / machining material , considering the demand for computational modeling and sample characteristics , developed on the basis of previous work , a chapter level prosodic labeling specifications . Rhythm label includes : prosodic hierarchy , accent , tone and intonation , enrich and improve the existing rhythm description ; implemented based on the enacted norms marked the the chapter level newscast database was constructed . 2 based on the original text processing workflow module , starting from its integral part , on the basis of comparative analysis module in the original text segmentation module with other word processing module , a segmentation module replacement ; using binary grammar training new text processing module , the core training new prosodic structure prediction module ; experimental test set and set outside to prove the new module to achieve a better rhythm structure prediction , and thus make the text processing effect to a certain extent been improvements.
|