Text to Speech Synthesis: New Paradigms and AdvancesText to speech synthesis (TTS) is a critical research and application area in the field of multimedia interfaces. Recent advances in TTS will impact is wide number of disciplines from education, business and entertainment applications to medical aids. Until recently, speech synthesis relied on models and rule-based approaches. While this had yielded intelligible sounding speech, the voice quality was unacceptable for widespread adoption. Fortunately, there has been a major technological paradigm shift recently in how speech synthesis is done: going from rule-based to explicit data-driven methods. Recent advances in computing and corpus driven methodologies have yielded exciting possibilities for research and development in this domain yielding highly natural sounding speech. The book focuses on recent advances and new paradigms in text to speech synthesis contributed by leading experts from both academia and industry from across the world. There is no book of this nature that documents in a comprehensive way the recent research trends. This is not only important for researchers and students of the field but potential customers and other benefactors of the results. The book's chapters address key current topic areas in text to speech synthesis (TTS): Data-driven systems, unit selection Hybrid Schemes: interplay between data-driven and knowledge-based techniques, prosody models and generation and expressive speech synthesis. |
Contents
REDUCING DISCONTINUITIES AT SYNTHESIS TIME | 1 |
viii | 5 |
TOWARD EXPRESSIVE SYNTHETIC SPEECH | 11 |
Copyright | |
10 other sections not shown
Common terms and phrases
accent acoustic measures algorithm analysis approach articulatory synthesis cepstral computed concatenation cost concatenative speech synthesis corpus-based correlation coefficient diphone discontinuities distance measures duration models emotional speech Euclidean distance Eurospeech evaluation expressive speech feature vectors Figure formant frames frequency glottal hidden Markov models HMM-based speech synthesis ICASSP ICSLP IEEE inventory join cost function listening tests LSFs Mahalanobis distance MBROLA MCA coefficients MFCC mora n-gram natural speech neutral optimal output palatalized consonant parameterizations parameters pause perceptual scores phoneme phoneme duration Pitch-connection pattern prediction Proc quantification theory quantification theory Type recorded Section segments sentences sequence Signal Processing smoothing speaker speaker recognition spectrum speech corpus speech rate speech recognition speech signal speech synthesis system speech units stimuli syllable synthetic speech Table techniques text to speech tion Tokuda TP-MBROLA TTS system unit selection unvoiced utterances values variables vocal tract voice quality difference vowel waveform word