Sociedade Brasileira de Telecomunicações · desde 1983 secretaria@sbrt.org.br
← SBrT2018

Speech Synthesis Based on Deep Neural Networks with Direct Modeling of Amplitude Spectra

Ranniery Maia, Rui Seara
Deep learningdeep neural networksspeech syn- thesistext-to-speech (TTS) systems

Resumo

In recent state-of-the-art text-to-speech systems, usually a sequence of graphemes is directly mapped onto the speech waveform using deep neural networks. Despite reaching very high quality, these approaches tend to be computationally costly at synthesis time and its training implementation is usually not trivial. In this paper, a method which can be interpreted as a simplified version of these systems is proposed. Here, framebased smoothed log spectra, fundamental frequency, and phase information are modeled at training time, while synthesis runs in a straightforward fashion. Experiments show that the proposed approach outperforms traditional ones using acoustic modeling of speech features.