← SBrT2018
Speech Synthesis Based on Deep Neural Networks with Direct Modeling of Amplitude Spectra
Deep learningdeep neural networksspeech syn- thesistext-to-speech (TTS) systems
Resumo
In recent state-of-the-art text-to-speech systems,
usually a sequence of graphemes is directly mapped onto the
speech waveform using deep neural networks. Despite reaching
very high quality, these approaches tend to be computationally
costly at synthesis time and its training implementation is usually
not trivial. In this paper, a method which can be interpreted as
a simplified version of these systems is proposed. Here, framebased smoothed log spectra, fundamental frequency, and phase
information are modeled at training time, while synthesis runs in
a straightforward fashion. Experiments show that the proposed
approach outperforms traditional ones using acoustic modeling
of speech features.