Abstract
This paper studies speech synthesis for languages with a high degree of phonemic orthography, focusing on the role of stressed sample labels in Tacotron 2 training. The work addresses the challenge of improving naturalness in lower-resource language settings.
Authors
Arnas Radzevičius, Aistis Raudys, Pijus Kasparaitis
What the study evaluates
The study evaluates whether adding stress information to the training data improves prosody and perceived speech quality. The focus is on Lithuanian, where grapheme-phoneme correspondence is relatively strong but stress and rhythm remain important for natural synthesis.
Main findings
Training with stressed sample labels improves the naturalness of generated speech as measured by mean opinion score. The result indicates that even in languages with comparatively transparent spelling, explicit stress information can materially improve neural TTS output.
The work also shows that targeted linguistic annotation can be valuable when large-scale speech resources are unavailable.
Why it matters
High-quality speech synthesis should be available beyond high-resource languages. This paper contributes to that goal by showing how language structure and focused annotation can improve neural synthesis in lower-resource settings.