Abstract
This paper studies emotional prosody control in text-to-speech systems without relying on predefined emotion labels. It combines StyleTTS 2 embeddings with HDBSCAN clustering over multimodal acoustic features to discover prosodic structure directly from data.
Authors
Arnas Radzevičius, Žygimantas Girdauskas, Rokas Sabaitis, Aistis Raudys
What the study evaluates
The work evaluates whether clustering methods can identify useful emotional or expressive speech patterns in a setting where manually labelled emotion categories are unavailable or too coarse for practical synthesis control.
Main findings
The proposed approach produces substantially richer grouping of expressive speech than the baseline. By combining neural embeddings with acoustic features, the method increases the number of discovered emotion clusters and reduces the amount of speech left unclustered.
These results suggest that unsupervised prosody discovery can provide a practical control layer for expressive speech synthesis, especially in languages or datasets where labelled emotional speech is limited.
Why it matters
Expressive voice systems need more than intelligible pronunciation. They need controllable rhythm, emphasis, and affect. This work moves toward synthesis systems that can learn those expressive dimensions from the speech signal itself.