AAI Labs
Main
Services
CasesResearchTeam
Company
More
Contact
Loading…

Services

  • AI for Energy
  • AI for Municipalities
  • AI for Transport
  • Generative AI
  • LLM for Business
View all services →

Company

  • About
  • Careers
  • Cases
  • Privacy Policy
  • Public R&D
  • Contact

More

  • EU AI Act hub
  • AI dictionary
  • Research
  • Blog
  • News

Products

  • AI Team for Hire
  • Merkys.AI
  • Cargobroker.AI
  • Klikt

UAB Taikomasis dirbtinis intelektas © 2026

[email protected]LinkedIn →
Research paper2026

Emotional Speech Synthesis Approach Using Prosody-Based Clustering

Controlling emotional prosody in TTS without predefined labels is an open challenge. We combine StyleTTS 2 embeddings with HDBSCAN clustering on multi-modal acoustic features, achieving 200% more emotion clusters and 21.4% less unclustered data versus the baseline.

Read the paper

Abstract

This paper studies emotional prosody control in text-to-speech systems without relying on predefined emotion labels. It combines StyleTTS 2 embeddings with HDBSCAN clustering over multimodal acoustic features to discover prosodic structure directly from data.

Authors

Arnas Radzevičius, Žygimantas Girdauskas, Rokas Sabaitis, Aistis Raudys

What the study evaluates

The work evaluates whether clustering methods can identify useful emotional or expressive speech patterns in a setting where manually labelled emotion categories are unavailable or too coarse for practical synthesis control.

Main findings

The proposed approach produces substantially richer grouping of expressive speech than the baseline. By combining neural embeddings with acoustic features, the method increases the number of discovered emotion clusters and reduces the amount of speech left unclustered.

These results suggest that unsupervised prosody discovery can provide a practical control layer for expressive speech synthesis, especially in languages or datasets where labelled emotional speech is limited.

Why it matters

Expressive voice systems need more than intelligible pronunciation. They need controllable rhythm, emphasis, and affect. This work moves toward synthesis systems that can learn those expressive dimensions from the speech signal itself.