Title:
|
Using voice-quality measurements with prosodic and spectral features for speaker diarization
|
Author:
|
Zewoudie, Abraham Woubie; Luque, Jordi; Hernando Pericás, Francisco Javier
|
Other authors:
|
Universitat Politècnica de Catalunya. Departament de Teoria del Senyal i Comunicacions; Universitat Politècnica de Catalunya. VEU - Grup de Tractament de la Parla |
Abstract:
|
Jitter and shimmer voice-quality measurements have been successfully used to detect voice pathologies and classify different speaking styles. In this paper, we investigate the usefulness of jitter and shimmer voice measurements in the framework of the speaker diarization task. The combination of jitter and shimmer voice-quality features with the long-term prosodic and short-term spectral features is explored in a subset of the Augmented Multi-party Interaction (AMI) corpus, a multi-party and spontaneous speech set of recordings. The best results have been obtained by fusing the voice-quality features with the prosodic ones at the feature level, and then fusing them with the spectral features at the score level. Experimental results show more than 20% relative DER improvement compared to the spectral baseline system. |
Abstract:
|
Peer Reviewed |
Subject(s):
|
-Àrees temàtiques de la UPC::Enginyeria de la telecomunicació::Processament del senyal::Processament de la parla i del senyal acústic -Voice Quality -Speaker diarization -Spectral features -Jitter -Shimmer -Prosody -Fusion -Veu, Processament de |
Rights:
|
|
Document type:
|
Article - Published version Conference Object |
Published by:
|
International Speech Communication Association (ISCA)
|
Share:
|
|