Dergiler / Turkish Journal of Electrical Engineering and Computer Sciences / 2019 / Cilt: 27 - Sayı: 4

Unsupervised deep feature embeddings for speaker diarization

Sayfa
3138–3149
DOI
—

Abstract

Speaker diarization aims to determine “who spoke when?” from multispeaker recording environments. Inthis paper, we propose to learn a set of high-level feature representations, referred to as feature embeddings, from anunsupervised deep architecture for speaker diarization. These sets of embeddings are learned through a deep autoencodermodel when trained on mel-frequency cepstral coefficients (MFCCs) of input speech frames. Learned embeddings are thenused in Gaussian mixture model based hierarchical clustering for diarization. The results show that these unsupervisedembeddings are better compared to MFCCs in reducing the diarization error rate. Experiments conducted on the popularsubset of the AMI meeting corpus consisting of 5.4 h of recordings show that the new embeddings decrease the averagediarization error rate by 2.96%. However, for individual recordings, maximum improvement of 8.05% is acquired.