Dergiler / Turkish Journal of Electrical Engineering and Computer Sciences / 2019 / Cilt: 27 - Sayı: 5
Scale-invariant MFCCs for speech/speaker recognition
- Sayfa
- 3758–3762
- DOI
- —
Abstract
The feature extraction process is a fundamental part of speech processing. Mel frequency cepstral coefficients(MFCCs) are the most commonly used feature types in the speech/speaker recognition literature. However, the MFCCframework may face numerical issues or dynamic range problems, which decreases their performance. A practicalsolution to these problems is adding a constant to filter-bank magnitudes before log compression, thus violating thescale-invariant property. In this work, a magnitude normalization and a multiplication constant are introduced to makethe MFCCs scale-invariant and to avoid dynamic range expansion of nonspeech frames. Speaker verification experimentsare conducted to show the effectiveness of the proposed scheme.