Dergiler / Gazi Üniversitesi Mühendislik Mimarlık Fakültesi Dergisi / 2009 / Cilt: 24 - Sayı: 4

Türkçe metinden konuşma sentezleme uygulamaları için bir veri sözlük seti ve yazılım çerçevesi

A lexicon set and software framework for Turkish text-to-speech synthesis applications

Sayfa
735–744
DOI
—

Özet

Bu çalışmada, Türkçe metinden konuşma sentezleme uygulamaları için altyapı sağlayacak olan bir veri sözlük seti ve yazılım çerçevesi tanıtılmaktadır. Söz konusu altyapı, Microsoft.NET 2.0 ortamında geliştirilmiş fonksiyonlar ve XML tabanlı kural betiklerinden oluşmaktadır. Altyapı fonksiyonları, verilen bir metnin doğru fonetik gösteriminin yapılabilmesi; metin içerisindeki kısaltmaların, sayısal değerlerin doğru bir şekilde okunabilmesi için uygun şekilde yazılı hale dönüştürülmesi vb. görevleri yerine getirmektedir. Bu çalışma kapsamında, fonetik gösterim için özel bir sembol kümesi de tanımlanmış olup bu makalede detaylı olarak tanımlanmıştır. Altyapının, gelecekte yapılacak olan metinden konuşma sentezleme uygulamaları için standardizasyon sağlaması hedeflenmektedir.

Abstract

In this work a lexicon set and software framework, which will provide an infrastructure for Turkish text-tospeech synthesis applications, are introduced. The infrastructure consists of functions developed at Microsoft.NET 2.0 environment and XML based rule scripts. The infrastructure functions perform operations such as the correct phonetic representation of a given text; the conversions of acronyms and abbreviations or numerical values inside a text to appropriate written form for correct speech synthesis, etc. In scope of this study, a proprietary phonetic representation has also been defined, and it is described in details in this paper. It is aimed that the infrastructure will provide standardization for the future text-to-speech applications.