Dergiler / Fırat Üniversitesi Mühendislik Bilimleri Dergisi / 2022 / Cilt: 34 - Sayı: 2

Görüntülerden Derin Öğrenmeye Dayalı Otomatik Metin Çıkarma Bir Görüntü Yakalama Sistemi

Automatic Text Extraction Based on Deep Learning from Images: An Image Capture System

Sayfa
829–837
DOI
—

Özet

Bilgisayarlı görme ve doğal dil işlemenin çalışma alanlarından biri olan görüntü yakalama (image capturing), doğal bir dil kullanarak görüntü içeriğini otomatik olarak tanımlama görevidir. Bu çalışmada, MS COCO veri seti üzerinde İngilizce dili için encoder-decoder tekniğine dayalı bir otomatik altyazı oluşturma yaklaşımı önerilmiştir. Önerilen yaklaşımda, görüntü özniteliklerini çıkarmak için encoder olarak Evrişimli Sinir Ağı (CNN) mimarisi ve görüntülerden altyazı oluşturmak için bir decoder olarak Tekrarlayan Sinir Ağı (RNN) mimarisi kullanılmıştır. Önerilen yaklaşımın performansı BLEU, METEOR ve ROUGE_L değerlendirme kriterleri kullanılarak değerlendirilmiş ve her bir görüntüden 5 cümle elde edilmiştir. Deneysel sonuçlar, modelin görüntülerdeki nesneleri doğru bir şekilde algılamada tatmin edici olduğunu göstermektedir.

Abstract

Image capturing, one of the working areas of computer vision and natural language processing, is the task of automatically identifying image content using a natural language. In this study, an encoder-decoder technique-based automatic captioning approach is proposed for the English language on the MS COCO dataset. In the proposed approach, Convolutional Neural Network (CNN) architecture is used as an encoder for extracting image features, and Recurrent Neural Network (RNN) architecture is used as a decoder for creating subtitles from images. The performance of the proposed approach was evaluated using BLEU, METEOR, and ROUGE_L evaluation criteria, and 5 sentences were obtained from each image. The experiment results show that the model is satisfactory in correctly detecting the objects in the images.