Dergiler / Bilgi Yönetimi / 2020 / Cilt: 3 - Sayı: 1
RapidMiner ile Twitter Verilerinin Konu Modellemesi
- Dergi
- Bilgi Yönetimi
- Sayfa
- 1–10
- DOI
- —
Özet
Bu çalışmada öncelikle RapidMiner kullanılarak Twitter’da belirli kelimeleriiçeren tweet verileri elde edildi, bu veriler ön işlemden geçirildi ve sonrasındatweetlerin konu modellemesi yapıldı. Ön işleme için “Search Twitter”, “SelectAttributes”, “Nominal to Text” blokları kullanıldı. Ön işlemden geçen Twitterverileri “Tokenize”, “Aggregate” ve “Discretize” operatörleri kullanılarak analizedildi. Tweetlerde en çok kullanılan kelimeler belirlendi ve kullanım sıklığınagöre kelime grupları oluşturuldu. Daha sonra Twitter verilerine nasıl konu bazlıkümeleme yapılacağı anlatıldı. Bu işlem için Latent Dirichlet Allocationmodelini kullanan “Extract Topics From Documents (LDA)” operatörükullanıldı. Tweetlerde en fazla kullanılan kelimeler ve kullanıcı başına atılantweet sayıları, grafik ve tablolarla incelendi, ayrıca konu modellemesi sonucundaelde edilen konuların kelime bulutu oluşturuldu.
Abstract
In this study, firstly, tweets containing specific words on the Twitter platform were obtained and pre-processed using the RapidMiner software. After that, the tweets are clustered based on the topic modeling approach. “Search Twitter”, “Select Attributes”, and “Nominal to Text” blocks were used for preprocessing. This preprocessed data is then analyzed using “Tokenize”, “Aggregate”, and “Discretize” operators. The most used words were determined, and tweets are grouped according to their frequencies. Then, it is explained how to perform topic-based modeling and clustering on Twitter data. “Extract Topics From Documents (LDA)” operator, which uses the Latent Dirichlet Allocation model, was used for this process. The most commonly used words in tweets, and the number of tweets per user were extracted and investigated via tables and graphical illustrations. In addition, the word cloud of each topic, obtained as a result of the topic modeling process, was created.