Journals / Havacılık ve Uzay Teknolojileri Dergisi / 2022 / Cilt: 15 - Sayı: 2
Deep Reinforcement Learning-based Cooperative Survivability Maximization for a UAV Fleet on an Air-to-Ground Mission
- Pages
- 94–107
- DOI
- —
Abstract
This study focuses on the cooperative strategy development of a UAV team that operates in a hostile environment in which the radar and weapon systems try to track and eliminate them. To simulate the hostile defense system, we present Markov models that generate the detecting and tracking probabilities of a radar system, and calculate the multiple-shot survivability of air vehicles that fly within the hostile environment. A cooperative strategy development procedure is presented based on proximal policy optimization algorithm, which is a deep reinforcement learning method. It is shown that the UAV team can develop cooperative strategies by exploiting enemy’s weakness to maximize team survivability in an air-to-ground mission after training with the proposed reinforcement learning scheme.
Özet
Çalışma kapsamında hasıma ait radar ve silah sistemlerinin bulunduğu bir ortamda operasyon yapan İHA takımının işbirlikçi strateji geliştirmesine odaklanılmıştır. Hasım savunma sisteminin benzetimini yapmak için Markov modelleri sunulmuştur. İlgili modeller radar sisteminin tespit ve takip olasılıklarının üretebilmekte ve hasım ortamında uçuş yapan hava araçlarının çoklu atış bekalarını hesaplayabilmektedir. Bir derin pekiştirmeli öğrenme metodu olan Proksimal Politika Algoritması vasıtasıyla bir işbirlikçi strateji geliştirme yöntemi sunulmuştur. Önerilen pekiştirmeli öğrenme yapısıyla eğitimin gerçekleştirilmesi ardılında İHA takımının rakibin zayıflığını kullanarak takım bekasını en iyilerken işbirlikçi strateji geliştirebildiği gösterilmiştir.