Optimizing UAV Path Following Using Dynamic Reward-Based Q-Learning


Hergul H. A., FIÇICI C., ÇATALBAŞ M. C.

11th International Conference on Recent Advances in Air and Space Technologies, Conference Program, RAST 2026, İstanbul, Türkiye, 13 - 15 Mayıs 2026, (Tam Metin Bildiri)

  • Yayın Türü: Bildiri / Tam Metin Bildiri
  • Doi Numarası: 10.1109/rast69551.2026.11672498
  • Basıldığı Şehir: İstanbul
  • Basıldığı Ülke: Türkiye
  • Anahtar Kelimeler: Obstacle Avoidance, Path Following, Q-Learning, Reinforcement Learning (RL), Unmanned Aerial Vehicle (UAV)
  • Ankara Üniversitesi Adresli: Evet

Özet

Unmanned Aerial Vehicles (UAVs) are increasingly used in environments where human access is limited and challenging due to their hardware and logistical advantages. Therefore, the system requires capable techniques to avoid obstacles and stay on a pre-determined route. Traditional path-following methods in variable and unfamiliar environments negatively impact the flight experience by introducing limitations such as static behavior, slow convergence rates, and low adaptability. In such scenarios, Reinforcement Learning (RL) is known for its ability to learn information by reacting with the environment without needing large pre-training datasets. In this study, a Grid-Based Q-Learning algorithm enhanced with a Dynamic Reward Mechanism is developed to optimize the course deviation measure and further improve obstacle avoidance capability. The UAV agent that is placed within a discretized three-dimensional (3D) simulation learns to navigate sequentially between intermediate points, adhering to the optimal route. Unlike standard techniques, the reward mechanism is used to support the agent staying close to the route without colliding with constantly moving both dynamic and static obstacles. Experimental results demonstrate that the algorithm guides the system to the final destination with 100% success rate after training, with minimized cross-track error and optimized path efficiency.