Aerospace Knowledge and Technology Journal

Aerospace Knowledge and Technology Journal

A comparative analysis of q-learning variants for real-time swarm robotic path planning in dynamic environments with static and mobile obstacles using monte carlo validation and an adaptive reward mechanism

Document Type : Research Paper

Authors
1 Assistant Professor, Aerospace Research Institute, Ministry of Science, Research and Technology, Tehran, Iran
2 Ph.D. Student, Aerospace Research Institute, Ministry of Science, Research and Technology, Tehran, Iran
Abstract
Real-time guidance of robotic swarms in dynamic environments with static and mobile obstacles is a complex challenge in autonomous robotic planning. This research presents an advanced reinforcement learning framework comparing Q-learning and Double Q-learning for cooperative guidance of aerial robot swarms, using a simulation with 10 aerial robots, 12 static obstacles, and 2 randomly moving mobile obstacles. Two motion structures were utilized: a 4-directional system (cardinal directions) and a 6-directional system (with diagonal movements). Innovative reward mechanisms, including progress reward (based on reduced Euclidean distance to target) and group success reward, were implemented. Results demonstrated that Double Q-learning with the 6-directional structure achieved the highest performance, with positional accuracy of 78.2% and average reward of 156.18. The Monte Carlo-based framework provides a reliable evaluation method for reinforcement learning algorithms in complex scenarios, applicable to autonomous delivery and search and rescue operations.
Keywords
Subjects

[1] P. Guo, R. Zhang, and B. Xu, "Safety Separation Distance Design for UAV Formation Based on System Performance," IEEE Transactions on Aerospace and Electronic Systems, pp. 1–16, 2025.
[2] E. C. Ozkat, "Vibration data-driven anomaly detection in UAVs: A deep learning approach," Engineering Science and Technology, an International Journal, vol. 54, p. 101702, 2024/06/01/ 2024.
[3] M. E. Elshaar, M. R. Elbalshy, A. Hussien, and M. Abido, "Path Planning in a dynamic environment using Spherical Particle Swarm Optimization," in 2024 IEEE Congress on Evolutionary Computation (CEC), pp. 1–8, July 2024.
[4] X. Olaz, D. Alaez, M. Prieto, J. Villadangos, and J. J. Astrain, "Quadcopter neural controller for take-off and landing in windy environments," Expert Systems with Applications, vol. 225, p. 120146, 2023.
[5] X. Ai, Y. Zhang, and Y.-Y. Chen, "Spherical Formation Tracking Control of Non-Holonomic UAVs with State Constraints and Time Delays," Aerospace, vol. 10, no. 2, 2023.
[6] S. Poudel, M. Y. Arafat, and S. Moh, "Bio-Inspired Optimization-Based Path Planning Algorithms in Unmanned Aerial Vehicles: A Survey," Sensors, vol. 23, no. 6, 2023.
[7] X. Wang, L. Wang, and Y. Hu, "Self-confirming Q-learning on unknown networks," Chaos, Solitons & Fractals, vol. 203, p. 117598, 2026.
[8] M. I. A. Shah, M. E. Cruz Victorio, M. Duffy, E. Barrett, and K. Mason, "Peer-to-peer energy trading in dairy farms using multi-agent reinforcement learning," Applied Energy, vol. 402, p. 127041, 2026.
[9] S. Lin, J. Wang, B. Huang, X. Kong, and H. Yang, "Bio particle swarm optimization and reinforcement learning algorithm for path planning of automated guided vehicles in dynamic industrial environments," Scientific Reports, vol. 15, no. 1, p. 463, 2025.
[10] C. Zhang, K. K. H. Ng, Z. Jin, S. Yao, and Y. Qin, "Q-learning-driven exact and meta-heuristic algorithms for the robust gate assignment problem," Advanced Engineering Informatics, vol. 67, 2025.
[11] B. Q. A. Nguyen, N. T. Dang, T. T. Le, and P. N. Dao, "On-policy and Off-policy Q-learning algorithms with policy iteration for two-wheeled inverted pendulum systems," Robotics and Autonomous Systems, p. 105111, 2025.
[12] C. Zhang, K. Ng, Z. Jin, S. Yao, and Y. Qin, "Q-learning-driven exact and meta-heuristic algorithms for the robust gate assignment problem," Advanced Engineering Informatics, vol. 67, p. 103551, 2025.
[13] C. Sun, L. Hou, S. Yu, and J. Shu, "HQA: Hybrid Q-Learning and AODV Multi-Path Routing Algorithm for Flying Ad-Hoc Networks," Vehicular Communications, p. 100947, 2025.
[14] H. Van Hasselt, A. Guez, and D. Silver, "Deep reinforcement learning with double q-learning," in Proceedings of the AAAI conference on artificial intelligence, vol. 30, no. 1, 2016.
[15] X. Guo, Q.-K. Pan, W. Zhang, Z.-H. Miao, X.-L. Jing, and H.-Y. Sang, "A Q-learning-assisted memetic algorithm for joint vehicle scheduling problem for harvesting and transportation in smart agriculture," Swarm and Evolutionary Computation, vol. 96, p. 102007, 2025.
[16] Z. Duan, Y. Zhang, Z. Shao, Z. Xu, and Z. Xiang, "Advanced robot path planning on rough terrain: A Q-learning-based multi-objective PSO algorithm," Applied Soft Computing, p. 113798, 2025.
[17] Z. Li, Y. Wang, Y. Han, K. Gao, and J. Li, "Q-Learning-Driven Accelerated Iterated Greedy Algorithm for Multi-Scenario Group Scheduling in Distributed Blocking Flowshops," Knowledge-Based Systems, vol. 317, p. 113424, 2025.
[18] A. Giuffrida, N. Basilico, and F. Amigoni, "An empirical evaluation of learning-based multi-agent path finding algorithms in warehouse environments," Robotics and Autonomous Systems, p. 105149, 2025.
[19] J. Gao, Y. Li, X. Shi, and K. Yan, "Customer satisfaction optimization in due-aware multi-agent path finding," Information Sciences, p. 122782, 2025.
[20] S. Venu and M. Gurusamy, "A Comprehensive Review of Path Planning Algorithms for Autonomous Navigation," Results in Engineering, p. 107750, 2025.
[21] S. Aslan and S. Demirci, "An immune plasma algorithm with Q-learning based pandemic management for path planning of unmanned aerial vehicles," Egyptian Informatics Journal, vol. 26, p. 100468, 2024.
[22] Y. Wang and G. Wang, "An adaptive and efficient path planning algorithm for UAV navigation in complex environments," Computers & Operations Research, vol. 185, p. 107296, 2026.
[23] X. Ni, W. Hu, Q. Fan, Y. Cui, and C. Qi, "A Q-learning based multi-strategy integrated artificial bee colony algorithm with application in unmanned vehicle path planning," Expert Systems with Applications, vol. 236, p. 121303, 2024.
[24] E. Mohammad, M. Jahromi, J. Pirkandi, M. Khazaee, and M. Mahmoodi, "Development of an intelligent jet engine controller using a model-based deep deterministic policy gradient technique,", Journal of Mechanical Engineering and Sciences, vol. 19, no. 3, pp. 10739 – 10755, 2025.