دانش و فناوری هوافضا

دانش و فناوری هوافضا

تحلیل یادگیری تقویتی Q برای مسیریابی گروهی بلادرنگ در محیط پویا با اعتبارسنجی مونت‌کارلو و پاداش تطبیقی

نوع مقاله : مقاله پژوهشی

نویسندگان
1 استادیار پژوهشگاه هوافضا، وزارت علوم، تحقیقات و فناوری، تهران، ایران
2 دانشجوی دکتری، پژوهشگاه هوافضا، وزارت علوم، تحقیقات و فناوری، تهران، ایران
چکیده
هدایت بلادرنگ گروه‌های رباتیک در محیط‌های پویا دارای موانع ثابت و متحرک، یکی از چالش‌های پیچیده در زمینه برنامه ریزی رباتیک خودران است. این پژوهش یک چارچوب پیشرفته یادگیری تقویتی ارائه می‌دهد که عملکرد دو الگوریتم یادگیری Q و یادگیری Q دوگانه را در هدایت گروهی ربات های هوایی مقایسه و بهبود می دهد. مطالعه حاضر از یک شبیه ‌سازی شامل ۱۰ ربات هوایی، 12 مانع ثابت با پراکندگی هوشمند و ۲ مانع متحرک با الگوی حرکت تصادفی است. دو ساختار حرکتی سیستم ۴ حرکتی و سیستم ۶ حرکتی (حرکت‌های مورب) بهره برده است. ساز و کارهای پاداش نوآورانه شامل پاداش پیشرفت (بر پایه کاهش فاصله اقلیدسی از هدف) و پاداش موفقیت گروهی پیاده‌ سازی شد. نتایج نشان داد یادگیری Q دوگانه با ساختار ۶ حرکتی بالاترین عملکرد را با دقت موقعیتی 78.2 درصد و میانگین پاداش 156.18 ارائه داد. چارچوب مبتنی بر روش مونت‌کارلو ارائه ‌شده، روشی قابل‌اطمینان برای ارزیابی الگوریتم‌های یادگیری تقویتی در سناریوهای پیچیده فراهم می‌کند. این چارچوب می‌تواند در کاربردهایی مانند تحویل خودران و عملیات جستجو و نجات به کار رود.
کلیدواژه‌ها
موضوعات

عنوان مقاله English

A comparative analysis of q-learning variants for real-time swarm robotic path planning in dynamic environments with static and mobile obstacles using monte carlo validation and an adaptive reward mechanism

نویسندگان English

Iman Shafieenejad 1
Faezeh Shokri 2
1 Assistant Professor, Aerospace Research Institute, Ministry of Science, Research and Technology, Tehran, Iran
2 Ph.D. Student, Aerospace Research Institute, Ministry of Science, Research and Technology, Tehran, Iran
چکیده English

Real-time guidance of robotic swarms in dynamic environments with static and mobile obstacles is a complex challenge in autonomous robotic planning. This research presents an advanced reinforcement learning framework comparing Q-learning and Double Q-learning for cooperative guidance of aerial robot swarms, using a simulation with 10 aerial robots, 12 static obstacles, and 2 randomly moving mobile obstacles. Two motion structures were utilized: a 4-directional system (cardinal directions) and a 6-directional system (with diagonal movements). Innovative reward mechanisms, including progress reward (based on reduced Euclidean distance to target) and group success reward, were implemented. Results demonstrated that Double Q-learning with the 6-directional structure achieved the highest performance, with positional accuracy of 78.2% and average reward of 156.18. The Monte Carlo-based framework provides a reliable evaluation method for reinforcement learning algorithms in complex scenarios, applicable to autonomous delivery and search and rescue operations.

کلیدواژه‌ها English

Real-Time Path Planning
Swarm Robotics
Double Q-Learning
Mobile Obstacles
Monte Carlo Simulation
[1] P. Guo, R. Zhang, and B. Xu, "Safety Separation Distance Design for UAV Formation Based on System Performance," IEEE Transactions on Aerospace and Electronic Systems, pp. 1–16, 2025.
[2] E. C. Ozkat, "Vibration data-driven anomaly detection in UAVs: A deep learning approach," Engineering Science and Technology, an International Journal, vol. 54, p. 101702, 2024/06/01/ 2024.
[3] M. E. Elshaar, M. R. Elbalshy, A. Hussien, and M. Abido, "Path Planning in a dynamic environment using Spherical Particle Swarm Optimization," in 2024 IEEE Congress on Evolutionary Computation (CEC), pp. 1–8, July 2024.
[4] X. Olaz, D. Alaez, M. Prieto, J. Villadangos, and J. J. Astrain, "Quadcopter neural controller for take-off and landing in windy environments," Expert Systems with Applications, vol. 225, p. 120146, 2023.
[5] X. Ai, Y. Zhang, and Y.-Y. Chen, "Spherical Formation Tracking Control of Non-Holonomic UAVs with State Constraints and Time Delays," Aerospace, vol. 10, no. 2, 2023.
[6] S. Poudel, M. Y. Arafat, and S. Moh, "Bio-Inspired Optimization-Based Path Planning Algorithms in Unmanned Aerial Vehicles: A Survey," Sensors, vol. 23, no. 6, 2023.
[7] X. Wang, L. Wang, and Y. Hu, "Self-confirming Q-learning on unknown networks," Chaos, Solitons & Fractals, vol. 203, p. 117598, 2026.
[8] M. I. A. Shah, M. E. Cruz Victorio, M. Duffy, E. Barrett, and K. Mason, "Peer-to-peer energy trading in dairy farms using multi-agent reinforcement learning," Applied Energy, vol. 402, p. 127041, 2026.
[9] S. Lin, J. Wang, B. Huang, X. Kong, and H. Yang, "Bio particle swarm optimization and reinforcement learning algorithm for path planning of automated guided vehicles in dynamic industrial environments," Scientific Reports, vol. 15, no. 1, p. 463, 2025.
[10] C. Zhang, K. K. H. Ng, Z. Jin, S. Yao, and Y. Qin, "Q-learning-driven exact and meta-heuristic algorithms for the robust gate assignment problem," Advanced Engineering Informatics, vol. 67, 2025.
[11] B. Q. A. Nguyen, N. T. Dang, T. T. Le, and P. N. Dao, "On-policy and Off-policy Q-learning algorithms with policy iteration for two-wheeled inverted pendulum systems," Robotics and Autonomous Systems, p. 105111, 2025.
[12] C. Zhang, K. Ng, Z. Jin, S. Yao, and Y. Qin, "Q-learning-driven exact and meta-heuristic algorithms for the robust gate assignment problem," Advanced Engineering Informatics, vol. 67, p. 103551, 2025.
[13] C. Sun, L. Hou, S. Yu, and J. Shu, "HQA: Hybrid Q-Learning and AODV Multi-Path Routing Algorithm for Flying Ad-Hoc Networks," Vehicular Communications, p. 100947, 2025.
[14] H. Van Hasselt, A. Guez, and D. Silver, "Deep reinforcement learning with double q-learning," in Proceedings of the AAAI conference on artificial intelligence, vol. 30, no. 1, 2016.
[15] X. Guo, Q.-K. Pan, W. Zhang, Z.-H. Miao, X.-L. Jing, and H.-Y. Sang, "A Q-learning-assisted memetic algorithm for joint vehicle scheduling problem for harvesting and transportation in smart agriculture," Swarm and Evolutionary Computation, vol. 96, p. 102007, 2025.
[16] Z. Duan, Y. Zhang, Z. Shao, Z. Xu, and Z. Xiang, "Advanced robot path planning on rough terrain: A Q-learning-based multi-objective PSO algorithm," Applied Soft Computing, p. 113798, 2025.
[17] Z. Li, Y. Wang, Y. Han, K. Gao, and J. Li, "Q-Learning-Driven Accelerated Iterated Greedy Algorithm for Multi-Scenario Group Scheduling in Distributed Blocking Flowshops," Knowledge-Based Systems, vol. 317, p. 113424, 2025.
[18] A. Giuffrida, N. Basilico, and F. Amigoni, "An empirical evaluation of learning-based multi-agent path finding algorithms in warehouse environments," Robotics and Autonomous Systems, p. 105149, 2025.
[19] J. Gao, Y. Li, X. Shi, and K. Yan, "Customer satisfaction optimization in due-aware multi-agent path finding," Information Sciences, p. 122782, 2025.
[20] S. Venu and M. Gurusamy, "A Comprehensive Review of Path Planning Algorithms for Autonomous Navigation," Results in Engineering, p. 107750, 2025.
[21] S. Aslan and S. Demirci, "An immune plasma algorithm with Q-learning based pandemic management for path planning of unmanned aerial vehicles," Egyptian Informatics Journal, vol. 26, p. 100468, 2024.
[22] Y. Wang and G. Wang, "An adaptive and efficient path planning algorithm for UAV navigation in complex environments," Computers & Operations Research, vol. 185, p. 107296, 2026.
[23] X. Ni, W. Hu, Q. Fan, Y. Cui, and C. Qi, "A Q-learning based multi-strategy integrated artificial bee colony algorithm with application in unmanned vehicle path planning," Expert Systems with Applications, vol. 236, p. 121303, 2024.
[24] E. Mohammad, M. Jahromi, J. Pirkandi, M. Khazaee, and M. Mahmoodi, "Development of an intelligent jet engine controller using a model-based deep deterministic policy gradient technique,", Journal of Mechanical Engineering and Sciences, vol. 19, no. 3, pp. 10739 – 10755, 2025.