دانش و فناوری هوافضا

دانش و فناوری هوافضا

هدایت و کنترل خودمختار ملاقات مداری با شش درجه آزادی با استفاده از فرایادگیری تقویتی و شبکه‌های ترنسفورمر

نوع مقاله : مقاله پژوهشی

نویسندگان
1 دانشجو دکتری، مجتمع دانشگاهی برق و سایبرنتیک، دانشگاه صنعتی مالک‌اشتر
2 دانشیار، مجتمع دانشگاهی برق و سایبرنتیک، دانشگاه صنعتی مالک‌اشتر
چکیده
افزایش تعداد ماهواره‌ها در مدار پایین زمین، خطر برخورد بین اجسام فضایی را به‌شدت افزایش داده است. مأموریت‌های سرویس‌دهی و حذف زباله‌های فضایی می‌توانند با افزایش عمر ماهواره‌ها و پاکسازی مدارها، این تهدید را کاهش دهند. در این پژوهش، یک رویکرد نوین برای هدایت و کنترل فضاپیما در سناریوهای ملاقات مداری با شش درجه آزادی ارائه شده است که مبتنی بر یادگیری تقویتی فرا-آموزشی و شبکه‌های ترنسفورمر است. این مدل با کمک شبکه‌های ترنسفورمر، امکان یادگیری روابط پیچیده زمانی و استنباط اطلاعات پنهان از محیط را برای فضاپیمای تعقیب‌کننده فراهم می‌کند. الگوریتم بهینه‌سازی سیاست مجاورتی (PPO) که برای آموزش مدل به‌کار گرفته شده، در کنترل پیوسته عملکرد بالایی دارد. نتایج شبیه‌سازی‌ها در محیط مجازی نشان می‌دهند که این رویکرد از لحاظ دقت و پایداری بر معماری‌های سنتی مانند LSTM برتری دارد. از سوی دیگر تعداد پارامترهای شبکه، خود چالشی مهم در پیاده سازی بر روی سخت افزارها هست که روش پیشنهادی با کاهش محسوس در تعداد پارامترهای شبکه در کنار افزایش انطباق‌پذیری و بهبود دقت در شرایط متغیر محیطی کمک می کند. این رویکرد می‌تواند به‌عنوان راهکاری مؤثر برای تسهیل مأموریت‌های سرویس‌دهی و مدیریت زباله‌های فضایی مورد استفاده قرار گیرد.
کلیدواژه‌ها
موضوعات

عنوان مقاله English

Autonomous six-degree-of-freedom orbital rendezvous guidance and control using meta-reinforcement learning and transformer networks

نویسندگان English

Mehrdad Mohseni 1
Iman Mohammadzaman 2
1 PhD Student, Faculty of Electrical and Cybernetic Engineering, Malek Ashtar University of Technology, Iran
2 Associate Professor, Faculty of Electrical and Cybernetic Engineering, Malek Ashtar University of Technology, Iran.
چکیده English

The increasing number of satellites in low Earth orbit has significantly heightened the risk of collisions between space objects. Servicing and debris removal missions offer a viable solution by extending satellite lifespans and clearing orbital pathways. This research presents an innovative approach for spacecraft guidance and control in six degrees-of-freedom orbital rendezvous scenarios, employing meta-reinforcement learning and transformer networks. Leveraging transformer networks, this model enables the chaser spacecraft to learn complex temporal relationships and infer hidden information from the environment. The Proximal Policy Optimization (PPO) algorithm, utilized for model training, demonstrates superior performance in continuous control tasks. Simulation results in a virtual environment indicate that this approach outperforms traditional architectures like LSTM in terms of accuracy and stability. Additionally, network parameter count poses a significant challenge for hardware implementation; the proposed method addresses this by achieving substantial parameter reduction alongside enhanced adaptability and improved precision under varying environmental conditions. This approach could serve as an effective solution for facilitating future on-orbit servicing and space debris management missions.

کلیدواژه‌ها English

Orbital rendezvous
Meta-reinforcement learning
Transformer
On-orbit servicing
[1] A. Richards, T. Schouwenaars, J. P. How, and E. Feron, "Spacecraft Trajectory Planning with Avoidance Constraints Using Mixed-Integer Linear Programming," Journal of Guidance, Control, and Dynamics, vol. 25, no. 4, pp. 755-764, 2002/07/01 2002, doi: 10.2514/2.4943.
 [2]S. Boyd and L. Vandenberghe, Convex Optimization. Cambridge: Cambridge University Press, 2004.
 [3]E. Hartley, A tutorial on model predictive control for spacecraft rendezvous. 2015, pp. 1355-1361.
[4] D. Woffinden and D. Geller, "Optimal Orbital Rendezvous Maneuvering for Angles-Only Navigation," Journal of Guidance Control and Dynamics - J GUID CONTROL DYNAM, vol. 32, pp. 1382-1387, 07/01 2009, doi: 10.2514/1.45006.
[5] L. Blackmore and D. Scharf, "Minimum-Landing-Error Powered-Descent Guidance for Mars Landing Using Convex Optimization," Journal of Guidance Control and Dynamics - J GUID CONTROL DYNAM, vol. 33, pp. 1161-1171, 07/01 2010, doi: 10.2514/1.47202.
 [6]L. Blackmore, "Autonomous precision landing of space rockets," vol. 46, pp. 15-20, 01/01 2016.
[7] U. Eren, A. Prach, B. B. Koçer, S. V. Raković, E. Kayacan, and B. Açıkmeşe, "Model Predictive Control in Aerospace Systems: Current State and Opportunities," Journal of Guidance, Control, and Dynamics, vol. 40, no. 7, pp. 1541-1566, 2017/07/01 2017, doi: 10.2514/1.G002507.
[8] B. Benedikter, A. Zavoli, G. Colasurdo, S. Pizzurro, and E. Cavallini, "Autonomous Upper Stage Guidance Using Convex Optimization and Model Predictive Control," in ASCEND 2020, (ASCEND: American Institute of Aeronautics and Astronautics, 2020.
[9] D. Miller, J. A. Englander, and R. Linares, "Interplanetary low-thrust design using proximal policy optimization," in 2019 AAS/AIAA Astrodynamics Specialist Conference, 2019, no. GSFC-E-DAA-TN71225.
[10] A. Zavoli and L. Federici, "Reinforcement Learning for Robust Trajectory Design of Interplanetary Missions," Journal of Guidance, Control, and Dynamics, vol. 44, no. 8, pp. 1440-1453, 2021, doi: 10.2514/1.G005794.
[11] C. J. Sullivan and N. Bosanac, "Using Reinforcement Learning to Design a Low-Thrust Approach into a Periodic Orbit in a Multi-Body System," in AIAA Scitech 2020 Forum, (AIAA SciTech Forum: American Institute of Aeronautics and Astronautics, 2020.
 [12]N. B. LaFarge, D. Miller, K. C. Howell, and R. Linares, "Guidance for Closed-Loop Transfers using Reinforcement Learning with Application to Libration Point Orbits," in AIAA Scitech 2020 Forum, (AIAA SciTech Forum: American Institute of Aeronautics and Astronautics, 2020.
[13] L. Federici, A. Scorsoglio, A. Zavoli, and R. Furfaro, "Autonomous guidance for cislunar orbit transfers via reinforcement learning," in AAS/AIAA Astrodynamics Specialist Conference, 2021: American Astronautical Society Big Sky, Montana (Virtual).
[14] R. Furfaro, A. Scorsoglio, R. Linares, and M. Massari, "Adaptive generalized ZEM-ZEV feedback guidance for planetary landing via a deep reinforcement learning approach," Acta Astronautica, vol. 171, pp. 156-171, 2020/06/01/ 2020, doi: https://doi.org/10.1016/j.actaastro.2020.02.051.
[15] B. Gaudet, R. Linares, and R. Furfaro, "Deep reinforcement learning for six degree-of-freedom planetary powered descent and landing," arXiv preprint arXiv:1810.08719, 2018.
[16] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, "Proximal policy optimization algorithms," arXiv preprint arXiv:1707.06347, 2017.
[17] T. P. Lillicrap et al., "Continuous control with deep reinforcement learning," arXiv e-prints, p. arXiv:1509.02971, 2015, doi: 10.48550/arXiv.1509.02971.
[18] C. E. Oestreich, R. Linares, and R. Gondhalekar, "Autonomous six-degree-of-freedom spacecraft docking maneuvers via reinforcement learning," arXiv preprint arXiv:2008.03215, 2020.
[19] L. Federici, A. Scorsoglio, A. Zavoli, and R. Furfaro, "Meta-reinforcement learning for adaptive spacecraft guidance during finite-thrust rendezvous missions," Acta Astronautica, vol. 201, pp. 129-141, 2022/12/01/ 2022, doi: https://doi.org/10.1016/j.actaastro.2022.08.047.
[20] N. Mishra, M. Rohaninejad, X. Chen, and P. Abbeel, "A simple neural attentive meta-learner," arXiv preprint arXiv:1707.03141, 2017.
 [21]L. C. Melo, "Transformers are meta-reinforcement learners," in international conference on machine learning, 2022: PMLR, pp. 15340-15359.
[22] J. Ventura, M. Ciarcià, M. Romano, and U. Walter, "Fast and Near-Optimal Guidance for Docking to Uncontrolled Spacecraft," Journal of Guidance, Control, and Dynamics, vol. 40, no. 12, pp. 3138-3154, 2017/12/01 2016, doi: 10.2514/1.G001843.
[23] G. Boyarko, O. Yakimenko, and M. Romano, "Optimal Rendezvous Trajectories of a Controlled Spacecraft and a Tumbling Object," Journal of Guidance, Control, and Dynamics, vol. 34, no. 4, pp. 1239-1252, 2011/07/01 2011, doi: 10.2514/1.47645.
 [24]K. J. Åström, "Optimal control of Markov processes with incomplete state information," Journal of Mathematical Analysis and Applications, vol. 10, pp. 174-205, 1965.
[25] J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz, "Trust region policy optimization," in International conference on machine learning, 2015: PMLR, pp. 1889-1897.
[26] T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, "Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor," presented at the Proceedings of the 35th International Conference on Machine Learning, Proceedings of Machine Learning Research, 2018. [Online]. Available: https://proceedings.mlr.press/v80/haarnoja18b.html.
 [27]S. Silvestrini et al., "Chapter Fifteen - Modern Spacecraft GNC," in Modern Spacecraft Guidance, Navigation, and Control, V. Pesce, A. Colagrossi, and S. Silvestrini Eds.: Elsevier, 2023, pp. 819-981.
[28] A. Brandonisio and M. Lavagna, Sensitivity Analysis of Adaptive Guidance via Deep Reinforcement Learning for Uncooperative Space Objects Smart Imaging. 2021.
[29] L. Capra, A. Brandonisio, and M. Lavagna, "Network architecture and action space analysis for deep reinforcement learning towards spacecraft autonomous guidance," Advances in Space Research, vol. 71, 12/01 2022, doi: 10.1016/j.asr.2022.11.048.
[30] T. Wolf et al., Transformers: State-of-the-Art Natural Language Processing. 2020, pp. 38-45.
 [31]J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding," in North American Chapter of the Association for Computational Linguistics, 2019.
 [32]H. Chen et al., "Pre-Trained Image Processing Transformer," in 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 20-25 June 2021 2021, pp. 12294-12305, doi: 10.1109/CVPR46437.2021.01212.
[33] A. Dosovitskiy, "An image is worth 16x16 words: Transformers for image recognition at scale," arXiv preprint arXiv:2010.11929, 2020.
[34] A. Vaswani et al., "Attention Is All You Need," 06/12 2017, doi: 10.48550/arXiv.1706.03762.
[35] M.-T. Luong, "Effective approaches to attention-based neural machine translation," arXiv preprint arXiv:1508.04025, 2015.
[36] S. Hochreiter, "Long Short-term Memory," Neural Computation MIT-Press, 1997.
 [37]L. Federici, A. Scorsoglio, A. Zavoli, and R. Furfaro, "Meta-Reinforcement Learning for Adaptive Spacecraft Guidance during Multi-Target Missions," in IAF Astrodynamics Symposium 2021 at the 72nd International Astronautical Congress, IAC 2021, 2021: International Astronautical Federation, IAF.
[38] S.-B. (SB3). "Reinforcement Learning Tips and Tricks." https://stable-baselines3.readthedocs.io/en/master/guide/rl_tips.html (accessed.
[39] B. Gaudet, R. Linares, and R. Furfaro, "Deep reinforcement learning for six degree-of-freedom planetary landing," Advances in Space Research, vol. 65, no. 7, pp. 1723-1741, 2020/04/01/ 2020, doi: https://doi.org/10.1016/j.asr.2019.12.030.