A preliminary approach for the asynchronous deployment of reinforcement learning software on physical systems.
DOI:
https://doi.org/10.17979/ja-cea.2026.47.13751Keywords:
Robots manipulators, Reinforcement learning, Asynchronous deployment, Variable timestep, Distributed Software ArchitectureAbstract
Reinforcement Learning (RL) usually assume fixed timesteps, ignoring temporal discrepancies between simulation and reality and the effects of system latencies. These factors can degrade performance in sim-to-real transfer. This work introduces a fully decoupled solution enabling asynchronous, algorithm-agnostic, and time-aware RL deployment. Our approach isolates the algorithm from the agent into independent loops, overcoming the limitations of traditional synchronous frameworks. This facilitates integration with physical robots and closed-source simulators, enabling variable-time RL strategies. Our tool is compatible with libraries like Stable-Baselines3, simulators such as MuJoCo, and robots like the Franka Emika Panda. A preliminary experimental validation has demonstrated its effectiveness with such robotics manipulator, and an initial version of the software has been made available to the community on GitHub: https://github.com/uncore-team/rl_spin_decoupler.
References
Al Mahmud, S., Kamarulariffin, A., Ibrahim, A. M., Mohideen, A. J. H., 2024. Advancements and challenges in mobile robot navigation: A comprehensive review of algorithms and potential for self-learning approaches. Journal of Intelligent & Robotic Systems 110 (3), 120. URL: https://doi.org/10.1007/s10846-024-02149-5 DOI: 10.1007/s10846-024-02149-5
Dulac-Arnold, G., Mankowitz, D. J., Hester, T., 2019. Challenges of real-world reinforcement learning. ArXiv abs/1904.12901.
Elguea-Aguinaco, Í., Serrano-Muñoz, A., Chrysostomou, D., Inziarte-Hidalgo, I., Bøgh, S., Arana-Arexolaleiba, N., 2023. A review on reinforcement learning for contact-rich robotic manipulation tasks. Robotics and Computer-Integrated Manufacturing 81, 102517.
Elsner, J., 2023. Taming the panda with python: A powerful duo for seamless robotics programming and integration. SoftwareX 24, 101532. URL: https://www.sciencedirect.com/science/article/pii/S2352711023002285 DOI: 10.1016/j.softx.2023.101532
Fernández-Madrigal, J.-A., Navarro, A., Asenjo, R., Cruz-Martín, A., 2020. Characterization, statistical analysis and method selection in the two-clocks synchronization problem for pairwise interconnected sensors. Sensors 20 (17). URL: https://www.mdpi.com/1424-8220/20/17/4808 DOI: 10.3390/s20174808
Gamma, E., Helm, R., Johnson, R., Vlissides, J., 1994. Design Patterns: Elements of Reusable Object-Oriented Software. Addison-Wesley, Reading, MA.
Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., Abbeel, P., Levine, S., 2019. Soft actor-critic algorithms and applications. URL: https://arxiv.org/abs/1812.05905
Haddadin, S., 2024. The franka emika robot: A standard platform in robotics research. IEEE Robotics & Automation Magazine.
Ibarz, J., Tan, J., Finn, C., Kalakrishnan, M., Pastor, P., Levine, S., 2021. How to train your robot with deep reinforcement learning: lessons we have learned. The International Journal of Robotics Research 40, 698 – 721.
IMECH.UMA, 2025. Instituto Universitario de Investigación en Ingeniería Mecatrónica y Sistemas Ciberfísicos, Universidad de Málaga, España. https://www.imech-uma.es/, accedido el 14 de julio de 2025.
Le, H., Saeedvand, S., Hsu, C.-C., 2024. A comprehensive review of mobile robot navigation using deep reinforcement learning algorithms in crowded environments. Journal of Intelligent & Robotic Systems 110 (4), 158. URL: https://doi.org/10.1007/s10846-024-02198-w DOI: 10.1007/s10846-024-02198-w
Mahmood, A. R., Korenkevych, D., Komer, B., Bergstra, J., 2018. Setting up a reinforcement learning task with a real-world robot. 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 4635–4640. URL: https://api.semanticscholar.org/CorpusID:3971262
Postel, J., 1981. Transmission control protocol. RFC 793, accessed: 2025-05-22. URL: https://www.rfc-editor.org/rfc/rfc793
Prasuna, R. G., Potturu, S. R., 2024. Deep reinforcement learning in mobile robotics – a concise review. Multimedia Tools and Applications 83 (28), 70815–70836.
Puterman, M. L., 2014. Markov Decision Processes: Discrete Stochastic Dynamic Programming, 1st Edition. John Wiley & Sons.
Raffin, A., Hill, A., Gleave, A., Kanervisto, A., Ernestus, M., Dormann, N., 2021. Stable-baselines3: Reliable reinforcement learning implementations. Journal of Machine Learning Research 22 (268), 1–8. URL: http://jmlr.org/papers/v22/20-1364.html
Robotics, C., 2024. Coppeliasim. https://www.coppeliarobotics.com/, accessed: 2025-05-08.
Salvato, E., Fenu, G., Medvet, E., Pellegrino, F. A., 2021. Crossing the reality gap: a survey on sim-to-real transferability of robot controllers in reinforcement learning. IEEE Access PP, 1–1. URL: https://api.semanticscholar.org/CorpusID:243882665
Todorov, E., Erez, T., Tassa, Y., 2012. Mujoco: A physics engine for model-based control. In: IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). pp. 5026–5033. DOI: 10.1109/IROS.2012.6386109
Towers, M., Kwiatkowski, A., Terry, J., Balis, J. U., De Cola, G., Deleu, T., Goulão, M., Kallinteris, A., Krimmel, M., KG, A., et al., 2024. Gymnasium: A standard interface for reinforcement learning environments. arXiv preprint arXiv:2407.17032.
Van Rossum, G., Drake, F. L., 2009. Python 3 Reference Manual. CreateSpace, Scotts Valley, CA.
Yuan, Y., Mahmood, A. R., 2022. Asynchronous reinforcement learning for real-time control of physical robots. URL: https://arxiv.org/abs/2203.12759
Zhao, W., Queralta, J. P., Westerlund, T., 2020. Sim-to-real transfer in deep reinforcement learning for robotics: a survey. 2020 IEEE Symposium Series on Computational Intelligence (SSCI), 737–744. URL: https://api.semanticscholar.org/CorpusID:221971078
Zhu, K., Zhang, T., 2021. Deep reinforcement learning based mobile robot navigation: A review. Tsinghua Science and Technology 26 (5), 674–691. DOI: 10.26599/TST.2021.9010012
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Adrián Bañuls Arias, Diego Caruana-Montes, Ana Cruz-Martín, Vicente Arévalo-Espejo, Cipriano Galindo, Juan-Manuel Gandarias

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.