{"kagg_id":"KAGG:2607.00001","slug":"2607.00001","title":"Cross-Embodiment Transfer in Vision-Based Robot Navigation: Challenges, Methods, and Open Problems","authors":[{"name":"Yiliang Zhang","affiliation":"Eastern Institute of Technology, Ningbo, Zhejiang, China"}],"affiliation":null,"abstract":"Cross-embodiment transfer in visual navigation is often described as a domain-generalization problem, but this framing is incomplete: changing the robot simultaneously changes the camera viewpoint, collision geometry, motion constraints, action semantics, and the consequences of perception error. This survey argues that reliable transfer therefore requires a factorized navigation stack with four explicit components: an environment-centric representation, an embodiment descriptor, a feasibility layer that converts scene understanding into body-aware constraints, and a deployment layer that handles dynamics and sim-to-real mismatch. We organize recent work into three complementary families: heterogeneous-data generalist policies, embodiment-randomized or embodiment-conditioned policies, and geometric abstraction with configuration-driven planning. Representative systems include GNM, ViNT, NoMaD, RING, NavFoM, EA-Nav, CeRLP, and AgniNav. We then connect these methods to evidence from sim-to-real navigation research, showing why simulator fidelity alone is insufficient and why collision behavior, action interfaces, and evaluation design strongly affect real-world conclusions. Finally, we propose an evaluation protocol that separates environment, sensor, geometry, dynamics, and reality gaps, and identify open problems in online embodiment estimation, adaptive safety margins, heterogeneous sensing, formally constrained learning, and cross-embodiment foundation models.","publication_date":"2026-07-15","registered_at":"2026-07-29T14:45:00Z","updated_at":"2026-07-29T17:45:08Z","status":"published","html_path":"data/articles/2607.00001/index.html","pdf_path":null,"source_type":"preprint","source_label":"Kaggleyes Academy Preprint","license":"CC BY 4.0","has_doi":false,"external_doi":null,"references":[{"text":"D. Shah, A. Sridhar, A. Bhorkar, N. Hirose, and S. Levine. GNM: A General Navigation Model to Drive Any Robot. arXiv:2210.03370, 2022.","identifier":"arXiv:2210.03370","url":"https://arxiv.org/abs/2210.03370"},{"text":"D. Shah, A. Sridhar, N. Dashora, K. Stachowicz, K. Black, N. Hirose, and S. Levine. ViNT: A Foundation Model for Visual Navigation. arXiv:2306.14846, 2023.","identifier":"arXiv:2306.14846","url":"https://arxiv.org/abs/2306.14846"},{"text":"A. Sridhar, D. Shah, C. Glossop, and S. Levine. NoMaD: Goal Masked Diffusion Policies for Navigation and Exploration. arXiv:2310.07896, 2023.","identifier":"arXiv:2310.07896","url":"https://arxiv.org/abs/2310.07896"},{"text":"A. Eftekhar et al. The One RING: a Robotic Indoor Navigation Generalist. arXiv:2412.14401, 2024.","identifier":"arXiv:2412.14401","url":"https://arxiv.org/abs/2412.14401"},{"text":"J. Zhang et al. Embodied Navigation Foundation Model. arXiv:2509.12129, 2025.","identifier":"arXiv:2509.12129","url":"https://arxiv.org/abs/2509.12129"},{"text":"J. Zhang, Y. Du, X. Guo, S. Sun, X. Liu, Y. Sun, G. Lu, W. Sui, and J. Li. EA-Nav: Learning Safe Visual Navigation Policies with Embodiment Awareness. arXiv:2607.19880, 2026.","identifier":"arXiv:2607.19880","url":"https://arxiv.org/abs/2607.19880"},{"text":"H. Xi, M. Tan, X. Zhang, S. Cheng, S. Wang, Y. Gu, X. Shen, and W. Zhang. CeRLP: A Cross-embodiment Robot Local Planning Framework for Visual Navigation. arXiv:2603.19602, 2026.","identifier":"arXiv:2603.19602","url":"https://arxiv.org/abs/2603.19602"},{"text":"T. Zang, S. Cheng, H. Huang, S. Wang, and W. Zhang. AgniNav: Configuration-Driven Cross-Embodiment Local Planning for Robot Navigation. arXiv:2606.10903, 2026.","identifier":"arXiv:2606.10903","url":"https://arxiv.org/abs/2606.10903"},{"text":"M. Savva et al. Habitat: A Platform for Embodied AI Research. arXiv:1904.01201, 2019.","identifier":"arXiv:1904.01201","url":"https://arxiv.org/abs/1904.01201"},{"text":"F. Xia, A. Zamir, Z.-Y. He, A. Sax, J. Malik, and S. Savarese. Gibson Env: Real-World Perception for Embodied Agents. arXiv:1808.10654, 2018.","identifier":"arXiv:1808.10654","url":"https://arxiv.org/abs/1808.10654"},{"text":"A. Kadian et al. Sim2Real Predictivity: Does Evaluation in Simulation Predict Real-World Performance? arXiv:1912.06321, 2019.","identifier":"arXiv:1912.06321","url":"https://arxiv.org/abs/1912.06321"},{"text":"P. Anderson, A. Shrivastava, J. Truong, A. Majumdar, D. Parikh, D. Batra, and S. Lee. Sim-to-Real Transfer for Vision-and-Language Navigation. Proceedings of Machine Learning Research, 155:671-681, 2021.","identifier":"PMLR 155","url":"https://proceedings.mlr.press/v155/anderson21a.html"},{"text":"J. Truong, M. Rudolph, N. H. Yokoyama, S. Chernova, D. Batra, and A. Rai. Rethinking Sim2Real: Lower Fidelity Simulation Leads to Higher Sim2Real Transfer in Navigation. Proceedings of Machine Learning Research, 205:859-870, 2023.","identifier":"PMLR 205","url":"https://proceedings.mlr.press/v205/truong23a.html"},{"text":"Z. Wang, X. Li, J. Yang, Y. Liu, and S. Jiang. Sim-to-Real Transfer via 3D Feature Fields for Vision-and-Language Navigation. Proceedings of Machine Learning Research, 270:2982-2995, 2025.","identifier":"PMLR 270","url":"https://proceedings.mlr.press/v270/wang25i.html"},{"text":"P. Anderson et al. On Evaluation of Embodied Navigation Agents. arXiv:1807.06757, 2018.","identifier":"arXiv:1807.06757","url":"https://arxiv.org/abs/1807.06757"},{"text":"N. Yokoyama, S. Ha, and D. Batra. Success Weighted by Completion Time: A Dynamics-Aware Evaluation Criteria for Embodied Navigation. arXiv:2103.08022, 2021.","identifier":"arXiv:2103.08022","url":"https://arxiv.org/abs/2103.08022"},{"text":"J. Hu, S. Yuan, G. C. R. Bethala, A. Tzes, and Y. Fang. Learning Adaptive Safety Margins for Visual Navigation. arXiv:2607.18200, 2026.","identifier":"arXiv:2607.18200","url":"https://arxiv.org/abs/2607.18200"}],"canonical_url":"https://academy.kaggleyes.top/abs/KAGG:2607.00001","metadata":{"title":"Cross-Embodiment Transfer in Vision-Based Robot Navigation: Challenges, Methods, and Open Problems","authors":[{"name":"Yiliang Zhang","affiliation":"Eastern Institute of Technology, Ningbo, Zhejiang, China"}],"publication_date":"2026-07-15","abstract":"Cross-embodiment transfer in visual navigation is often described as a domain-generalization problem, but this framing is incomplete: changing the robot simultaneously changes the camera viewpoint, collision geometry, motion constraints, action semantics, and the consequences of perception error. This survey argues that reliable transfer therefore requires a factorized navigation stack with four explicit components: an environment-centric representation, an embodiment descriptor, a feasibility layer that converts scene understanding into body-aware constraints, and a deployment layer that handles dynamics and sim-to-real mismatch. We organize recent work into three complementary families: heterogeneous-data generalist policies, embodiment-randomized or embodiment-conditioned policies, and geometric abstraction with configuration-driven planning. Representative systems include GNM, ViNT, NoMaD, RING, NavFoM, EA-Nav, CeRLP, and AgniNav. We then connect these methods to evidence from sim-to-real navigation research, showing why simulator fidelity alone is insufficient and why collision behavior, action interfaces, and evaluation design strongly affect real-world conclusions. Finally, we propose an evaluation protocol that separates environment, sensor, geometry, dynamics, and reality gaps, and identify open problems in online embodiment estimation, adaptive safety margins, heterogeneous sensing, formally constrained learning, and cross-embodiment foundation models.","references":[{"text":"D. Shah, A. Sridhar, A. Bhorkar, N. Hirose, and S. Levine. GNM: A General Navigation Model to Drive Any Robot. arXiv:2210.03370, 2022.","identifier":"arXiv:2210.03370","url":"https://arxiv.org/abs/2210.03370"},{"text":"D. Shah, A. Sridhar, N. Dashora, K. Stachowicz, K. Black, N. Hirose, and S. Levine. ViNT: A Foundation Model for Visual Navigation. arXiv:2306.14846, 2023.","identifier":"arXiv:2306.14846","url":"https://arxiv.org/abs/2306.14846"},{"text":"A. Sridhar, D. Shah, C. Glossop, and S. Levine. NoMaD: Goal Masked Diffusion Policies for Navigation and Exploration. arXiv:2310.07896, 2023.","identifier":"arXiv:2310.07896","url":"https://arxiv.org/abs/2310.07896"},{"text":"A. Eftekhar et al. The One RING: a Robotic Indoor Navigation Generalist. arXiv:2412.14401, 2024.","identifier":"arXiv:2412.14401","url":"https://arxiv.org/abs/2412.14401"},{"text":"J. Zhang et al. Embodied Navigation Foundation Model. arXiv:2509.12129, 2025.","identifier":"arXiv:2509.12129","url":"https://arxiv.org/abs/2509.12129"},{"text":"J. Zhang, Y. Du, X. Guo, S. Sun, X. Liu, Y. Sun, G. Lu, W. Sui, and J. Li. EA-Nav: Learning Safe Visual Navigation Policies with Embodiment Awareness. arXiv:2607.19880, 2026.","identifier":"arXiv:2607.19880","url":"https://arxiv.org/abs/2607.19880"},{"text":"H. Xi, M. Tan, X. Zhang, S. Cheng, S. Wang, Y. Gu, X. Shen, and W. Zhang. CeRLP: A Cross-embodiment Robot Local Planning Framework for Visual Navigation. arXiv:2603.19602, 2026.","identifier":"arXiv:2603.19602","url":"https://arxiv.org/abs/2603.19602"},{"text":"T. Zang, S. Cheng, H. Huang, S. Wang, and W. Zhang. AgniNav: Configuration-Driven Cross-Embodiment Local Planning for Robot Navigation. arXiv:2606.10903, 2026.","identifier":"arXiv:2606.10903","url":"https://arxiv.org/abs/2606.10903"},{"text":"M. Savva et al. Habitat: A Platform for Embodied AI Research. arXiv:1904.01201, 2019.","identifier":"arXiv:1904.01201","url":"https://arxiv.org/abs/1904.01201"},{"text":"F. Xia, A. Zamir, Z.-Y. He, A. Sax, J. Malik, and S. Savarese. Gibson Env: Real-World Perception for Embodied Agents. arXiv:1808.10654, 2018.","identifier":"arXiv:1808.10654","url":"https://arxiv.org/abs/1808.10654"},{"text":"A. Kadian et al. Sim2Real Predictivity: Does Evaluation in Simulation Predict Real-World Performance? arXiv:1912.06321, 2019.","identifier":"arXiv:1912.06321","url":"https://arxiv.org/abs/1912.06321"},{"text":"P. Anderson, A. Shrivastava, J. Truong, A. Majumdar, D. Parikh, D. Batra, and S. Lee. Sim-to-Real Transfer for Vision-and-Language Navigation. Proceedings of Machine Learning Research, 155:671-681, 2021.","identifier":"PMLR 155","url":"https://proceedings.mlr.press/v155/anderson21a.html"},{"text":"J. Truong, M. Rudolph, N. H. Yokoyama, S. Chernova, D. Batra, and A. Rai. Rethinking Sim2Real: Lower Fidelity Simulation Leads to Higher Sim2Real Transfer in Navigation. Proceedings of Machine Learning Research, 205:859-870, 2023.","identifier":"PMLR 205","url":"https://proceedings.mlr.press/v205/truong23a.html"},{"text":"Z. Wang, X. Li, J. Yang, Y. Liu, and S. Jiang. Sim-to-Real Transfer via 3D Feature Fields for Vision-and-Language Navigation. Proceedings of Machine Learning Research, 270:2982-2995, 2025.","identifier":"PMLR 270","url":"https://proceedings.mlr.press/v270/wang25i.html"},{"text":"P. Anderson et al. On Evaluation of Embodied Navigation Agents. arXiv:1807.06757, 2018.","identifier":"arXiv:1807.06757","url":"https://arxiv.org/abs/1807.06757"},{"text":"N. Yokoyama, S. Ha, and D. Batra. Success Weighted by Completion Time: A Dynamics-Aware Evaluation Criteria for Embodied Navigation. arXiv:2103.08022, 2021.","identifier":"arXiv:2103.08022","url":"https://arxiv.org/abs/2103.08022"},{"text":"J. Hu, S. Yuan, G. C. R. Bethala, A. Tzes, and Y. Fang. Learning Adaptive Safety Margins for Visual Navigation. arXiv:2607.18200, 2026.","identifier":"arXiv:2607.18200","url":"https://arxiv.org/abs/2607.18200"}],"affiliation":null,"source_type":"preprint","source_label":"Kaggleyes Academy Preprint","license":"CC BY 4.0","has_doi":false,"external_doi":null,"status":"published","kagg_id":"KAGG:2607.00001","slug":"2607.00001","registered_at":"2026-07-29T14:45:00Z","updated_at":"2026-07-29T17:45:08Z","canonical_url":"https://academy.kaggleyes.top/abs/KAGG:2607.00001","resolver_url":"https://academy.kaggleyes.top/id/KAGG:2607.00001","html_path":"data/articles/2607.00001/index.html"},"resolver_url":"https://academy.kaggleyes.top/id/KAGG:2607.00001"}