Beyond Teleoperation: Enhancing VLA Robustness via Explicit Kinematic Retargeting of Human Demonstrations

Authors

DOI:

https://doi.org/10.31181/dma412026180

Keywords:

Robot Learning, Vision-Language-Action Models, Motion Retargeting, Inverse Kinematics, Terminal State Ambiguity

Abstract

Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in generalized robotic control, yet their scalability is fundamentally bottlenecked by the high cost and low diversity of teleoperated data. While abundant, human demonstration videos cannot be directly utilized for policy training due to the severe morphological differences between human anatomy and robotic manipulators. To bridge this embodiment gap, this work proposes a lightweight retargeting pipeline that kinematically retargets human interaction data (DexYCB) onto six-degree-of-freedom manipulator trajectories to fine-tune policies based on the pi0.5 architecture. By prioritizing Cartesian positional alignment via constrained Inverse Kinematics (IK) and introducing an object-based grasping heuristic, smooth geometric priors are generated without relying on computationally heavy visual synthesis. Physical evaluations demonstrate that retargeted models significantly outperform standard teleoperation (40.6% success rate), achieving 65.6% success via co-training and a peak 78.1% success rate via two-stage cross-embodiment co-training. Furthermore, evaluations under extreme visual clutter reveal that explicitly retargeted policies exhibit immunity to semantic visual distractors. Finally, we examined and analysed Terminal State Ambiguity, a temporal failure mode where generative models fail to terminate the task when exposed to scenarios similar to the nature of human video priors.

Downloads

Download data is not yet available.

References

Bjorck, J., Castañeda, F., Cherniadev, N., Da, X., Ding, R., Fan, L., ... & Zhu, Y. (2025). Gr00t n1: An open foundation model for generalist humanoid robots. arXiv preprint arXiv:2503.14734. (arXiv:2503.14734). arXiv. https://doi.org/10.48550/arXiv.2503.14734.

Black, K., Brown, N., Darpinian, J., Dhabalia, K., Driess, D., Esmail, A., ... & Zhilinsky, U. (2025). π0.5: A Vision-Language-Action Model with Open-World Generalization (arXiv:2504.16054). arXiv. https://doi.org/10.15607/RSS.2025.XXI.010.

Kawaharazuka, K., Oh, J., Yamada, J., Posner, I., & Zhu, Y. (2025). Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications. IEEE Access, 13, 162467-162504. https://doi.org/10.1109/ACCESS.2025.3609980.

Sapkota, R., Cao, Y., Roumeliotis, K. I., & Karkee, M. (2025). Vision-Language-Action Models: Concepts, Progress, Applications and Challenges (arXiv:2505.04769). arXiv. https://doi.org/10.48550/arXiv.2505.04769.

Xiao, X., Liu, J., Wang, Z., Zhou, Y., Qi, Y., Jiang, S., He, B., & Cheng, Q. (2025). Robot learning in the era of foundation models: A survey. Neurocomputing, 638, 129963. https://doi.org/10.1016/j.neucom.2025.129963.

Lepert, M., Fang, J., & Bohg, J. (2025). Masquerade: Learning from In-the-wild Human Videos using Data-Editing (arXiv:2508.09976). arXiv. https://doi.org/10.48550/arXiv.2508.09976.

Yang, R., Yu, Q., Wu, Y., Yan, R., Li, B., Cheng, A. C., ... & Wang, X. (2025). EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos (arXiv:2507.12440). arXiv. https://doi.org/10.48550/arXiv.2507.12440.

Feng, Z., Li, Q., Liang, H., Yang, R., Shen, Y., Du, Z., Zhang, Z., Deng, Y., Zhao, L., Zhao, H., Lu, Z., Mees, O., Pollefeys, M., Yang, J., & Guo, B. (2026). From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data. https://doi.org/10.36227/techrxiv.177126525.54038135/v1.

Chao, Y. W., Yang, W., Xiang, Y., Molchanov, P., Handa, A., Tremblay, J., ... & Fox, D. (2021). Dexycb: A benchmark for capturing hand grasping of objects. In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 9040-9049). IEEE. https://doi.org/10.1109/CVPR46437.2021.00893.

Walke, H. R., Black, K., Zhao, T. Z., Vuong, Q., Zheng, C., Hansen-Estruch, P., ... & Levine, S. (2023). Bridgedata v2: A dataset for robot learning at scale. In Conference on robot learning (pp. 1723-1736). PMLR.

O’Neill, A., Rehman, A., Maddukuri, A., Gupta, A., Padalkar, A., Lee, A., ... & Chen, M. (2024, May). Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0. In 2024 IEEE International Conference on Robotics and Automation (ICRA) (pp. 6892-6903). IEEE. https://doi.org/10.1109/ICRA57147.2024.10611477.

Qiu, R.-Z., Yang, S., Cheng, X., Chawla, C., Li, J., He, T., ... & Wang, X. (2025). Humanoid Policy ~ Human Policy (arXiv:2503.13441). arXiv. https://doi.org/10.48550/arXiv.2503.13441.

Darvish, K., Tirupachuri, Y., Romualdi, G., Rapetti, L., Ferigo, D., Chavez, F. J. A., & Pucci, D. (2019). Whole-Body Geometric Retargeting for Humanoid Robots. In 2019 IEEE-RAS 19th International Conference on Humanoid Robots (Humanoids), 679-686. https://doi.org/10.1109/Humanoids43949.2019.9035059.

Yin, Z.-H., Wang, C., Pineda, L., Bodduluri, K., Wu, T., Abbeel, P., & Mukadam, M. (2025). Geometric Retargeting: A Principled, Ultrafast Neural Hand Retargeting Algorithm (arXiv:2503.07541). arXiv. https://doi.org/10.1109/IROS60139.2025.11247700.

Fu, Z., Zhao, Q., Wu, Q., Wetzstein, G., & Finn, C. (2024). HumanPlus: Humanoid Shadowing and Imitation from Humans. (arXiv:2406.10454). arXiv. https://doi.org/10.48550/arXiv.2406.10454.

Kareer, S., Patel, D., Punamiya, R., Mathur, P., Cheng, S., Wang, C., ... & Xu, D. (2025). Egomimic: Scaling imitation learning via egocentric video. In 2025 IEEE International Conference on Robotics and Automation (ICRA) (pp. 13226-13233). IEEE. https://doi.org/10.1109/ICRA55743.2025.11127989.

Jeong, T., Byun, T., Kim, J., Choi, K., Oh, J., Lee, S., Darwish, O., Kim, J., & Choi, S. (2024). Robust Robot Motion Retargeting: Rig Unification and Application to Diverse Robots. Research Square. https://doi.org/10.21203/rs.3.rs-4544618/v1.

Huang, H., Bethala, G. C. R., Yuan, S., Wen, C., Tzes, A., & Fang, Y. (2025). One-shot Humanoid Whole-body Motion Learning (arXiv:2510.25241). arXiv. https://doi.org/10.48550/arXiv.2510.25241.

Yan, Y., Mascaro, E. V., & Lee, D. (2024). ImitationNet: Unsupervised Human-to-Robot Motion Retargeting via Shared Latent Space (arXiv:2309.05310). arXiv. https://doi.org/10.1109/Humanoids57100.2023.10375150.

Park, S., Lee, S., Choi, M., Lee, J., Kim, J., Kim, J., & Joo, H. (2025). Learning to Transfer Human Hand Skills for Robot Manipulations (arXiv:2501.04169). arXiv. https://doi.org/10.48550/arXiv.2501.04169.

Canh, T. N., Tran, T.-T., Zhang, H., Gao, Z., Chong, N. Y., & HoangVan, X. (2026). Human-to-Robot Interaction: Learning from Video Demonstration for Robot Imitation (arXiv:2602.19184). arXiv. https://doi.org/10.48550/arXiv.2602.19184.

Lang, X., Wang, Y., Zhou, Y., Ni, C., Li, K., Zhu, J., … & Zhu, Z. (2026). VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis (arXiv:2604.09330). arXiv. https://doi.org/10.48550/arXiv.2604.09330.

Li, H., Zhang, I., Ouyang, R., Wang, X., Zhu, Z., Yang, Z., …. & Wang, X. (2025). MimicDreamer: Aligning Human and Robot Demonstrations for Scalable VLA Training (arXiv:2509.22199). arXiv. https://doi.org/10.48550/arXiv.2509.22199.

Lepert, M., Fang, J., & Bohg, J. (2025). Phantom: Training Robots Without Robots Using Only Human Videos (arXiv:2503.00779). arXiv. https://doi.org/10.48550/arXiv.2503.00779.

Black, K., Brown, N., Driess, D., Esmail, A., Equi, M., Finn, C., … & Zhilinsky, U. (2024). π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv:2410.24164; Version 1). arXiv. https://doi.org/10.15607/RSS.2025.XXI.010.

Luo, H., Feng, Y., Zhang, W., Zheng, S., Wang, Y., Yuan, H., Liu, J., Xu, C., Jin, Q., & Lu, Z. (2025). Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos (arXiv:2507.15597). arXiv. https://doi.org/10.48550/arXiv.2507.15597.

Published

2026-08-07

How to Cite

Nunes Andrade, J. A., & Castelli, M. (2026). Beyond Teleoperation: Enhancing VLA Robustness via Explicit Kinematic Retargeting of Human Demonstrations. Decision Making Advances, 4(1), 131–153. https://doi.org/10.31181/dma412026180