🔥 Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
📅 Published on Jul 16
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.15330
• PDF: https://arxiv.org/pdf/2607.15330
• Project Page: https://robotics.xiaomi.com/xiaomi-robotics-1.html
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#VisionLanguageModels #MobileManipulation #RoboticsLearning #RealWorldTrajectories #LanguageActionInterfaces
💡 The paper presents Xiaomi-Robotics-1, a foundational vision-language-action model that can follow diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments. The model is trained using a two-stage training recipe consisting of pre-training and post-training. During pre-training, the model is trained on over 100,000 hours of real-world manipulation trajectories collected via UM devices, and a scalable auto-labeling pipeline is developed to annotate trajectory clips with natural language descriptions of scene state transitions. This provides rich and precise conditioning for action learning.
During post-training, the model is fine-tuned to align with robot embodiment and imperative instructions that humans naturally use to prompt robots. The experiments demonstrate strong scaling behavior, with the model consistently improving with increased data scales and model sizes during pre-training. This scaling behavior directly transfers to post-training, where stronger pre-training models yield better out-of-the-box real-robot performance in unseen environments.
The results show that Xiaomi-Robotics-1 serves as a strong robot foundation policy that can be efficiently fine-tuned on complex, dexterous tasks with high data efficiency. Across multiple simulation benchmarks, Xiaomi-Robotics-1 outperforms state-of-the-art methods, establishing a new state-of-the-art with a 57.6% success rate on RoboCasa 365, surpassing the previous best of 46.6%. Additionally, it achieves an average score of 20.07 on RoboDojo, significantly outperforming the prior state-of-the-art. The code and model checkpoints will be released.
Overall, the paper contributes to the development of a robust and scalable vision-language-action model that can be applied to a wide range of robotic tasks, and demonstrates the effectiveness of the proposed two-stage training recipe and auto-labeling pipeline in achieving state-of-the-art performance.
📅 Published on Jul 16
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.15330
• PDF: https://arxiv.org/pdf/2607.15330
• Project Page: https://robotics.xiaomi.com/xiaomi-robotics-1.html
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#VisionLanguageModels #MobileManipulation #RoboticsLearning #RealWorldTrajectories #LanguageActionInterfaces
GitHub
Hugging Face
The AI community building the future. Hugging Face has 458 repositories available. Follow their code on GitHub.