✨VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation
📝 Summary:
VLA-4D enhances robotic manipulation by integrating 4D spatial-temporal awareness into visual and action representations. This enables smoother and more coherent robot control for complex tasks by embedding time into 3D positions and extending action planning with temporal information.
🔹 Publication Date: Published on Nov 21
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2511.17199
• PDF: https://arxiv.org/pdf/2511.17199
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#Robotics #AI #VLAModels #SpatialTemporalAI #RobotManipulation
📝 Summary:
VLA-4D enhances robotic manipulation by integrating 4D spatial-temporal awareness into visual and action representations. This enables smoother and more coherent robot control for complex tasks by embedding time into 3D positions and extending action planning with temporal information.
🔹 Publication Date: Published on Nov 21
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2511.17199
• PDF: https://arxiv.org/pdf/2511.17199
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#Robotics #AI #VLAModels #SpatialTemporalAI #RobotManipulation
✨Xiaomi-Robotics-0: An Open-Sourced Vision-Language-Action Model with Real-Time Execution
📝 Summary:
Xiaomi-Robotics-0 is an open-sourced vision-language-action model enabling real-time, high-performance robot manipulation. It leverages large-scale pre-training and specialized methods for fast execution on real robots, achieving SOTA simulation and high real-robot success.
🔹 Publication Date: Published on Feb 13
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2602.12684
• PDF: https://arxiv.org/pdf/2602.12684
• Project Page: https://xiaomi-robotics-0.github.io/
• Github: https://github.com/XiaomiRobotics/Xiaomi-Robotics-0
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#Robotics #AI #VisionLanguageModels #OpenSource #RobotManipulation
📝 Summary:
Xiaomi-Robotics-0 is an open-sourced vision-language-action model enabling real-time, high-performance robot manipulation. It leverages large-scale pre-training and specialized methods for fast execution on real robots, achieving SOTA simulation and high real-robot success.
🔹 Publication Date: Published on Feb 13
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2602.12684
• PDF: https://arxiv.org/pdf/2602.12684
• Project Page: https://xiaomi-robotics-0.github.io/
• Github: https://github.com/XiaomiRobotics/Xiaomi-Robotics-0
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#Robotics #AI #VisionLanguageModels #OpenSource #RobotManipulation
This media is not supported in your browser
VIEW IN TELEGRAM
✨EgoPush: Learning End-to-End Egocentric Multi-Object Rearrangement for Mobile Robots
📝 Summary:
EgoPush allows mobile robots to rearrange multiple objects in cluttered spaces using a single egocentric camera. It uses an object-centric latent space and stage-decomposed rewards for long-horizon tasks, outperforming end-to-end baselines and demonstrating sim-to-real transfer.
🔹 Publication Date: Published on Feb 20
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2602.18071
• PDF: https://arxiv.org/pdf/2602.18071
• Project Page: https://ai4ce.github.io/EgoPush/
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#Robotics #ComputerVision #AI #MachineLearning #RobotManipulation
📝 Summary:
EgoPush allows mobile robots to rearrange multiple objects in cluttered spaces using a single egocentric camera. It uses an object-centric latent space and stage-decomposed rewards for long-horizon tasks, outperforming end-to-end baselines and demonstrating sim-to-real transfer.
🔹 Publication Date: Published on Feb 20
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2602.18071
• PDF: https://arxiv.org/pdf/2602.18071
• Project Page: https://ai4ce.github.io/EgoPush/
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#Robotics #ComputerVision #AI #MachineLearning #RobotManipulation
AI & ML Papers
Photo
🔥 RepWAM: World Action Modeling with Representation Visual-Action Tokenizers
📅 Published on Jun 11
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2606.13674
• PDF: https://arxiv.org/pdf/2606.13674
• Project Page: https://wdrink.github.io/RepWAM/
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#RobotManipulation #WorldActionModeling #VisualActionTokenizers #LanguageGuidedControl #FutureStatePrediction
💡 The paper introduces RepWAM, a representation-centric world action model that improves robot manipulation performance through language-guided future state prediction and action modeling. The problem with existing world action models is that they use reconstruction-oriented video tokenizers that prioritize visual fidelity over instruction-following dynamics, limiting their ability to connect future prediction with robot control. To address this, the authors propose a semantic visual-action latent space that maps visual inputs into aligned visual and latent action tokens. They train a representation visual-action tokenizer and pretrain their world action model to jointly model future visual states and latent actions under language instructions. The model is then adapted to real robot trajectories for closed-loop manipulation. The results show that RepWAM delivers strong performance across diverse manipulation settings, outperforming reconstruction-oriented alternatives. The authors highlight the value of semantic visual-action tokenization as a promising foundation for world action models and a step toward generalist robot policies. The code and weights for RepWAM will be made available, allowing for further development and application of this technology. Overall, the paper contributes a new approach to world action modeling that prioritizes instruction-following dynamics and semantic understanding, leading to improved robot manipulation performance.
📅 Published on Jun 11
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2606.13674
• PDF: https://arxiv.org/pdf/2606.13674
• Project Page: https://wdrink.github.io/RepWAM/
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#RobotManipulation #WorldActionModeling #VisualActionTokenizers #LanguageGuidedControl #FutureStatePrediction
GitHub
Hugging Face
The AI community building the future. Hugging Face has 467 repositories available. Follow their code on GitHub.