AI & ML Papers
Photo
🔥 Masked Visual Actions for Unified World Modeling
📅 Published on Jul 21
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.19343
• PDF: https://arxiv.org/pdf/2607.19343
• Project Page: https://masked-visual-actions.github.io/
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#MaskedVisualActions #UnifiedWorldModeling #VideoModels #RoboticWorldModeling #ForwardDynamicsModel
💡 The paper introduces Masked Visual Actions for unified world modeling, a method that enables video models to learn how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge addressed is how to communicate action to such models in a form aligned with the visual space in which they learned interaction priors, yet still grounded in physical manipulation.
The proposed method, Masked Visual Actions, expresses action as a partially revealed trajectory of an arbitrary entity in a video, using a pixel-space control interface. This allows the model to act as a forward dynamics model that predicts the scene's response to low-level robot actions, while also recovering robot behavior consistent with a desired outcome.
The method is fine-tuned with only 15 hours of masked examples from real videos and simulation, and achieves strong visual fidelity and controllability across diverse scenes and multiple embodiments. The model produces imagined rollouts whose outcomes correlate with real-world execution for policy evaluation, improves decision-making by ranking candidate futures in model-based planning, and supports inverse modeling by synthesizing robot motion from desired object motion.
The contributions of the paper include a novel method for communicating action to video models, a pixel-space control interface for expressing action, and a model that can predict scene responses to robot actions and recover robot behavior consistent with desired outcomes. The results demonstrate the effectiveness of the method in achieving strong visual fidelity and controllability, and its potential applications in downstream manipulation settings, such as policy evaluation, model-based planning, and inverse modeling.
📅 Published on Jul 21
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.19343
• PDF: https://arxiv.org/pdf/2607.19343
• Project Page: https://masked-visual-actions.github.io/
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#MaskedVisualActions #UnifiedWorldModeling #VideoModels #RoboticWorldModeling #ForwardDynamicsModel
GitHub
Hugging Face
The AI community building the future. Hugging Face has 458 repositories available. Follow their code on GitHub.
❤1