AI & ML Papers
33.3K subscribers
7.17K photos
551 videos
24 files
7.86K links
Advancing research in Machine Learning – practical insights, tools, and techniques for researchers.

Admin: @HusseinSheikho || @Hussein_Sheikho
Download Telegram
AI & ML Papers
Photo
🔥 Masked Visual Actions for Unified World Modeling

💡 The paper introduces Masked Visual Actions for unified world modeling, a method that enables video models to learn how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge addressed is how to communicate action to such models in a form aligned with the visual space in which they learned interaction priors, yet still grounded in physical manipulation.

The proposed method, Masked Visual Actions, expresses action as a partially revealed trajectory of an arbitrary entity in a video, using a pixel-space control interface. This allows the model to act as a forward dynamics model that predicts the scene's response to low-level robot actions, while also recovering robot behavior consistent with a desired outcome.

The method is fine-tuned with only 15 hours of masked examples from real videos and simulation, and achieves strong visual fidelity and controllability across diverse scenes and multiple embodiments. The model produces imagined rollouts whose outcomes correlate with real-world execution for policy evaluation, improves decision-making by ranking candidate futures in model-based planning, and supports inverse modeling by synthesizing robot motion from desired object motion.

The contributions of the paper include a novel method for communicating action to video models, a pixel-space control interface for expressing action, and a model that can predict scene responses to robot actions and recover robot behavior consistent with desired outcomes. The results demonstrate the effectiveness of the method in achieving strong visual fidelity and controllability, and its potential applications in downstream manipulation settings, such as policy evaluation, model-based planning, and inverse modeling.


📅 Published on Jul 21

🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.19343
• PDF: https://arxiv.org/pdf/2607.19343
• Project Page: https://masked-visual-actions.github.io/

━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus

#MaskedVisualActions #UnifiedWorldModeling #VideoModels #RoboticWorldModeling #ForwardDynamicsModel
1