AI & ML Papers
Photo
🔥 Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation
📅 Published on Jul 14
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.13125
• PDF: https://arxiv.org/pdf/2607.13125
• Project Page: https://boogu.org/
🤖 Models citing this paper:
• https://huggingface.co/Boogu/Boogu-Image-0.1-Edit
• https://huggingface.co/Boogu/Boogu-Image-0.1-Base
• https://huggingface.co/Boogu/Boogu-Image-0.1-Turbo
🚀 Spaces citing this paper:
• https://huggingface.co/spaces/RioShiina/ImageGen
• https://huggingface.co/spaces/multimodalart/Boogu-Image
• https://huggingface.co/spaces/Tinchote/ImageGen
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#MultimodalUnderstanding #TextToImageGeneration #OpenSourceAI #MultimodalGeneration #InstructionBasedEditing
💡 The paper introduces Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family that delivers competitive performance in high-quality text-to-image generation, fast inference, instruction-based editing, and bilingual text rendering. The model family consists of Base, Turbo, Edit, and Edit-Turbo variants.
The problem addressed in the paper is that closed-source multimodal systems achieve strong performance through system-level integration rather than a single model, yet their internal practices remain largely undisclosed. The authors demonstrate that targeted improvements in model understanding, data quality, and training pipelines, coupled with genetic inference-time scaling, can substantially enhance generation and editing performance even under highly constrained compute budgets.
The method used to achieve this involves making improvements in model understanding, data quality, and training pipelines. The authors also use genetic inference-time scaling to enhance performance. The model is trained on a dataset of 208.62 million unique images, and the theoretical training cost is approximately 400,000 dollars.
The results show that Boogu-Image-0.1 consistently matches or surpasses other open-source models across standard benchmarks and achieves results approaching leading closed-source systems. The authors share practical discussions and release weights, code, and recipes under Apache 2.0 to advance the open ecosystem for unified multimodal understanding and generation. The code is available on GitHub.
Overall, the paper contributes to the development of open-source multimodal models that can achieve competitive performance with closed-source systems, and provides a valuable resource for the research community.
📅 Published on Jul 14
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.13125
• PDF: https://arxiv.org/pdf/2607.13125
• Project Page: https://boogu.org/
🤖 Models citing this paper:
• https://huggingface.co/Boogu/Boogu-Image-0.1-Edit
• https://huggingface.co/Boogu/Boogu-Image-0.1-Base
• https://huggingface.co/Boogu/Boogu-Image-0.1-Turbo
🚀 Spaces citing this paper:
• https://huggingface.co/spaces/RioShiina/ImageGen
• https://huggingface.co/spaces/multimodalart/Boogu-Image
• https://huggingface.co/spaces/Tinchote/ImageGen
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#MultimodalUnderstanding #TextToImageGeneration #OpenSourceAI #MultimodalGeneration #InstructionBasedEditing
GitHub
Hugging Face
The AI community building the future. Hugging Face has 458 repositories available. Follow their code on GitHub.
AI & ML Papers
Photo
🔥 SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
📅 Published on Jul 16
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.14777
• PDF: https://arxiv.org/pdf/2607.14777
• Project Page: https://jinyangwu.github.io/seed/
🤖 Models citing this paper:
• https://huggingface.co/Jinyang23/Seed-AlfWorld-3B
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#AgenticReinforcementLearning #SelfEvolvingDistillation #OnPolicyDistillation #MultiTurnInteraction #ReinforcementLearningForLanguageModels
💡 The paper proposes a self-evolving framework called SEED, which stands for Self-Evolving On-Policy Distillation, to improve the performance of large language models in interactive tasks involving multi-turn interaction, tool use, and environment feedback. The problem addressed is that outcome-based reinforcement learning provides limited guidance on intermediate decisions, leaving a supervision gap between episode-level outcomes and token-level policy learning.
SEED converts completed on-policy trajectories into training-time hindsight skills and distills their behavioral effect back into the policy model. The framework first fine-tunes the policy to analyze completed trajectories and generate natural language skills that capture reusable workflows, decisive observations, or failure-avoidance rules. During reinforcement learning, the current policy both collects trajectories and serves as the analyzer that extracts hindsight skills from them.
The policy updates therefore improve subsequent decision-making and skill analysis together, allowing hindsight supervision to evolve with the policy. SEED then re-scores the sampled actions under ordinary and skill-augmented contexts, converting the skill-induced probability shift into a dense token-level on-policy distillation signal. This signal is jointly optimized with outcome-based reinforcement learning, keeping the auxiliary supervision aligned with the current trajectory distribution.
The results show that SEED consistently improves performance and sample efficiency, exhibiting robust generalization to unseen scenarios, in extensive experiments on text-based and vision-based agentic tasks. The code for SEED is available, making it possible for others to implement and build upon the framework. Overall, SEED provides a novel approach to addressing the supervision gap in outcome-based reinforcement learning, with promising results in a range of tasks.
📅 Published on Jul 16
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.14777
• PDF: https://arxiv.org/pdf/2607.14777
• Project Page: https://jinyangwu.github.io/seed/
🤖 Models citing this paper:
• https://huggingface.co/Jinyang23/Seed-AlfWorld-3B
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#AgenticReinforcementLearning #SelfEvolvingDistillation #OnPolicyDistillation #MultiTurnInteraction #ReinforcementLearningForLanguageModels
GitHub
Hugging Face
The AI community building the future. Hugging Face has 458 repositories available. Follow their code on GitHub.
AI & ML Papers
Photo
🔥 VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
📅 Published on Jul 16
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.14935
• PDF: https://arxiv.org/pdf/2607.14935
• Project Page: https://mcg-nju.github.io/VideoChat3
🤖 Models citing this paper:
• https://huggingface.co/MCG-NJU/VideoChat3-4B
• https://huggingface.co/MCG-NJU/I3D-ViT
📊 Datasets citing this paper:
• https://huggingface.co/datasets/MCG-NJU/VideoChat3-LV116k
• https://huggingface.co/datasets/MCG-NJU/VideoChat3-Academic2M
• https://huggingface.co/datasets/MCG-NJU/VideoChat3-OL617k
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#VideoUnderstanding #MultimodalLearning #LargeLanguageModels #VideoCentricAI #EfficientComputerVision
💡 The paper introduces VideoChat3, a fully open, efficient, and generalist video-centric multimodal large language model for video understanding. Current open-source models are limited in several ways, struggling to generalize across diverse video types and being computationally demanding, which restricts their efficiency and scalability. Most models are also only partially open, with key components such as training code, strategy, or datasets unavailable, hindering reproducibility and slowing community-driven development.
To address these issues, VideoChat3 advances video understanding through two complementary designs. For efficiency, it introduces the Inflated 3D Vision Transformer and Adaptive Frame Resolution for Streaming Video Perception, enabling efficient spatiotemporal representation and reducing the cost of processing video inputs during training and inference. For effectiveness, it develops a scalable video data synthesis pipeline that curates three diverse, high-quality training datasets, covering general, long-form, and streaming video scenarios, which improves the model's generalization across domains.
By integrating these designs, VideoChat3 achieves a rare balance of broad generalization and computational efficiency. Experiments across general, long-form, and streaming benchmarks demonstrate that VideoChat3 surpasses prior open-source models with equal or larger parameter counts, using only 4B parameters and achieving higher efficiency. The paper's contributions include a fully open and efficient video understanding model, a scalable video data synthesis pipeline, and state-of-the-art results on various video understanding benchmarks.
📅 Published on Jul 16
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.14935
• PDF: https://arxiv.org/pdf/2607.14935
• Project Page: https://mcg-nju.github.io/VideoChat3
🤖 Models citing this paper:
• https://huggingface.co/MCG-NJU/VideoChat3-4B
• https://huggingface.co/MCG-NJU/I3D-ViT
📊 Datasets citing this paper:
• https://huggingface.co/datasets/MCG-NJU/VideoChat3-LV116k
• https://huggingface.co/datasets/MCG-NJU/VideoChat3-Academic2M
• https://huggingface.co/datasets/MCG-NJU/VideoChat3-OL617k
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#VideoUnderstanding #MultimodalLearning #LargeLanguageModels #VideoCentricAI #EfficientComputerVision
GitHub
Hugging Face
The AI community building the future. Hugging Face has 458 repositories available. Follow their code on GitHub.
AI & ML Papers
Photo
🔥 BadWAM: When World-Action Models Dream Right but Act Wrong
📅 Published on Jul 16
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.15207
• PDF: https://arxiv.org/pdf/2607.15207
• Project Page: https://liqiiiii.github.io/BadWAM/
📊 Datasets citing this paper:
• https://huggingface.co/datasets/cy0307/awesome-egocentric-atlas
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#WorldActionModels #EmbodiedControl #AdversarialAttacks #WorldActionDrift #RobustnessInAI
💡 This paper introduces BadWAM, a unified framework for modeling and evaluating World-Action Drift Attacks on World-Action Models. World-Action Models are a promising foundation for embodied control, learning representations that couple action generation with future world prediction. However, this paper shows that the assumption that World-Action Models are robust and safe is fragile. The authors introduce a new class of attacks that use small visual perturbations to break the alignment between what a World-Action Model imagines and what it executes. The BadWAM framework characterizes this attack surface along two natural criteria: attack strength and stealthiness. The authors evaluate BadWAM across different variants of World-Action Models and show that their attacks substantially reduce task success rates under closed-loop execution. For example, their action-only attack reduces model performance from 96.5 percent to 43.1 percent success. The results of their imagination-preserving attack further expose a World-Action Model-specific vulnerability, where moderate future-preserving regularization can maintain strong attack performance while reducing future imagination drift. The paper's contributions highlight the potential risks and vulnerabilities of World-Action Models and provide a framework for evaluating and improving their robustness and safety.
📅 Published on Jul 16
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.15207
• PDF: https://arxiv.org/pdf/2607.15207
• Project Page: https://liqiiiii.github.io/BadWAM/
📊 Datasets citing this paper:
• https://huggingface.co/datasets/cy0307/awesome-egocentric-atlas
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#WorldActionModels #EmbodiedControl #AdversarialAttacks #WorldActionDrift #RobustnessInAI
GitHub
Hugging Face
The AI community building the future. Hugging Face has 458 repositories available. Follow their code on GitHub.
❤1
🔥 SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration
📅 Published on Jul 16
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.15257
• PDF: https://arxiv.org/pdf/2607.15257
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#InformationSeekingAgents #OpenDomainSearch #MultiAgentSystems #RelationalSchemaCompletion #CollaborativeSearchFrameworks
💡 The paper introduces SearchOS, a system level multi agent framework that aims to improve the collaboration of information seeking agents in open domain search tasks. The problem addressed is that current single and multi agent systems struggle to track task progress as interaction histories grow, leading to repetitive loops, wasted search budgets, and compromised output quality.
To address this, the authors formulate open domain information seeking as relational schema completion with grounded citations, where agents discover entities, populate attributes across linked tables, and anchor each value to source evidence. They design Search Oriented Context Management, which externalizes the evolving state into a Frontier Task, an Evidence Graph, a Coverage Map, and a Failure Memory.
Built on this, SearchOS applies a pipeline parallel scheduling mechanism that overlaps the execution of sub agents and continuously refills free slots with tasks targeting unresolved coverage gaps to improve utilization and throughput. The system also introduces a Search Tool Middleware Harness that intercepts model and tool interactions to record grounded evidence and react to stalls or budget exhaustion, and provides a reusable hierarchical skill system comprising strategy and access skills to augment the agents' search process and avoid repeating failed search patterns across runs.
The results show that SearchOS leads all metrics among the evaluated single and multi agent baselines on Wide Search and GISa, paving the way towards robust information seeking collaboration. The paper's contributions include the formulation of open domain information seeking as relational schema completion, the design of Search Oriented Context Management, and the introduction of a pipeline parallel scheduling mechanism and a Search Tool Middleware Harness, all of which work together to improve the collaboration of information seeking agents.
📅 Published on Jul 16
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.15257
• PDF: https://arxiv.org/pdf/2607.15257
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#InformationSeekingAgents #OpenDomainSearch #MultiAgentSystems #RelationalSchemaCompletion #CollaborativeSearchFrameworks
GitHub
Hugging Face
The AI community building the future. Hugging Face has 458 repositories available. Follow their code on GitHub.
Forwarded from Machine Learning
Boost me and we both win! Sign up on Kimi and we each get a guaranteed benefit — up to 1-Year Membership Credits: https://kimi-bot.com/activities/viral-referral/share?scenario=invite&from=share_poster&invitation_code=PJMK9U
❤2🔥1
Follow the Ai Tools Daily channel on WhatsApp:
https://whatsapp.com/channel/0029VbChm8XAojYoblmIW60h
https://whatsapp.com/channel/0029VbChm8XAojYoblmIW60h
AI & ML Papers
Photo
🔥 Moonshine: Speech Recognition for Live Transcription and Voice Commands
📅 Published on Oct 21, 2024
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2410.15608
• PDF: https://arxiv.org/pdf/2410.15608
🤖 Models citing this paper:
• https://huggingface.co/UsefulSensors/moonshine
• https://huggingface.co/UsefulSensors/moonshine-base
• https://huggingface.co/UsefulSensors/moonshine-tiny
🚀 Spaces citing this paper:
• https://huggingface.co/spaces/Taf2023/higgs-audio-v3-tts
• https://huggingface.co/spaces/multimodalart/MisoTTS
• https://huggingface.co/spaces/microsoft/paza-bench
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#SpeechRecognitionTechnology #LiveTranscriptionSystems #VoiceCommandProcessing #TransformerArchitecture #RotaryPositionEmbedding
💡 The paper introduces Moonshine, a speech recognition model designed for live transcription and voice command processing. The problem addressed is the high computational requirements of traditional speech recognition models, which can be a barrier for real-time and resource-constrained applications. To solve this, the authors propose an encoder-decoder transformer architecture that uses Rotary Position Embedding instead of traditional absolute position embeddings. This approach allows the model to be trained on speech segments of various lengths without using zero-padding, making it more efficient during inference time. The results show that the Moonshine model, specifically the Moonshine Tiny version, achieves a 5x reduction in compute requirements for transcribing a 10-second speech segment compared to OpenAI's Whisper tiny.en model, without increasing word error rates on standard evaluation datasets. This makes Moonshine a promising solution for real-time and resource-constrained speech recognition applications.
📅 Published on Oct 21, 2024
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2410.15608
• PDF: https://arxiv.org/pdf/2410.15608
🤖 Models citing this paper:
• https://huggingface.co/UsefulSensors/moonshine
• https://huggingface.co/UsefulSensors/moonshine-base
• https://huggingface.co/UsefulSensors/moonshine-tiny
🚀 Spaces citing this paper:
• https://huggingface.co/spaces/Taf2023/higgs-audio-v3-tts
• https://huggingface.co/spaces/multimodalart/MisoTTS
• https://huggingface.co/spaces/microsoft/paza-bench
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#SpeechRecognitionTechnology #LiveTranscriptionSystems #VoiceCommandProcessing #TransformerArchitecture #RotaryPositionEmbedding
GitHub
Hugging Face
The AI community building the future. Hugging Face has 458 repositories available. Follow their code on GitHub.
❤2
AI & ML Papers
Photo
🔥 Flavors of Moonshine: Tiny Specialized ASR Models for Edge Devices
📅 Published on Sep 2, 2025
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2509.02523
• PDF: https://arxiv.org/pdf/2509.02523
🤖 Models citing this paper:
• https://huggingface.co/UsefulSensors/moonshine-tiny-ja
• https://huggingface.co/UsefulSensors/moonshine-tiny-ar
• https://huggingface.co/UsefulSensors/moonshine-tiny-zh
🚀 Spaces citing this paper:
• https://huggingface.co/spaces/wmoto-ai/moonshine-tiny-ja-demo
• https://huggingface.co/spaces/Haitam03/whisper-tiny-ar-quran
• https://huggingface.co/spaces/Maycsz/video-transcribe-mcp
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#AutomaticSpeechRecognition #MultilingualModeling #EdgeDeviceComputing #SpecializedASR #TinyMLModels
💡 The paper challenges the common assumption that multilingual automatic speech recognition models are superior to monolingual models. Instead, the authors demonstrate that training monolingual models on a balanced mix of high-quality human-labeled, pseudo-labeled, and synthetic data can achieve better performance for small model sizes. The proposed approach, called Flavors of Moonshine, involves training tiny specialized models for underrepresented languages. The results show that these models, with only 27 million parameters, outperform larger multilingual models, including the Whisper Tiny, Whisper Small, and Whisper Medium models. On average, the Moonshine models achieve error rates 48 percent lower than the Whisper Tiny model. The authors release models for six languages, including Arabic, Chinese, Japanese, Korean, Ukrainian, and Vietnamese, under an open-source license. The contributions of this paper advance the state of the art for small automatic speech recognition models, enabling accurate on-device speech recognition for languages that previously had limited support.
📅 Published on Sep 2, 2025
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2509.02523
• PDF: https://arxiv.org/pdf/2509.02523
🤖 Models citing this paper:
• https://huggingface.co/UsefulSensors/moonshine-tiny-ja
• https://huggingface.co/UsefulSensors/moonshine-tiny-ar
• https://huggingface.co/UsefulSensors/moonshine-tiny-zh
🚀 Spaces citing this paper:
• https://huggingface.co/spaces/wmoto-ai/moonshine-tiny-ja-demo
• https://huggingface.co/spaces/Haitam03/whisper-tiny-ar-quran
• https://huggingface.co/spaces/Maycsz/video-transcribe-mcp
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#AutomaticSpeechRecognition #MultilingualModeling #EdgeDeviceComputing #SpecializedASR #TinyMLModels
GitHub
Hugging Face
The AI community building the future. Hugging Face has 458 repositories available. Follow their code on GitHub.
AI & ML Papers
Photo
🔥 Self-Improvements in Modern Agentic Systems: A Survey
📅 Published on Jul 14
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.13104
• PDF: https://arxiv.org/pdf/2607.13104
• Project Page: https://selfimproving-agent.github.io/
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#AgenticSystems #AutonomousAgents #SelfImprovingSystems #ArtificialIntelligenceEvolution #AdaptiveSystemDesign
💡 The paper presents a survey of self-improving autonomous agents that are transitioning from research prototypes to deployed systems. The primary goal of these agents is to achieve controllable evolution or adaptation from experience with minimal or no human input. The authors frame modern self-improving agents as adaptive systems that convert experience into accumulated capability gains. They propose a system-level framework that represents a modern agent as a configuration coupling a foundation model with an operational scaffold of prompts, memory, tools, and control logic. Within this framework, self-improvement is formalized as a self-induced update operator that obtains and commits updates to model parameters or scaffold components.
The authors organize prior work by update target and by the signals that drive change, and then review applications and discuss evaluation. They also identify open problems and future directions. The paper provides a comprehensive overview of the current state of self-improving agents and offers a framework for understanding and developing these systems. The authors also provide a resource for tracking technical updates on self-improving agents.
The paper's contributions include a systematic framework for understanding self-improving agents, a review of prior work in the field, and an identification of open problems and future directions. The authors also provide a resource for tracking technical updates in the field. Overall, the paper provides a comprehensive overview of self-improving agents and offers a framework for understanding and developing these systems. The paper's results highlight the potential of self-improving agents to achieve controllable evolution or adaptation from experience with minimal or no human input, and identify areas for future research and development.
📅 Published on Jul 14
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.13104
• PDF: https://arxiv.org/pdf/2607.13104
• Project Page: https://selfimproving-agent.github.io/
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#AgenticSystems #AutonomousAgents #SelfImprovingSystems #ArtificialIntelligenceEvolution #AdaptiveSystemDesign
GitHub
Hugging Face
The AI community building the future. Hugging Face has 458 repositories available. Follow their code on GitHub.
❤1
AI & ML Papers
Photo
🔥 SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
📅 Published on Jun 2, 2025
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2506.01844
• PDF: https://arxiv.org/pdf/2506.01844
• Project Page: https://huggingface.co/blog/smolvla
🤖 Models citing this paper:
• https://huggingface.co/lerobot/smolvla_base
• https://huggingface.co/HuggingFaceVLA/smolvla_libero
• https://huggingface.co/jadechoghari/smolvla_metaworld
📊 Datasets citing this paper:
• https://huggingface.co/datasets/lerobot/community_dataset_v3
• https://huggingface.co/datasets/HuggingFaceVLA/community_dataset_v1
• https://huggingface.co/datasets/HuggingFaceVLA/community_dataset_v2
🚀 Spaces citing this paper:
• https://huggingface.co/spaces/HuggingFaceVLA/libero-vla-leaderboard
• https://huggingface.co/spaces/lerobot/robot-learning-tutorial
• https://huggingface.co/spaces/arpitg1304/lerobot_scripts_simplified
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#VisionLanguageAction #EfficientRobotics #CompactNeuralNetworks #RoboticsForAll #AffordableAI
💡 The paper introduces SmolVLA, a compact and efficient vision-language-action model designed for affordable and efficient robotics. The problem with existing vision-language-action models is that they are typically very large, with billions of parameters, which leads to high training costs and limited deployability on consumer-grade hardware. These models also rely on academic and industrial datasets, overlooking community-collected data from affordable robotic platforms.
To address this issue, the authors developed SmolVLA, a small and efficient model that can be trained on a single GPU and deployed on consumer-grade GPUs or even CPUs. SmolVLA is designed to be community-driven, using community-collected data from affordable robotic platforms, which reduces both training and inference costs while retaining competitive performance.
The method used to achieve this involves adapting vision-language models into vision-language-action models, but with a much smaller size. The authors also introduced an asynchronous inference stack that decouples perception and action prediction from action execution, allowing for higher control rates with chunked action generation.
The results show that SmolVLA achieves performance comparable to vision-language-action models that are 10 times larger, despite its compact size. The model is evaluated on a range of simulated and real-world robotic benchmarks, and the authors release all code, pretrained models, and training data. Overall, SmolVLA provides a more efficient and affordable solution for robotics, making it possible to deploy vision-language-action models on consumer-grade hardware.
📅 Published on Jun 2, 2025
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2506.01844
• PDF: https://arxiv.org/pdf/2506.01844
• Project Page: https://huggingface.co/blog/smolvla
🤖 Models citing this paper:
• https://huggingface.co/lerobot/smolvla_base
• https://huggingface.co/HuggingFaceVLA/smolvla_libero
• https://huggingface.co/jadechoghari/smolvla_metaworld
📊 Datasets citing this paper:
• https://huggingface.co/datasets/lerobot/community_dataset_v3
• https://huggingface.co/datasets/HuggingFaceVLA/community_dataset_v1
• https://huggingface.co/datasets/HuggingFaceVLA/community_dataset_v2
🚀 Spaces citing this paper:
• https://huggingface.co/spaces/HuggingFaceVLA/libero-vla-leaderboard
• https://huggingface.co/spaces/lerobot/robot-learning-tutorial
• https://huggingface.co/spaces/arpitg1304/lerobot_scripts_simplified
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#VisionLanguageAction #EfficientRobotics #CompactNeuralNetworks #RoboticsForAll #AffordableAI
GitHub
Hugging Face
The AI community building the future. Hugging Face has 458 repositories available. Follow their code on GitHub.
❤1
🔥 Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories
📅 Published on Jul 16
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.15330
• PDF: https://arxiv.org/pdf/2607.15330
• Project Page: https://robotics.xiaomi.com/xiaomi-robotics-1.html
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#VisionLanguageModels #MobileManipulation #RoboticsLearning #RealWorldTrajectories #LanguageActionInterfaces
💡 The paper presents Xiaomi-Robotics-1, a foundational vision-language-action model that can follow diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments. The model is trained using a two-stage training recipe consisting of pre-training and post-training. During pre-training, the model is trained on over 100,000 hours of real-world manipulation trajectories collected via UM devices, and a scalable auto-labeling pipeline is developed to annotate trajectory clips with natural language descriptions of scene state transitions. This provides rich and precise conditioning for action learning.
During post-training, the model is fine-tuned to align with robot embodiment and imperative instructions that humans naturally use to prompt robots. The experiments demonstrate strong scaling behavior, with the model consistently improving with increased data scales and model sizes during pre-training. This scaling behavior directly transfers to post-training, where stronger pre-training models yield better out-of-the-box real-robot performance in unseen environments.
The results show that Xiaomi-Robotics-1 serves as a strong robot foundation policy that can be efficiently fine-tuned on complex, dexterous tasks with high data efficiency. Across multiple simulation benchmarks, Xiaomi-Robotics-1 outperforms state-of-the-art methods, establishing a new state-of-the-art with a 57.6% success rate on RoboCasa 365, surpassing the previous best of 46.6%. Additionally, it achieves an average score of 20.07 on RoboDojo, significantly outperforming the prior state-of-the-art. The code and model checkpoints will be released.
Overall, the paper contributes to the development of a robust and scalable vision-language-action model that can be applied to a wide range of robotic tasks, and demonstrates the effectiveness of the proposed two-stage training recipe and auto-labeling pipeline in achieving state-of-the-art performance.
📅 Published on Jul 16
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.15330
• PDF: https://arxiv.org/pdf/2607.15330
• Project Page: https://robotics.xiaomi.com/xiaomi-robotics-1.html
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#VisionLanguageModels #MobileManipulation #RoboticsLearning #RealWorldTrajectories #LanguageActionInterfaces
GitHub
Hugging Face
The AI community building the future. Hugging Face has 458 repositories available. Follow their code on GitHub.