AI & ML Papers
34.2K subscribers
7.39K photos
616 videos
24 files
8.16K links
Advancing research in Machine Learning – practical insights, tools, and techniques for researchers.

Admin: @HusseinSheikho || @Hussein_Sheikho
Download Telegram
AI & ML Papers
Photo
πŸ”₯ From Foundation to Application: Improving VLA Models in Practice

πŸ’‘ The paper presents LingBot-VLA 2.0, an improved version of the VLA foundation model, which aims to bridge the gap between laboratory conditions and real-world applications. The main problem addressed is the disparity between the two environments, which hinders the practical implementation of VLA models. To solve this, the authors propose three main improvements.

First, they enhance generalization across tasks and embodiments by expanding the data preprocessing pipeline and training the model on a large dataset of around 60,000 hours of data, including 50,000 hours of robot trajectories from 20 robot configurations and 10,000 hours of human videos.

Second, they extend the action space to include whole-body degrees of freedom, allowing the robots to perform more complex tasks. This is achieved by accommodating degrees of freedom for the heads, waists, mobile bases, and dexterous hands, in addition to dual-arm hardware platforms.

Third, they incorporate predictive dynamics modeling for improved temporal reasoning. This is done by formulating future prediction as a proxy task, using a video representation model for semantic priors and a depth estimation model for geometric cues.

The results show that these modifications have a beneficial impact, as evaluated on the GM-100 benchmark in a generalist setting. Additionally, the expanded pretraining data enables LingBot-VLA 2.0 to demonstrate strong cross-embodiment long-horizon mobile manipulation capability across two robotic platforms. Overall, the paper presents significant improvements to the VLA foundation model, making it more suitable for real-world applications.


πŸ“… Published on Jul 7

πŸ”— Links:
β€’ GitHub: https://github.com/huggingface
β€’ arXiv: https://arxiv.org/abs/2607.06403
β€’ PDF: https://arxiv.org/pdf/2607.06403
β€’ Project Page: https://technology.robbyant.com/lingbot-vla-v2

πŸ“Š Datasets citing this paper:
β€’ https://huggingface.co/datasets/cy0307/awesome-egocentric-atlas

━━━━━━━━━━━━━━━━━━━━━━━━
πŸ“’ By: https://xn--r1a.website/PaperNexus

#VisualLearningAgents #FoundationModels #RobotLearning #EmbodiedAI #VLAmodels
AI & ML Papers
Photo
πŸ”₯ SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe

πŸ’‘ The paper proposes a minimal viable pipeline for skill optimization in autonomous agents, called SkillOpt-Lite, which eliminates redundancies while maintaining convergence and generalization. The existing methods for skill optimization rely on complex pipelines, leaving a fundamental question unaddressed: what constitutes a minimal viable pipeline for skill optimization. The authors formalize skill optimization via Zeroth-Order optimization, mapping classical counterparts to recent literature, and establish three principles for convergence and generalization: trajectory exploration, consensus attribute mining, and independent validation gating.

The proposed SkillOpt-Lite pipeline accelerates convergence and outperforms the full SkillOpt pipeline, improving performance on various benchmarks. For example, it improves LiveMath by 8.8 points on GPT-5.5 and 25.4 points on GPT-5.4-nano, allowing the nano model to surpass the standard GPT-5.4 optimized by SkillOpt. The framework is also integrated into production coding agents like VSCode Copilot, enabling developers to evolve agent skills via a simple interface.

The authors further extend their framework to full harness optimization, called HarnessOpt, which enables the GPT-5.4-nano model to achieve 0.7758 accuracy on the SpreadsheetBench benchmark, outperforming the larger GPT-5.5 model running standard pipelines. The code for SkillOpt-Lite is made available, and the minimal pipeline naturally generalizes to full harness optimization. Overall, the paper contributes a more efficient and effective pipeline for skill optimization, which can be applied to various autonomous agents and production coding systems.


πŸ“… Published on Jul 3

πŸ”— Links:
β€’ GitHub: https://github.com/huggingface
β€’ arXiv: https://arxiv.org/abs/2607.03451
β€’ PDF: https://arxiv.org/pdf/2607.03451
β€’ Project Page: https://evolvinglmms-lab.github.io/SkillOpt-Lite/

πŸ“Š Datasets citing this paper:
β€’ https://huggingface.co/datasets/cy0307/awesome-loop-engineering

━━━━━━━━━━━━━━━━━━━━━━━━
πŸ“’ By: https://xn--r1a.website/PaperNexus

#AutonomousAgentLearning #SkillOptimization #ZerothOrderOptimization #AgentSelfEvolution #MinimalViablePipeline
❀1
AI & ML Papers
Photo
πŸ”₯ Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling

πŸ’‘ The paper proposes a new attention mechanism called Hierarchical Landmark Sparse Attention, which enables efficient long-context language modeling. The problem with current large language models is that they are limited by the quadratic computation cost and poor length extrapolation of dense attention, making it difficult to scale them to long contexts. Existing chunk-wise sparse attention methods have inaccurate chunk selection, which falls short of full attention performance.

The proposed method, Hierarchical Landmark Sparse Attention, learns chunk selection end-to-end under the language modeling loss. It factorizes attention hierarchically, where each query performs attention independently with each retrieved chunk to extract chunk-specific information, and the resulting outputs are fused according to chunk retrieval scores. This allows the model to optimize retrieval scores directly with the language modeling loss, enabling end-to-end retrieval learning and native sparse training.

The experimental results show that Hierarchical Landmark Sparse Attention achieves performance comparable to, and in some cases better than, full attention at in-domain context lengths. Moreover, it extrapolates more than 64 times the training context length with 90 percent retrieval accuracy, far beyond full attention. The method also allows existing full-attention models to be converted to Hierarchical Landmark Sparse Attention with lightweight continued pretraining, preserving in-domain performance while acquiring ultra-long-context extrapolation.

Overall, the proposed method breaks the usual efficiency-performance trade-off, enabling long-context language models that are both more efficient and more effective on general long-context tasks than their full-attention counterparts.


πŸ“… Published on Jul 3

πŸ”— Links:
β€’ GitHub: https://github.com/huggingface
β€’ arXiv: https://arxiv.org/abs/2607.02980
β€’ PDF: https://arxiv.org/pdf/2607.02980

━━━━━━━━━━━━━━━━━━━━━━━━
πŸ“’ By: https://xn--r1a.website/PaperNexus

#HierarchicalAttentionMechanisms #SparseAttentionMethods #LongContextLanguageModeling #EfficientLanguageModeling #InfiniteContextModeling
AI & ML Papers
Photo
πŸ”₯ Vision as Unified Multimodal Generation

πŸ’‘ The paper introduces a unified multimodal model that formulates computer vision tasks as generation problems using natural language and visual prompts. This approach allows for a single model to perform a wide range of vision tasks without requiring task-specific architectures. The model, called SenseNova-Vision, uses natural-language instructions and optional visual prompts to specify tasks and generates responses as text, images, or mixed text-and-image outputs. To support large-scale training, the authors created the SenseNova-Vision Corpus, a computer-vision instruction-response corpus that spans text, image, and mixed targets. The model is trained on this corpus, along with auxiliary multimodal data, and achieves performance comparable to specialized systems across diverse vision tasks, including detection, OCR, keypoint estimation, segmentation, and camera pose estimation. The results demonstrate that a single unified model can match leading task-specialized systems, suggesting that unified multimodal generation is a scalable route for integrating computer vision capabilities into general-purpose foundation models. The model and corpus are publicly available, providing a valuable resource for the research community. Overall, the paper presents a significant contribution to the field of computer vision, offering a unified and flexible approach to tackling a wide range of vision tasks.


πŸ“… Published on Jul 7

πŸ”— Links:
β€’ GitHub: https://github.com/huggingface
β€’ arXiv: https://arxiv.org/abs/2607.06560
β€’ PDF: https://arxiv.org/pdf/2607.06560

πŸ€– Models citing this paper:
β€’ https://huggingface.co/sensenova/SenseNova-Vision-7B-MoT

πŸ“Š Datasets citing this paper:
β€’ https://huggingface.co/datasets/sensenova/SenseNova-Vision-Corpus-50M
β€’ https://huggingface.co/datasets/sensenova/SenseNova-Vision-Benchmark

━━━━━━━━━━━━━━━━━━━━━━━━
πŸ“’ By: https://xn--r1a.website/PaperNexus

#MultimodalGeneration #VisionTasks #NaturalLanguageProcessing #ComputerVision #MultimodalLearning
πŸ”₯ Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning

πŸ’‘ The paper proposes a parallelized autoregressive framework for dense video captioning, which aims to generate temporally grounded descriptions of video events. The existing autoregressive models have limitations in terms of inference efficiency and scalability due to the token-by-token generation paradigm. To address this issue, the authors propose a framework that exploits weak local dependencies across temporally distinct events to enable parallel generation of tokens. The key insight is to restructure the causal dependency graph to allow for lossless parallel decoding, where tokens with weak cross-event dependencies can be decoded in parallel, while tightly coupled tokens within each event retain sequential decoding.

The authors introduce two key components to realize this insight: a latent global planning mechanism and an event-factorized parallel decoding mechanism. The latent global planning mechanism learns the event-level structure and produces compact tokens encoding global inter-event causality, while adaptively aggregating event-level audio-visual semantics. The event-factorized parallel decoding mechanism balances local focus with global inter-event awareness, enabling effective parallel decoding.

The proposed framework improves generation efficiency and enhances temporally grounded captioning performance. The experiments on various benchmarks demonstrate the clear advantage of the approach in both efficiency and performance in omni-modal event grounding and captioning. The results show that the parallelized autoregressive framework can generate dense captions more efficiently and accurately, making it a promising solution for large-scale video captioning tasks. Overall, the paper contributes to the development of more efficient and effective dense video captioning models, which can benefit both event-level video understanding and generation.


πŸ“… Published on Jul 3

πŸ”— Links:
β€’ GitHub: https://github.com/huggingface
β€’ arXiv: https://arxiv.org/abs/2607.02963
β€’ PDF: https://arxiv.org/pdf/2607.02963
β€’ Project Page: https://github.com/showlab/PadCaptioner

━━━━━━━━━━━━━━━━━━━━━━━━
πŸ“’ By: https://xn--r1a.website/PaperNexus

#DenseVideoCaptioning #AutoregressiveDecoding #ParallelizedInference #OmniModalLearning #TemporalEventModeling
This media is not supported in your browser
VIEW IN TELEGRAM
Shock: A Chinese laboratory has put half of the paid video generation industry in an awkward position. 😲

You upload a photo and an audio recording β€” and you get a talking avatar with lip synchronization, which can maintain quality for several minutes. And all this with open source code. πŸ€–

What used to require a camera and editing now increasingly resembles a GitHub repository. πŸ’»
It's called LongCat-Avatar. 🐱

GitHub - https://github.com/meituan-longcat/LongCat-Video
Weights - https://huggingface.co/meituan-longcat/LongCat-Video-Avatar-1.5

#AI #OpenSource #VideoGeneration #LongCat #TechNews #Innovation

✨ Join Best TG Channels https://xn--r1a.website/addlist/0f6vfFbEMdAwODBk

⭐️ Join Our WhatsApp Channel https://whatsapp.com/channel/0029VaC7Weq29753hpcggW2A
πŸš€ Looking for a portfolio-ready NLP project?

I recently published an end-to-end walkthrough on Towards Data Science using Kaggle’s Spooky Author Identification dataset.

You’ll see how far classical NLP can go with:

πŸ“ Bag-of-Words and TF-IDF
πŸ”€ Character n-grams
πŸ“Š Model comparison
🧩 Ensemble stacking

It’s a practical project for anyone preparing for an ML/DS role, with no deep learning required. I walk through the entire workflow step by step:

πŸ”— https://towardsdatascience.com/how-far-can-classical-nlp-go-from-bag-of-words-to-stacking-on-spooky-author-identification/
❀1
πŸ”₯ Infinite Worlds with Versatile Interactions

πŸ’‘ The paper presents LingBot-World 2.0, an advanced world modeling system that offers enhanced interaction capabilities and real-time processing for collaborative virtual environments. The system addresses the problem of limited interaction horizons and slow response times in previous world modeling systems. To overcome these limitations, the authors propose a carefully crafted causal pretraining paradigm that enables the model to achieve an unbounded interaction horizon while maintaining consistent output quality.

The system features four distinct upgrades. First, it achieves an unbounded interaction horizon, allowing for more complex and dynamic interactions. Second, the system guarantees rapid response time, sufficient to drive high-quality video streams. Third, it introduces a broader spectrum of interactive elements, including diverse actions and text-driven events. Fourth, the system integrates an agentic harness, where a pilot agent plans and executes character behaviors, and a director agent synthesizes novel environmental elements.

The authors demonstrate the effectiveness of their system through a 14B model and a lightweight 1.3B counterpart, which can be deployed on a single GPU. The system also features a multi-player interface, allowing multiple players to simultaneously immerse themselves in the virtual environment. The results show that the system can generate high-quality interactive elements and respond rapidly to user input, making it suitable for a wide range of applications, including collaborative virtual environments and interactive storytelling.

Overall, the paper presents a significant contribution to the field of world modeling, offering a powerful and flexible system for creating interactive and immersive virtual environments. The system's ability to achieve an unbounded interaction horizon, guarantee rapid response time, and integrate diverse interactive elements makes it a valuable tool for researchers and developers in the field of data science and artificial intelligence.


πŸ“… Published on Jul 8

πŸ”— Links:
β€’ GitHub: https://github.com/huggingface
β€’ arXiv: https://arxiv.org/abs/2607.07534
β€’ PDF: https://arxiv.org/pdf/2607.07534
β€’ Project Page: https://technology.robbyant.com/lingbot-world-v2

πŸ“Š Datasets citing this paper:
β€’ https://huggingface.co/datasets/cy0307/awesome-egocentric-atlas

━━━━━━━━━━━━━━━━━━━━━━━━
πŸ“’ By: https://xn--r1a.website/PaperNexus

#VirtualWorldModeling #CollaborativeVirtualEnvironments #CausalPretraining #AdvancedWorldModeling #RealTimeProcessing
πŸ”₯ Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

πŸ’‘ The paper introduces LingBot-Video, a video pretraining framework designed for embodied intelligence applications, such as robot control. The current video generative models are not suitable for these applications due to their focus on content creation, prioritizing visual fidelity and creativity over computational efficiency and physical realism. To address this issue, the authors propose a framework that combines a Mixture-of-Experts architecture with specialized data augmentation and a multi-dimensional reward system.

The Mixture-of-Experts architecture is used instead of a dense framework to achieve a better balance between modeling capacity and inference efficiency. The authors also construct a data profiling engine that augments standard internet videos with robot-oriented footage, including manipulation, navigation, and egocentric perspectives, to equip the base model with an understanding of actions and world dynamics.

The training process involves a multi-dimensional reward system that enforces physical rationality and task completion, in addition to standard criteria such as aesthetics and motion consistency. The authors evaluate the performance and efficiency of LingBot-Video and validate its effectiveness as a video foundation model.

The main contributions of the paper are the introduction of LingBot-Video, a large-scale, open-source Mixture-of-Experts video foundation model, and the development of a framework that bridges digital creativity and physical actuation. The authors provide a pioneering effort to address the domain mismatch between video generative models and embodied intelligence applications, and their work has the potential to improve the performance of robots in various tasks.


πŸ“… Published on Jul 8

πŸ”— Links:
β€’ GitHub: https://github.com/huggingface
β€’ arXiv: https://arxiv.org/abs/2607.07675
β€’ PDF: https://arxiv.org/pdf/2607.07675
β€’ Project Page: https://technology.robbyant.com/lingbot-video

πŸ“Š Datasets citing this paper:
β€’ https://huggingface.co/datasets/cy0307/awesome-egocentric-atlas

━━━━━━━━━━━━━━━━━━━━━━━━
πŸ“’ By: https://xn--r1a.website/PaperNexus

#MixtureOfExpertsArchitecture #EmbodiedIntelligence #VideoPretraining #RobotControlSystems #MixtureOfExpertsModels
πŸ”₯ RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies

πŸ’‘ The paper introduces RoboDojo, a unified sim-and-real benchmark for evaluating generalist robot manipulation policies. The existing benchmarks have limitations as they rely on simple, short-horizon, or skill-narrow tasks with limited capability coverage and are often conducted only in simulation or only in the real world. Simulation enables scalable feedback but misses physical deployment challenges, while real-world evaluation is costly, time-consuming, and difficult to reproduce.

RoboDojo addresses these limitations by providing a comprehensive evaluation framework that includes 42 simulation tasks and 18 real-world tasks covering diverse and complementary manipulation capabilities. The simulation benchmark evaluates five dimensions: generalization, memory, precision, long-horizon execution, and open-vocabulary instruction following, while the real-world benchmark exposes policies to challenging physical-world deployment conditions.

The RoboDojo framework supports scalable evaluation through heterogeneous parallel simulation in Isaac Sim and provides RoboDojo-RealEval, a reproducible real-world evaluation system with remote cloud access, standardized hardware, scene reset, evaluation protocol, and deployment interface. The paper also integrates 30 policies into XPolicyLab and evaluates them on RoboDojo, establishing a public leaderboard and systematic analysis of current policy performance.

The contributions of the paper are the introduction of a unified sim-and-real benchmark for evaluating generalist robot manipulation policies, a comprehensive evaluation framework that covers diverse manipulation capabilities, and a scalable and reproducible evaluation system. The paper provides a systematic analysis of current policy performance and establishes a public leaderboard, which can facilitate the development and evaluation of generalist robot manipulation policies.


πŸ“… Published on Jul 7

πŸ”— Links:
β€’ GitHub: https://github.com/huggingface
β€’ arXiv: https://arxiv.org/abs/2607.04434
β€’ PDF: https://arxiv.org/pdf/2607.04434
β€’ Project Page: https://robodojo-benchmark.com/

πŸ“Š Datasets citing this paper:
β€’ https://huggingface.co/datasets/cy0307/awesome-egocentric-atlas

━━━━━━━━━━━━━━━━━━━━━━━━
πŸ“’ By: https://xn--r1a.website/PaperNexus

#RobotManipulationPolicies #SimToRealRobotics #RobotLearningBenchmarks #GeneralistRobotics #RoboticManipulationTasks
πŸ”₯ AlayaWorld: Long-Horizon and Playable Video World Generation

πŸ’‘ The paper presents AlayaWorld, an open source framework for creating interactive generative worlds that can be used in real time. Traditionally, video game worlds are built using labor intensive production pipelines, which are costly to develop, difficult to customize, and expensive to modify after deployment. Recent advances in video world models offer a new approach, where models can synthesize future observations conditioned on the current world state and user interactions, allowing playable worlds to be generated online.

The AlayaWorld framework addresses this problem by providing a modular architecture that enables real time user interaction and supports diverse actions. The framework is trained on both gameplay recordings and real world videos, which allows it to capture diverse visual appearances and physical dynamics. This enables the creation of interactive applications beyond gaming, including embodied intelligence.

The AlayaWorld framework provides a full stack solution, including data preparation, model architecture, model training, inference acceleration, and deployment. The framework is modular and extensible, allowing users to easily customize and extend it. The authors also release reproducible pipelines, reference implementations, evaluation tools, and comprehensive documentation, making it easier for others to use and build upon the framework.

The contributions of the paper include the development of the AlayaWorld framework, which enables open ended real time interaction and allows users to freely navigate and perform diverse actions. The framework also provides a practical foundation for future research and real time applications of generative world models. Overall, the paper presents a significant contribution to the field of video world generation, enabling the creation of interactive and dynamic worlds that can be used in a variety of applications.


πŸ“… Published on Jul 7

πŸ”— Links:
β€’ GitHub: https://github.com/huggingface
β€’ arXiv: https://arxiv.org/abs/2607.06291
β€’ PDF: https://arxiv.org/pdf/2607.06291
β€’ Project Page: https://alaya-lab.github.io/AlayaWorld/

━━━━━━━━━━━━━━━━━━━━━━━━
πŸ“’ By: https://xn--r1a.website/PaperNexus

#GenerativeWorlds #VideoWorldGeneration #InteractiveWorldModels #RealTimeWorldSynthesis #PlayableWorldCreation