AI & ML Papers
33.4K subscribers
7.17K photos
556 videos
24 files
7.87K links
Advancing research in Machine Learning – practical insights, tools, and techniques for researchers.

Admin: @HusseinSheikho || @Hussein_Sheikho
Download Telegram
This media is not supported in your browser
VIEW IN TELEGRAM
TokenDial: Continuous Attribute Control in Text-to-Video via Spatiotemporal Token Offsets

📝 Summary:
TokenDial enables precise attribute control in text-to-video models by using additive offsets in spatiotemporal token space for coherent edits without retraining. AI-generated summary We present Token...

🔹 Publication Date: Published on Mar 29

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2603.27520
• PDF: https://arxiv.org/pdf/2603.27520
• Project Page: https://tokendial.github.io/
• Github: https://github.com/ariannaliu/TokenDial

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#TextToVideo #GenerativeAI #AIControl #VideoGeneration #DeepLearning
This media is not supported in your browser
VIEW IN TELEGRAM
VOID: Video Object and Interaction Deletion

📝 Summary:
VOID is a video object removal framework designed for complex scenarios involving significant object interactions. It uses vision-language and video diffusion models, leveraging causal reasoning to generate physically plausible counterfactual scenes. VOID better preserves consistent scene dynamic...

🔹 Publication Date: Published on Apr 2

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.02296
• PDF: https://arxiv.org/pdf/2604.02296
• Project Page: https://void-model.github.io/
• Github: https://github.com/Netflix/void-model

🔹 Models citing this paper:
https://huggingface.co/netflix/void-model

Spaces citing this paper:
https://huggingface.co/spaces/sam-motamed/VOID

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#VideoEditing #DiffusionModels #ComputerVision #GenerativeAI #DeepLearning
Salt: Self-Consistent Distribution Matching with Cache-Aware Training for Fast Video Generation

📝 Summary:
The paper introduces Salt, a method for fast video generation. It proposes Self-Consistent Distribution Matching Distillation SC-DMD to improve low-NFE quality by regularizing denoising updates. Cache-Distribution-Aware training further optimizes real-time autoregressive generation using KV cache.

🔹 Publication Date: Published on Apr 3

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.03118
• PDF: https://arxiv.org/pdf/2604.03118
• Github: https://github.com/XingtongGe/Salt

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#VideoGeneration #GenerativeAI #DeepLearning #AIResearch #RealTimeAI
Media is too big
VIEW IN TELEGRAM
AvatarPointillist: AutoRegressive 4D Gaussian Avatarization

📝 Summary:
AvatarPointillist creates dynamic 4D Gaussian avatars from a single image using an autoregressive Transformer. It builds point clouds with adaptive density and binding info for realistic animation, producing high-quality, controllable results.

🔹 Publication Date: Published on Apr 6

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.04787
• PDF: https://arxiv.org/pdf/2604.04787
• Project Page: https://kumapowerliu.github.io/AvatarPointillist/
• Github: https://github.com/KumapowerLIU/AvatarPointillist

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#AI #ComputerVision #3DAvatars #GenerativeAI #MachineLearning
Think in Strokes, Not Pixels: Process-Driven Image Generation via Interleaved Reasoning

📝 Summary:
This paper introduces process-driven image generation, an iterative method with interleaved textual and visual reasoning. It decomposes synthesis into planning, drafting, reflecting, and refining steps. Dense step-wise supervision ensures consistency and interpretability of intermediate states.

🔹 Publication Date: Published on Apr 8

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.04746
• PDF: https://arxiv.org/pdf/2604.04746

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#ImageGeneration #GenerativeAI #ArtificialIntelligence #DeepLearning #ComputerVision
ViVa: A Video-Generative Value Model for Robot Reinforcement Learning

📝 Summary:
ViVa is a video-generative value model for robot reinforcement learning. It estimates values by leveraging pretrained video generators to predict future robot dynamics, moving beyond static observations. This approach improves robot manipulation and generalizes to novel objects.

🔹 Publication Date: Published on Apr 9

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.08168
• PDF: https://arxiv.org/pdf/2604.08168
• Project Page: https://viva-value-model.github.io/
• Github: https://github.com/GigaAI-research/ViVa

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#Robotics #ReinforcementLearning #GenerativeAI #MachineLearning #AI
RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details

📝 Summary:
RefineAnything is a multimodal diffusion model for region-specific image refinement. It fixes local detail collapse while strictly preserving backgrounds using a Focus-and-Refine strategy and boundary-aware loss. This provides a practical solution for high-precision local editing.

🔹 Publication Date: Published on Apr 8

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.06870
• PDF: https://arxiv.org/pdf/2604.06870
• Project Page: https://limuloo.github.io/RefineAnything/
• Github: https://github.com/limuloo/RefineAnything

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#DiffusionModels #ImageEditing #ComputerVision #DeepLearning #GenerativeAI
Media is too big
VIEW IN TELEGRAM
Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory

📝 Summary:
Matrix-Game 3.0 is a memory-augmented diffusion model achieving real-time 720p interactive video generation with long-term temporal consistency. It uses an advanced data engine, a self-correction training framework with memory, and efficient inference strategies. This enables practical, industria...

🔹 Publication Date: Published on Apr 10

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.08995
• PDF: https://arxiv.org/pdf/2604.08995
• Project Page: https://matrix-game-v3.github.io/

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#DiffusionModels #VideoGeneration #RealTimeAI #GenerativeAI #MachineLearning
MixFlow: Mixed Source Distributions Improve Rectified Flows

📝 Summary:
Rectified flows and diffusion models are improved through κ-FC formulation that conditions the source distribution and MixFlow training strategy that reduces generative path curvatures and enhances sa...

🔹 Publication Date: Published on Apr 10

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.09181
• PDF: https://arxiv.org/pdf/2604.09181
• Github: https://github.com/NazirNayal8/MixFlow

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#RectifiedFlows #DiffusionModels #GenerativeAI #MachineLearning #AIResearch
This media is not supported in your browser
VIEW IN TELEGRAM
Narrative-Driven Paper-to-Slide Generation via ArcDeck

📝 Summary:
ArcDeck is a multi-agent framework for paper-to-slide generation that models a paper's logical flow through discourse trees. It uses an iterative refinement process to ensure narrative coherence and improve presentations over direct summarization methods.

🔹 Publication Date: Published on Apr 13

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.11969
• PDF: https://arxiv.org/pdf/2604.11969
• Project Page: https://arcdeck.org/
• Github: https://github.com/RehgLab/ArcDeck

Datasets citing this paper:
https://huggingface.co/datasets/ArcDeck/ArcBench

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#AIPresentations #NLP #GenerativeAI #ResearchTools #AcademicPublishing
1
Qwen3.5-Omni Technical Report

📝 Summary:
Qwen3.5-Omni is a large multimodal model excelling in audio-visual understanding and generation, achieving SOTA results across many benchmarks. It features a Hybrid Attention MoE architecture, introduces ARIA for improved speech synthesis, and exhibits a new Audio-Visual Vibe Coding capability.

🔹 Publication Date: Published on Apr 17

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.15804
• PDF: https://arxiv.org/pdf/2604.15804

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#MultimodalAI #AIResearch #DeepLearning #GenerativeAI #SpeechSynthesis
This media is not supported in your browser
VIEW IN TELEGRAM
Repurposing 3D Generative Model for Autoregressive Layout Generation

📝 Summary:
LaviGen is a 3D layout generation framework that repurposes 3D generative models. It uses an adapted 3D diffusion model for autoregressive generation, explicitly modeling geometric relations and physical constraints. This achieves superior, more plausible 3D layouts 65% faster than previous methods.

🔹 Publication Date: Published on Apr 17

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.16299
• PDF: https://arxiv.org/pdf/2604.16299
• Project Page: https://fenghora.github.io/LaviGen-Page/
• Github: https://github.com/fenghora/LaviGen

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#3DGeneration #DiffusionModels #GenerativeAI #ComputerGraphics #DeepLearning
Media is too big
VIEW IN TELEGRAM
Hierarchical Codec Diffusion for Video-to-Speech Generation

📝 Summary:
HiCoDiT generates speech from videos by leveraging the hierarchical structure of discrete speech tokens, achieving better audio-visual alignment through coarse-to-fine conditioning with dual-scale nor...

🔹 Publication Date: Published on Apr 17

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.15923
• PDF: https://arxiv.org/pdf/2604.15923

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#VideoToSpeech #DiffusionModels #GenerativeAI #SpeechSynthesis #DeepLearning
UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models

📝 Summary:
UDM-GRPO integrates Uniform Discrete Diffusion Models with reinforcement learning, solving training instability issues. It optimizes using final samples as actions and reconstructed trajectories. This achieves state-of-the-art performance in text-to-image generation and OCR tasks.

🔹 Publication Date: Published on Apr 20

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.18518
• PDF: https://arxiv.org/pdf/2604.18518
• Project Page: https://yovecent.github.io/UDM-GRPO.github.io/
• Github: https://github.com/Yovecent/UDM-GRPO

🔹 Models citing this paper:
https://huggingface.co/Yovecents/URSA-1.7B-IBQ512-UDMGRPO-GenEval
https://huggingface.co/Yovecents/URSA-1.7B-IBQ512-UDMGRPO-PickScore

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#DiffusionModels #ReinforcementLearning #GenerativeAI #TextToImage #DeepLearning
1
Media is too big
VIEW IN TELEGRAM
CityRAG: Stepping Into a City via Spatially-Grounded Video Generation

📝 Summary:
CityRAG generates long-term, physically grounded video sequences that maintain environmental consistency and support complex navigation through real-world geography using geo-registered data as contex...

🔹 Publication Date: Published on Apr 21

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.19741
• PDF: https://arxiv.org/pdf/2604.19741
• Project Page: https://cityrag.github.io/

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#VideoGeneration #GenerativeAI #SpatialAI #ComputerVision #UrbanSimulation
1
FlowAnchor: Stabilizing the Editing Signal for Inversion-Free Video Editing

📝 Summary:
FlowAnchor stabilizes inversion-free video editing by addressing signal instability in high-dimensional latent spaces. It uses spatial-aware attention refinement and adaptive magnitude modulation to ensure precise localization and sufficient editing strength, leading to faithful and coherent vide...

🔹 Publication Date: Published on Apr 24

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.22586
• PDF: https://arxiv.org/pdf/2604.22586

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#VideoEditing #DeepLearning #ComputerVision #GenerativeAI #AIResearch
This media is not supported in your browser
VIEW IN TELEGRAM
Video Analysis and Generation via a Semantic Progress Function

📝 Summary:
Researchers developed a Semantic Progress Function to analyze and correct non-linear semantic evolution in generated media. This function identifies uneven pacing, enabling a linearization procedure that re-times sequences for smoother, more coherent transitions at a constant semantic rate.

🔹 Publication Date: Published on Apr 24

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.22554
• PDF: https://arxiv.org/pdf/2604.22554
• Project Page: https://sagipolaczek.github.io/semantic-progress-function/
• Github: https://github.com/SagiPolaczek/semantic-progress-function

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#VideoAI #GenerativeAI #ComputerVision #SemanticAnalysis #AIResearch
MAIC-UI: Making Interactive Courseware with Generative UI

📝 Summary:
MAIC-UI is a zero-code system for educators to rapidly create and edit interactive STEM courseware using structured knowledge analysis and incremental generation. It significantly improves editing efficiency, student learnability, and STEM learning outcomes.

🔹 Publication Date: Published on Apr 28

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.25806
• PDF: https://arxiv.org/pdf/2604.25806
• Project Page: https://open.maic.chat/
• Github: https://github.com/THU-MAIC/MAIC-UI

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#EdTech #GenerativeAI #STEMEducation #NoCode #Courseware
Large Language Models Explore by Latent Distilling

📝 Summary:
Exploratory Sampling ESamp boosts LLM diversity beyond lexical variation. It uses a lightweight Distiller to predict hidden representations, biasing decoding towards novel semantic patterns via prediction error. ESamp boosts reasoning efficiency and creative writing, with low overhead.

🔹 Publication Date: Published on Apr 27

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.24927
• PDF: https://arxiv.org/pdf/2604.24927
• Github: https://github.com/LinesHogan/tllm

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#LLM #AI #NLP #DeepLearning #GenerativeAI
1
Instruction-Guided Poetry Generation in Arabic and Its Dialects

📝 Summary:
A new instruction-based dataset and fine-tuned LLMs enable controllable Arabic poetry generation across Modern Standard Arabic and dialects. This work allows users to create, revise, and continue poems effectively, moving beyond just analysis, as confirmed by strong evaluations.

🔹 Publication Date: Published on Apr 30

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.27766
• PDF: https://arxiv.org/pdf/2604.27766
• Github: https://github.com/mbzuai-nlp/instructpoet-ar

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#LLM #NLP #ArabicAI #GenerativeAI #PoetryGeneration