✨Structured Causal Video Reasoning via Multi-Objective Alignment
📝 Summary:
This paper introduces Structured Event Facts for explicit causal video reasoning, moving beyond unstructured methods. It uses a multi-objective reinforcement learning pipeline to balance training goals, leading to Factum-4B. This model achieves reliable, stronger performance on complex temporal v...
🔹 Publication Date: Published on Apr 6
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.04415
• PDF: https://arxiv.org/pdf/2604.04415
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#CausalAI #VideoReasoning #ReinforcementLearning #ComputerVision #AIResearch
📝 Summary:
This paper introduces Structured Event Facts for explicit causal video reasoning, moving beyond unstructured methods. It uses a multi-objective reinforcement learning pipeline to balance training goals, leading to Factum-4B. This model achieves reliable, stronger performance on complex temporal v...
🔹 Publication Date: Published on Apr 6
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.04415
• PDF: https://arxiv.org/pdf/2604.04415
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#CausalAI #VideoReasoning #ReinforcementLearning #ComputerVision #AIResearch
✨3DTV: A Feedforward Interpolation Network for Real-Time View Synthesis
📝 Summary:
3DTV is a feedforward network combining lightweight geometry and learning for real-time, robust sparse-view interpolation. It generates novel views efficiently without scene-specific optimization, making it practical for interactive applications.
🔹 Publication Date: Published on Apr 13
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.11211
• PDF: https://arxiv.org/pdf/2604.11211
• Project Page: https://stefanmschulz.github.io/3DTV_webpage/
• Github: https://github.com/StefanMSchulz/3DTV
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#ViewSynthesis #DeepLearning #ComputerVision #NeuralNetworks #RealTimeAI
📝 Summary:
3DTV is a feedforward network combining lightweight geometry and learning for real-time, robust sparse-view interpolation. It generates novel views efficiently without scene-specific optimization, making it practical for interactive applications.
🔹 Publication Date: Published on Apr 13
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.11211
• PDF: https://arxiv.org/pdf/2604.11211
• Project Page: https://stefanmschulz.github.io/3DTV_webpage/
• Github: https://github.com/StefanMSchulz/3DTV
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#ViewSynthesis #DeepLearning #ComputerVision #NeuralNetworks #RealTimeAI
✨ReconPhys: Reconstruct Appearance and Physical Attributes from Single Video
📝 Summary:
ReconPhys is the first feedforward framework to jointly learn physical attribute estimation and 3D Gaussian Splatting reconstruction from a single video. It offers significantly faster inference and superior reconstruction quality for non-rigid objects compared to prior optimization-based methods...
🔹 Publication Date: Published on Apr 9
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.07882
• PDF: https://arxiv.org/pdf/2604.07882
• Project Page: https://chuanshuogushi.github.io/ReconPhys/
• Github: https://chuanshuogushi.github.io/ReconPhys/
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#ComputerVision #3DReconstruction #GaussianSplatting #DeepLearning #AIResearch
📝 Summary:
ReconPhys is the first feedforward framework to jointly learn physical attribute estimation and 3D Gaussian Splatting reconstruction from a single video. It offers significantly faster inference and superior reconstruction quality for non-rigid objects compared to prior optimization-based methods...
🔹 Publication Date: Published on Apr 9
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.07882
• PDF: https://arxiv.org/pdf/2604.07882
• Project Page: https://chuanshuogushi.github.io/ReconPhys/
• Github: https://chuanshuogushi.github.io/ReconPhys/
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#ComputerVision #3DReconstruction #GaussianSplatting #DeepLearning #AIResearch
✨VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects
📝 Summary:
VEFX-Bench offers a large human-annotated video editing dataset and VEFX-Reward, a specialized model for quality assessment. This benchmark allows standardized comparison, showing current models struggle with instruction following and edit locality.
🔹 Publication Date: Published on Apr 17
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.16272
• PDF: https://arxiv.org/pdf/2604.16272
• Project Page: https://xiangbogaobarry.github.io/VEFX-Bench/
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#VideoEditing #VFX #AI #ComputerVision #Benchmarks
📝 Summary:
VEFX-Bench offers a large human-annotated video editing dataset and VEFX-Reward, a specialized model for quality assessment. This benchmark allows standardized comparison, showing current models struggle with instruction following and edit locality.
🔹 Publication Date: Published on Apr 17
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.16272
• PDF: https://arxiv.org/pdf/2604.16272
• Project Page: https://xiangbogaobarry.github.io/VEFX-Bench/
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#VideoEditing #VFX #AI #ComputerVision #Benchmarks
✨NTIRE 2026 Challenge on Video Saliency Prediction: Methods and Results
📝 Summary:
This paper overviews the NTIRE 2026 Challenge on Video Saliency Prediction. Participants developed automatic saliency map prediction for videos using a novel 2,000-video dataset with crowdsourced fixations. Over 20 teams submitted, and all challenge data is now publicly available.
🔹 Publication Date: Published on Apr 16
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.14816
• PDF: https://arxiv.org/pdf/2604.14816
• Project Page: https://www.codabench.org/competitions/12842/
• Github: https://github.com/msu-video-group/NTIRE26_Saliency_Prediction
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#VideoSaliency #ComputerVision #NTIRE #MachineLearning #SaliencyPrediction
📝 Summary:
This paper overviews the NTIRE 2026 Challenge on Video Saliency Prediction. Participants developed automatic saliency map prediction for videos using a novel 2,000-video dataset with crowdsourced fixations. Over 20 teams submitted, and all challenge data is now publicly available.
🔹 Publication Date: Published on Apr 16
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.14816
• PDF: https://arxiv.org/pdf/2604.14816
• Project Page: https://www.codabench.org/competitions/12842/
• Github: https://github.com/msu-video-group/NTIRE26_Saliency_Prediction
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#VideoSaliency #ComputerVision #NTIRE #MachineLearning #SaliencyPrediction
✨Concrete Jungle: Towards Concreteness Paved Contrastive Negative Mining for Compositional Understanding
📝 Summary:
This paper improves vision-language models for compositional reasoning by using concreteness-based negative sample selection and a novel margin-based loss. Their framework, Slipform, achieves state-of-the-art accuracy on compositional benchmarks and cross-modal retrieval.
🔹 Publication Date: Published on Apr 14
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.13313
• PDF: https://arxiv.org/pdf/2604.13313
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#VisionLanguage #DeepLearning #AIResearch #ComputerVision #NLP
📝 Summary:
This paper improves vision-language models for compositional reasoning by using concreteness-based negative sample selection and a novel margin-based loss. Their framework, Slipform, achieves state-of-the-art accuracy on compositional benchmarks and cross-modal retrieval.
🔹 Publication Date: Published on Apr 14
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.13313
• PDF: https://arxiv.org/pdf/2604.13313
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#VisionLanguage #DeepLearning #AIResearch #ComputerVision #NLP
Media is too big
VIEW IN TELEGRAM
✨CityRAG: Stepping Into a City via Spatially-Grounded Video Generation
📝 Summary:
CityRAG generates long-term, physically grounded video sequences that maintain environmental consistency and support complex navigation through real-world geography using geo-registered data as contex...
🔹 Publication Date: Published on Apr 21
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.19741
• PDF: https://arxiv.org/pdf/2604.19741
• Project Page: https://cityrag.github.io/
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#VideoGeneration #GenerativeAI #SpatialAI #ComputerVision #UrbanSimulation
📝 Summary:
CityRAG generates long-term, physically grounded video sequences that maintain environmental consistency and support complex navigation through real-world geography using geo-registered data as contex...
🔹 Publication Date: Published on Apr 21
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.19741
• PDF: https://arxiv.org/pdf/2604.19741
• Project Page: https://cityrag.github.io/
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#VideoGeneration #GenerativeAI #SpatialAI #ComputerVision #UrbanSimulation
❤1
This media is not supported in your browser
VIEW IN TELEGRAM
✨DeVI: Physics-based Dexterous Human-Object Interaction via Synthetic Video Imitation
📝 Summary:
DeVI enables physically plausible dexterous robot control by leveraging text-conditioned synthetic videos through a hybrid tracking reward that combines 3D and 2D tracking for improved hand-object int...
🔹 Publication Date: Published on Apr 22
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.20841
• PDF: https://arxiv.org/pdf/2604.20841
• Project Page: https://snuvclab.github.io/devi/
• Github: https://github.com/snuvclab/devi
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#Robotics #AI #ComputerVision #HumanRobotInteraction #DeepLearning
📝 Summary:
DeVI enables physically plausible dexterous robot control by leveraging text-conditioned synthetic videos through a hybrid tracking reward that combines 3D and 2D tracking for improved hand-object int...
🔹 Publication Date: Published on Apr 22
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.20841
• PDF: https://arxiv.org/pdf/2604.20841
• Project Page: https://snuvclab.github.io/devi/
• Github: https://github.com/snuvclab/devi
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#Robotics #AI #ComputerVision #HumanRobotInteraction #DeepLearning
✨3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decoding
📝 Summary:
3D-VCD is a new inference-time framework that reduces hallucinations in 3D embodied agents. It constructs distorted 3D scene graphs and contrasts predictions to suppress ungrounded tokens. This improves reasoning on 3D benchmarks without retraining.
🔹 Publication Date: Published on Apr 9
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.08645
• PDF: https://arxiv.org/pdf/2604.08645
• Project Page: https://plan-lab.github.io/projects/3d-vcd
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#3DLLM #EmbodiedAI #HallucinationMitigation #ComputerVision #AIResearch
📝 Summary:
3D-VCD is a new inference-time framework that reduces hallucinations in 3D embodied agents. It constructs distorted 3D scene graphs and contrasts predictions to suppress ungrounded tokens. This improves reasoning on 3D benchmarks without retraining.
🔹 Publication Date: Published on Apr 9
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.08645
• PDF: https://arxiv.org/pdf/2604.08645
• Project Page: https://plan-lab.github.io/projects/3d-vcd
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#3DLLM #EmbodiedAI #HallucinationMitigation #ComputerVision #AIResearch
arXiv.org
3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through...
Large multimodal models are increasingly used as the reasoning core of embodied agents operating in 3D environments, yet they remain prone to hallucinations that can produce unsafe and ungrounded...
✨FlowAnchor: Stabilizing the Editing Signal for Inversion-Free Video Editing
📝 Summary:
FlowAnchor stabilizes inversion-free video editing by addressing signal instability in high-dimensional latent spaces. It uses spatial-aware attention refinement and adaptive magnitude modulation to ensure precise localization and sufficient editing strength, leading to faithful and coherent vide...
🔹 Publication Date: Published on Apr 24
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.22586
• PDF: https://arxiv.org/pdf/2604.22586
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#VideoEditing #DeepLearning #ComputerVision #GenerativeAI #AIResearch
📝 Summary:
FlowAnchor stabilizes inversion-free video editing by addressing signal instability in high-dimensional latent spaces. It uses spatial-aware attention refinement and adaptive magnitude modulation to ensure precise localization and sufficient editing strength, leading to faithful and coherent vide...
🔹 Publication Date: Published on Apr 24
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.22586
• PDF: https://arxiv.org/pdf/2604.22586
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#VideoEditing #DeepLearning #ComputerVision #GenerativeAI #AIResearch
This media is not supported in your browser
VIEW IN TELEGRAM
✨Video Analysis and Generation via a Semantic Progress Function
📝 Summary:
Researchers developed a Semantic Progress Function to analyze and correct non-linear semantic evolution in generated media. This function identifies uneven pacing, enabling a linearization procedure that re-times sequences for smoother, more coherent transitions at a constant semantic rate.
🔹 Publication Date: Published on Apr 24
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.22554
• PDF: https://arxiv.org/pdf/2604.22554
• Project Page: https://sagipolaczek.github.io/semantic-progress-function/
• Github: https://github.com/SagiPolaczek/semantic-progress-function
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#VideoAI #GenerativeAI #ComputerVision #SemanticAnalysis #AIResearch
📝 Summary:
Researchers developed a Semantic Progress Function to analyze and correct non-linear semantic evolution in generated media. This function identifies uneven pacing, enabling a linearization procedure that re-times sequences for smoother, more coherent transitions at a constant semantic rate.
🔹 Publication Date: Published on Apr 24
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.22554
• PDF: https://arxiv.org/pdf/2604.22554
• Project Page: https://sagipolaczek.github.io/semantic-progress-function/
• Github: https://github.com/SagiPolaczek/semantic-progress-function
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#VideoAI #GenerativeAI #ComputerVision #SemanticAnalysis #AIResearch