✨InstructMix2Mix: Consistent Sparse-View Editing Through Multi-View Model Personalization
📝 Summary:
InstructMix2Mix I-Mix2Mix improves multi-view image editing from sparse inputs, which often lack consistency. It distills a 2D diffusion model into a multi-view diffusion model, leveraging its 3D prior for cross-view coherence. This framework significantly enhances multi-view consistency and per-...
🔹 Publication Date: Published on Nov 18
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2511.14899
• PDF: https://arxiv.org/pdf/2511.14899
• Project Page: https://danielgilo.github.io/instruct-mix2mix/
• Github: https://danielgilo.github.io/instruct-mix2mix/
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#MultiViewEditing #DiffusionModels #ComputerVision #3DVision #ImageSynthesis
📝 Summary:
InstructMix2Mix I-Mix2Mix improves multi-view image editing from sparse inputs, which often lack consistency. It distills a 2D diffusion model into a multi-view diffusion model, leveraging its 3D prior for cross-view coherence. This framework significantly enhances multi-view consistency and per-...
🔹 Publication Date: Published on Nov 18
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2511.14899
• PDF: https://arxiv.org/pdf/2511.14899
• Project Page: https://danielgilo.github.io/instruct-mix2mix/
• Github: https://danielgilo.github.io/instruct-mix2mix/
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#MultiViewEditing #DiffusionModels #ComputerVision #3DVision #ImageSynthesis
❤1
✨SUCCESS-GS: Survey of Compactness and Compression for Efficient Static and Dynamic Gaussian Splatting
📝 Summary:
This survey overviews efficient 3D and 4D Gaussian Splatting. It categorizes parameter and restructuring compression methods to reduce memory and computation while maintaining reconstruction quality. It also covers current limitations and future research.
🔹 Publication Date: Published on Dec 8
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2512.07197
• PDF: https://arxiv.org/pdf/2512.07197
• Project Page: https://cmlab-korea.github.io/Awesome-Efficient-GS/
• Github: https://cmlab-korea.github.io/Awesome-Efficient-GS/
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#GaussianSplatting #3DVision #ComputerGraphics #DeepLearning #Efficiency
📝 Summary:
This survey overviews efficient 3D and 4D Gaussian Splatting. It categorizes parameter and restructuring compression methods to reduce memory and computation while maintaining reconstruction quality. It also covers current limitations and future research.
🔹 Publication Date: Published on Dec 8
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2512.07197
• PDF: https://arxiv.org/pdf/2512.07197
• Project Page: https://cmlab-korea.github.io/Awesome-Efficient-GS/
• Github: https://cmlab-korea.github.io/Awesome-Efficient-GS/
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#GaussianSplatting #3DVision #ComputerGraphics #DeepLearning #Efficiency
✨Sharp Monocular View Synthesis in Less Than a Second
📝 Summary:
SHARP synthesizes photorealistic 3D views from a single image using a 3D Gaussian representation. It achieves state-of-the-art quality with rapid processing, taking less than a second, and supports metric camera movements.
🔹 Publication Date: Published on Dec 11
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2512.10685
• PDF: https://arxiv.org/pdf/2512.10685
• Project Page: https://apple.github.io/ml-sharp/
• Github: https://github.com/apple/ml-sharp
🔹 Models citing this paper:
• https://huggingface.co/apple/Sharp
✨ Spaces citing this paper:
• https://huggingface.co/spaces/ronedgecomb/ml-sharp
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#ViewSynthesis #3DVision #ComputerVision #RealtimeAI #GaussianSplats
📝 Summary:
SHARP synthesizes photorealistic 3D views from a single image using a 3D Gaussian representation. It achieves state-of-the-art quality with rapid processing, taking less than a second, and supports metric camera movements.
🔹 Publication Date: Published on Dec 11
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2512.10685
• PDF: https://arxiv.org/pdf/2512.10685
• Project Page: https://apple.github.io/ml-sharp/
• Github: https://github.com/apple/ml-sharp
🔹 Models citing this paper:
• https://huggingface.co/apple/Sharp
✨ Spaces citing this paper:
• https://huggingface.co/spaces/ronedgecomb/ml-sharp
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#ViewSynthesis #3DVision #ComputerVision #RealtimeAI #GaussianSplats
❤1
✨Dream2Flow: Bridging Video Generation and Open-World Manipulation with 3D Object Flow
📝 Summary:
Dream2Flow bridges video generation and robotic control using 3D object flow. It reconstructs 3D object motions from generated videos, enabling zero-shot manipulation of diverse objects through trajectory tracking without task-specific demonstrations.
🔹 Publication Date: Published on Dec 31, 2025
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2512.24766
• PDF: https://arxiv.org/pdf/2512.24766
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#VideoGeneration #Robotics #3DVision #AI #ZeroShotLearning
📝 Summary:
Dream2Flow bridges video generation and robotic control using 3D object flow. It reconstructs 3D object motions from generated videos, enabling zero-shot manipulation of diverse objects through trajectory tracking without task-specific demonstrations.
🔹 Publication Date: Published on Dec 31, 2025
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2512.24766
• PDF: https://arxiv.org/pdf/2512.24766
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#VideoGeneration #Robotics #3DVision #AI #ZeroShotLearning
✨InfiniteVGGT: Visual Geometry Grounded Transformer for Endless Streams
📝 Summary:
InfiniteVGGT enables continuous 3D visual geometry understanding for infinite streams. It uses a causal transformer with adaptive rolling memory for long-term stability, outperforming existing streaming methods. A new Long3D benchmark is introduced for rigorous evaluation of such systems.
🔹 Publication Date: Published on Jan 5
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2601.02281
• PDF: https://arxiv.org/pdf/2601.02281
• Github: https://github.com/AutoLab-SAI-SJTU/InfiniteVGGT
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#VisualGeometry #3DVision #Transformers #StreamingAI #DeepLearning
📝 Summary:
InfiniteVGGT enables continuous 3D visual geometry understanding for infinite streams. It uses a causal transformer with adaptive rolling memory for long-term stability, outperforming existing streaming methods. A new Long3D benchmark is introduced for rigorous evaluation of such systems.
🔹 Publication Date: Published on Jan 5
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2601.02281
• PDF: https://arxiv.org/pdf/2601.02281
• Github: https://github.com/AutoLab-SAI-SJTU/InfiniteVGGT
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#VisualGeometry #3DVision #Transformers #StreamingAI #DeepLearning
✨MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources
📝 Summary:
Metric Anything introduces a scalable pretraining framework for metric depth using Sparse Metric Prompts to handle diverse, noisy 3D data. It shows clear scaling trends and achieves state-of-the-art performance across various depth estimation and spatial intelligence tasks.
🔹 Publication Date: Published on Jan 29
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2601.22054
• PDF: https://arxiv.org/pdf/2601.22054
• Project Page: https://metric-anything.github.io/metric-anything-io/
• Github: https://github.com/metric-anything/metric-anything
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#MetricDepth #ComputerVision #MachineLearning #DeepLearning #3DVision
📝 Summary:
Metric Anything introduces a scalable pretraining framework for metric depth using Sparse Metric Prompts to handle diverse, noisy 3D data. It shows clear scaling trends and achieves state-of-the-art performance across various depth estimation and spatial intelligence tasks.
🔹 Publication Date: Published on Jan 29
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2601.22054
• PDF: https://arxiv.org/pdf/2601.22054
• Project Page: https://metric-anything.github.io/metric-anything-io/
• Github: https://github.com/metric-anything/metric-anything
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#MetricDepth #ComputerVision #MachineLearning #DeepLearning #3DVision
✨SpatialBoost: Enhancing Visual Representation through Language-Guided Reasoning
📝 Summary:
SpatialBoost improves the 3D spatial awareness of vision encoders by integrating linguistic 3D spatial knowledge. It achieves this through a multi-turn Chain-of-Thought reasoning process using Large Language Models, converting 3D spatial information from 2D images into linguistic descriptions. Th...
🔹 Publication Date: Published on Mar 23
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2603.22057
• PDF: https://arxiv.org/pdf/2603.22057
• Project Page: https://rootyjeon.github.io/spatial-boost/
• Github: https://github.com/rootyJeon/SpatialBoost
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#SpatialBoost #ComputerVision #LLM #3DVision #AI
📝 Summary:
SpatialBoost improves the 3D spatial awareness of vision encoders by integrating linguistic 3D spatial knowledge. It achieves this through a multi-turn Chain-of-Thought reasoning process using Large Language Models, converting 3D spatial information from 2D images into linguistic descriptions. Th...
🔹 Publication Date: Published on Mar 23
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2603.22057
• PDF: https://arxiv.org/pdf/2603.22057
• Project Page: https://rootyjeon.github.io/spatial-boost/
• Github: https://github.com/rootyJeon/SpatialBoost
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#SpatialBoost #ComputerVision #LLM #3DVision #AI
✨One View Is Enough! Monocular Training for In-the-Wild Novel View Generation
📝 Summary:
OVIE enables monocular novel-view synthesis from single images by generating pseudo-target views via a geometric scaffold. This eliminates the need for multi-view supervision, allowing training on massive unpaired datasets. OVIE achieves superior zero-shot performance and is significantly faster ...
🔹 Publication Date: Published on Mar 24
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2603.23488
• PDF: https://arxiv.org/pdf/2603.23488
• Github: https://github.com/AdrienRR/ovie
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#NovelViewSynthesis #MonocularVision #ComputerVision #DeepLearning #3DVision
📝 Summary:
OVIE enables monocular novel-view synthesis from single images by generating pseudo-target views via a geometric scaffold. This eliminates the need for multi-view supervision, allowing training on massive unpaired datasets. OVIE achieves superior zero-shot performance and is significantly faster ...
🔹 Publication Date: Published on Mar 24
🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2603.23488
• PDF: https://arxiv.org/pdf/2603.23488
• Github: https://github.com/AdrienRR/ovie
==================================
For more data science resources:
✓ https://xn--r1a.website/DataScienceT
#NovelViewSynthesis #MonocularVision #ComputerVision #DeepLearning #3DVision
❤1