AI with Papers - Artificial Intelligence & Deep Learning

🌻MLLMs Fine Segmentation🌻

👉SimpleSeg: MLLMs with native pixel-level perception. Repo & Model available💙

👉Review https://t.ly/eVguh
👉Paper arxiv.org/pdf/2601.19228
👉Project simpleseg.github.io/
👉Repo github.com/songtianhui/SimpleSeg

🔥4👍3❤2👏1

3.52K viewsedited 07:39

🔥 DeepSeek-OCR 2 is out 🔥

👉DeepSeek-AI announced the new version of its powerful SOTA OCR. A new architectural approach with the potential to achieve genuine 2D reasoning. Codes & weights💙

👉Review https://t.ly/gX4bX
👉Paper https://arxiv.org/pdf/2601.20552
👉Repo github.com/deepseek-ai/DeepSeek-OCR-2

🔥7❤4👏1

3.35K views07:33

AI with Papers - Artificial Intelligence & Deep Learning

This media is not supported in your browser

VIEW IN TELEGRAM

📊 SOTA Style Transfer 📊

👉TeleAI unveils TeleStyle, a lightweight yet effective model for image/video stylization. Built upon Qwen-Image-Edit, TeleStyle leverages the base model’s robust capabilities in content preservation & style customization. Code & Model released💙

👉Review https://t.ly/viVR0
👉Paper arxiv.org/pdf/2601.20175
👉Project tele-ai.github.io/TeleStyle/
👉Repo github.com/Tele-AI/TeleStyle

❤10👍2🔥1🤯1🤣1

3.58K views13:01

AI with Papers - Artificial Intelligence & Deep Learning

This media is not supported in your browser

VIEW IN TELEGRAM

🍑 Metric Anything is out 🍑

👉Metric Anything (Li Auto inc.) is a simple and scalable pretraining framework that learns metric depth from noisy, diverse 3D sources without manually engineered prompts, camera-specific modeling, or task-specific architectures. Impressive. Code announced 💙

👉Review https://t.ly/54Ccr
👉Paper arxiv.org/pdf/2601.22054
👉Project metric-anything.github.io/metric-anything-io/
👉Repo github.com/metric-anything/metric-anything

🔥10❤5👏1

3.65K views08:02

AI with Papers - Artificial Intelligence & Deep Learning

Still in love with this channel?

Anonymous Poll

❤6

320 voters3.25K views21:37

AI with Papers - Artificial Intelligence & Deep Learning

The hottest website on Earth right now: https://www.moltbook.com

What do you think about it?

moltbook

moltbook - the front page of the agent internet

A social network built exclusively for AI agents. Where AI agents share, discuss, and upvote. 🦞🤖

❤1🤯1😢1

3.72K views22:03

AI with Papers - Artificial Intelligence & Deep Learning

This media is not supported in your browser

VIEW IN TELEGRAM

🌈Segment Any Events by Language🌈

👉SEAL (by NUS) is the first Semantic-aware Segment Any Events framework that addresses Open-Vocabulary Event Instance Segmentation. Code announced💙

👉Review https://t.ly/1ZMF0
👉Paper https://arxiv.org/pdf/2601.23159
👉Project https://0nandon.github.io/SEAL/
👉Repo https://github.com/0nandon/SEAL

🔥7❤4👏1🤯1

3.54K views08:06

AI with Papers - Artificial Intelligence & Deep Learning

👉RAM prices skyrocketing

👉Me acting like a rich kid.

Let's talk: https://www.linkedin.com/posts/visionarynet_ai-ram-ddr5-activity-7424127924020072448-NbaO

🤣21❤4🔥1👏1

3.54K viewsedited 16:36

AI with Papers - Artificial Intelligence & Deep Learning

This media is not supported in your browser

VIEW IN TELEGRAM

🐮CoWTracker: Track-Warping🐮

👉CoWTracker (VGG + META) is a novel dense point tracker that eschews cost volumes in favor of warping. Code/Models under FAIR NC💙

👉Review https://t.ly/6bAn9
👉Paper https://arxiv.org/pdf/2602.04877
👉Project https://cowtracker.github.io/
👉Repo https://github.com/facebookresearch/cowtracker

🔥4❤2

3.25K views07:12

AI with Papers - Artificial Intelligence & Deep Learning

0:02

This media is not supported in your browser

VIEW IN TELEGRAM

🌈TrajVG Trajectory-Geometry🌈

👉TrajVG is a novel reconstruction framework that makes cross-frame 3D correspondence an explicit prediction by estimating camera-coordinate 3D trajectories. Code announced💙

👉Review https://t.ly/yVi01
👉Paper arxiv.org/pdf/2602.04439
👉Project xingy038.github.io/TrajVG/
👉Repo github.com/xingy038/TrajVG

❤7🔥1👏1

3.44K viewsedited 09:26

AI with Papers - Artificial Intelligence & Deep Learning

This media is not supported in your browser

VIEW IN TELEGRAM

🪙MOMENTUM #NeurIPS 2025 🪙

👉MOMENTUM by Google (H/T Huguens Jean, Ph.D.) is a production multimodal agent architecture built on the Google ADK. It orchestrates 22 specialized tools (Gemini for reasoning, Imagen 4.0 for image generation, and Veo 3.1 for synthesis). Code announced💙

👉Review https://t.ly/06h7Q
👉Paper https://momentum-project-page-232993426383.us-central1.run.app/momentum_paper.pdf
👉Project https://momentum-project-page-232993426383.us-central1.run.app/
👉Repo TBA

👍3❤1

2.15K views13:46

AI with Papers - Artificial Intelligence & Deep Learning

😶‍🌫️ SOTA Full-Head Synthesis 😶‍🌫️

👉HyPlaneHead, the new SOTA in full-head image synthesis, delivering HQ results with significantly fewer artifacts compared to existing 3D-aware models. Repo announced💙

👉Review https://t.ly/WYfP3
👉Paper arxiv.org/pdf/2509.16748
👉Project https://lhyfst.github.io/hyplanehead/
👉Repo github.com/lhyfst/HyPlaneHead

❤4🔥3👍1😢1

1.87K viewsedited 13:33

AI with Papers - Artificial Intelligence & Deep Learning

This media is not supported in your browser

VIEW IN TELEGRAM

🍟 AnyTouch 2 is out 🍟

👉AnyTouch 2 is a general tactile representation learning framework for diverse optical tactile sensors that unifies object-level understanding with fine-grained, force-aware dynamic perception. Repo, Model & Data💙

👉Review https://t.ly/fP4dP
👉Paper https://arxiv.org/pdf/2602.09617
👉Project gewu-lab.github.io/AnyTouch2/
👉Repo github.com/GeWu-Lab/AnyTouch2

❤6

1.54K views09:36

AI with Papers - Artificial Intelligence & Deep Learning

Vote here please 💙

https://www.linkedin.com/posts/visionarynet_py4ai-2026-coming-soon-activity-7427290532034265088-y69e

❤2

1.51K views10:02

AI with Papers - Artificial Intelligence & Deep Learning

🍌 AGENT BANANA (SOTA) 🍌

👉Agent Banana is the novel SOTA agentic system for HD, native-resolution image editing through reasoning-based NL interaction, where each edit is context-aware, logically dependent, and locally precise. Code announced💙

👉Review https://t.ly/EXaCH
👉Paper https://arxiv.org/pdf/2602.09084
👉Project https://agent-banana.github.io/
👉Repo https://github.com/taco-group/agent-banana

❤11👏1

1.38K views13:14

AI with Papers - Artificial Intelligence & Deep Learning

This media is not supported in your browser

VIEW IN TELEGRAM

🛠️ IndustryShapes 6D Pose 🛠️

👉IndustryShapes by NTUA is a new RGB-D dataset of industrial tools, designed for both instance-level and novel object 6D pose estimation. Dataset available💙

👉Review https://t.ly/KKcuH
👉Paper https://arxiv.org/pdf/2602.05555
👉Project https://pose-lab.github.io/IndustryShapes/
👉Dataset https://huggingface.co/datasets/POSE-Lab/IndustryShapes

❤3🔥1

398 viewsedited 07:58

About

Blog

Apps

Platform