Github Top Repositories
14.5K subscribers
3.88K photos
61 videos
10 files
3.52K links
Top GitHub repositories in one place πŸš€
Explore the best projects in programming, AI, data science, and more.
Download Telegram
πŸ”– ImageBind: One Embedding Space To Bind Them All

πŸ“ This project is a significant step forward in understanding and connecting information from diverse sources like images, text, audio, video, and even motion sensor data.

βš™οΈ Supports 6 Modalities:

πŸ“· Image
πŸ“ Text
πŸ”ˆ Audi
πŸŽ₯ Video
🦴 IMU sensor data (e.g., accelerometer)
πŸ™„ Depth/Thermal & 3D data
Interestingly, only some modalities had labels, yet ImageBind learned to align them through self-supervised learning.


πŸ’« Key Features:

..No need for paired data (e.g., images and audio don’t have to be aligned)..Leverages contrastive learning for learning joint embedding space
..Competes with CLIP and AudioCLIP, but with better accuracy and coverage..Enables zero-shot retrieval (e.g., finding relevant video using just a sentence)


πŸ“Œ Repo: https://github.com/facebookresearch/ImageBind

πŸ” By: https://xn--r1a.website/DataScienceN 🌟

#ImageBind #MultimodalAI #MetaAI #DeepLearning #SelfSupervised
Please open Telegram to view this post
VIEW IN TELEGRAM
πŸ‘3πŸ”₯2
This media is not supported in your browser
VIEW IN TELEGRAM
πŸˆβ€β¬› TTT Long Video Generation πŸˆβ€β¬›

▢️ A novel architecture for video generation, adapting the #CogVideoX 5B model by incorporating #TestTimeTraining (TTT) layers.
Adding TTT layers into a pre-trained Transformer enables generating a one-minute clip from text storyboards.
Videos, code & annotations released πŸ’™

πŸ”— Review: https://t.ly/mhlTN
πŸ“„ Paper: arxiv.org/pdf/2504.05298
🌐 Project: test-time-training.github.io/video-dit
πŸ§‘β€πŸ’» Repo: github.com/test-time-training/ttt-video-dit

#AI #VideoGeneration #MachineLearning #DeepLearning #Transformers #TTT #GenerativeAI

πŸ” By: https://xn--r1a.website/DataScienceN5
Please open Telegram to view this post
VIEW IN TELEGRAM
πŸ‘3πŸ₯°2