AI & ML Papers
Photo
🔥 OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation
📅 Published on Jul 26
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.23855
• PDF: https://arxiv.org/pdf/2607.23855
• Project Page: https://openmoss.ai/OmniVAE.github.io/
🤖 Models citing this paper:
• https://huggingface.co/OpenMOSS-Team/OmniVAE
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#AudioVideoGeneration #CrossModalAlignment #VariationalAutoencoders #MultimodalLearning #JointGenerationModels
💡 The paper introduces OmniVAE, a joint audio-video variational autoencoder that learns fine-grained semantic alignment between audio and video latent representations. Recent generative models have moved beyond silent video or standalone audio synthesis towards joint generation of synchronized audio and video, but this remains challenging due to the fundamental structural differences between the two modalities. Most existing methods use audio and video VAEs trained separately, resulting in a lack of cross-modal alignment, which leaves the downstream generative model to learn cross-modal synchronization from scratch.
OmniVAE addresses this issue by jointly training an audio-video VAE that captures temporal-semantic correspondence and aligns the two latent spaces. It uses a segment-level audio-video contrastive objective to achieve this alignment and also distills features from pre-trained modality-specific semantic encoders into each modality, improving the downstream learnability of both latent spaces.
The results of extensive experiments show that OmniVAE consistently improves the learnability of the latent spaces, translating into higher generation quality and more accurate cross-modal synchronization in downstream text-to-audio-video generation tasks. The findings underscore the importance of learning unified representations as a foundation for omnimodal modeling. Overall, OmniVAE provides a novel approach to joint audio-video generation, enabling more effective and synchronized generation of audio and video content.
📅 Published on Jul 26
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.23855
• PDF: https://arxiv.org/pdf/2607.23855
• Project Page: https://openmoss.ai/OmniVAE.github.io/
🤖 Models citing this paper:
• https://huggingface.co/OpenMOSS-Team/OmniVAE
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#AudioVideoGeneration #CrossModalAlignment #VariationalAutoencoders #MultimodalLearning #JointGenerationModels
GitHub
Hugging Face
The AI community building the future. Hugging Face has 458 repositories available. Follow their code on GitHub.