AI & ML Papers
Photo
🔥 What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion
📅 Published on May 8
🔗 Links:
• arXiv: https://arxiv.org/abs/2605.07915
• PDF: https://arxiv.org/pdf/2605.07915
• Project Page: https://zhengrongyue.github.io/pae.github.io/
• GitHub: https://github.com/ZhengrongYue/PAE ⭐ 29
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#LatentDiffusionModels #GenerativeModeling #AutoencoderArchitecture #LatentManifoldLearning #DiffusionBasedGenerativeModels
💡 This paper investigates the properties of a latent manifold that are favorable for diffusion models, which are a type of generative model. The authors argue that existing methods for defining the latent space, known as tokenizers, are primarily designed to improve reconstruction fidelity or inherit pre-trained representations, but do not necessarily produce a latent space that is well-suited for generative modeling. To address this issue, the authors study the properties of a diffusion-friendly latent manifold and identify three key properties: coherent spatial structure, local manifold continuity, and global manifold semantics. They find that these properties are more closely related to downstream generation quality than reconstruction fidelity.
To explicitly shape the latent manifold with these desirable properties, the authors propose a new method called the Prior-Aligned AutoEncoder, or PAE. The PAE uses refined priors derived from variational autoencoders and perturbation-based regularization to turn the desired properties of the latent manifold into explicit training objectives. This approach allows the PAE to directly optimize the latent space structure for improved generative modeling.
The authors evaluate the PAE on the ImageNet 256x256 dataset and find that it improves both training efficiency and generation quality compared to existing tokenizers. Specifically, the PAE achieves comparable performance to the state-of-the-art method, RAE, but with up to 13 times faster convergence under the same training setup. Additionally, the PAE achieves a new state-of-the-art result, with a generative fidelity score of 1.03. These results highlight the importance of organizing the latent manifold for latent diffusion models and demonstrate the effectiveness of the PAE in producing high-quality generative models.
📅 Published on May 8
🔗 Links:
• arXiv: https://arxiv.org/abs/2605.07915
• PDF: https://arxiv.org/pdf/2605.07915
• Project Page: https://zhengrongyue.github.io/pae.github.io/
• GitHub: https://github.com/ZhengrongYue/PAE ⭐ 29
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#LatentDiffusionModels #GenerativeModeling #AutoencoderArchitecture #LatentManifoldLearning #DiffusionBasedGenerativeModels
arXiv.org
What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned...
Tokenizers are a crucial component of latent diffusion models, as they define the latent space in which diffusion models operate. However, existing tokenizers are primarily designed to improve...
❤1
AI & ML Papers
Photo
🔥 DiffusionBench: On Holistic Evaluation of Diffusion Transformers
📅 Published on Jun 23
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2606.24888
• PDF: https://arxiv.org/pdf/2606.24888
• Project Page: https://end2end-diffusion.github.io/diffusion-bench/
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#DiffusionTransformers #ImageGenerationTasks #TextToImageGeneration #GenerativeModeling #DiffusionBasedArchitectures
💡 The paper introduces a unified framework called NanoGen for training and evaluating diffusion transformers, which are used in image generation tasks. The current evaluation setup for diffusion transformers is limited to class-conditional generation on ImageNet, which may not reflect real progress in generative modeling. The authors argue that text-to-image generation is a more comprehensive task, but it is often skipped due to perceived high costs and inconvenience. However, the authors show that with NanoGen, training and evaluating text-to-image models requires comparable compute to ImageNet.
The NanoGen framework supports various diffusion methods and can be easily configured to train models on both ImageNet and text-to-image tasks. The authors trained 21 latent diffusion models using NanoGen and found that the ranking of methods on ImageNet and text-to-image tasks shows no strong correlation. This suggests that a method that improves performance on ImageNet may not necessarily improve performance on text-to-image generation.
To address this issue, the authors propose a holistic benchmark called DiffusionBench, which summarizes results on both ImageNet and text-to-image tasks. The authors recommend reporting DiffusionBench in place of ImageNet alone, as methods that improve DiffusionBench are more likely to reflect broader progress in generative modeling. The main contribution of the paper is the introduction of NanoGen and DiffusionBench, which provide a more comprehensive evaluation setup for diffusion transformers and can help to advance research in generative modeling.
📅 Published on Jun 23
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2606.24888
• PDF: https://arxiv.org/pdf/2606.24888
• Project Page: https://end2end-diffusion.github.io/diffusion-bench/
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#DiffusionTransformers #ImageGenerationTasks #TextToImageGeneration #GenerativeModeling #DiffusionBasedArchitectures
GitHub
Hugging Face
The AI community building the future. Hugging Face has 469 repositories available. Follow their code on GitHub.
🔥 JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents
📅 Published on Jul 26
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.23588
• PDF: https://arxiv.org/pdf/2607.23588
• Project Page: https://www.jarvishub.site/
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#MultimodalCreativeAgents #CanvasNativeGeneration #GenerativeModeling #CreativeArtificialIntelligence #MultimodalProductionSystems
💡 The paper introduces JarvisHub, an open harness for canvas-native multimodal creative agents, which aims to address the limitations of existing generative models in supporting long-horizon multimodal creative production. Current models can synthesize high-quality images, videos, audio clips, and other creative assets, but real-world creative work requires more than isolated prompt-output interactions. It involves references, drafts, alternatives, edits, failed attempts, version relations, tool actions, evaluation signals, and human feedback, which together form an evolving project state.
Existing prompt-based, chat-based, and node-based generation systems only partially support this state, as they often discard intermediate context, rely on linear conversations, or require manually specified workflows. Recent commercial systems indicate a shift towards agent-assisted creative production, but their closed architectures make it difficult to study how agents represent context, choose tools, revise artifacts, recover from failures, and maintain consistency over time.
JarvisHub treats an editable canvas as the user workspace, the agent's external memory, action space, and shared project state, representing multimodal artifacts, dependencies, versions, and feedback as typed canvas nodes and links. Through a three-layer architecture of canvas state, protocol bridge, and agent runtime, JarvisHub enables agents to act within an inspectable and editable creative state. This design moves creative agents beyond isolated tool use towards sustained, human-steerable creative automation, where agents can progressively plan, generate, revise, and organize multimodal projects while users remain able to inspect, guide, and intervene throughout the process.
The paper's contributions include the introduction of JarvisHub as an open harness for canvas-native multimodal creative agents, which provides a flexible and inspectable framework for long-horizon multimodal creative production. The system allows agents to act within an editable canvas, representing the user workspace, agent's external memory, action space, and shared project state. The three-layer architecture of JarvisHub enables agents to progressively plan, generate, revise, and organize multimodal projects, while users can inspect, guide, and intervene throughout the process. Overall, JarvisHub has the potential to advance the field of creative AI by providing a more flexible and sustainable approach to multimodal creative production.
📅 Published on Jul 26
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.23588
• PDF: https://arxiv.org/pdf/2607.23588
• Project Page: https://www.jarvishub.site/
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#MultimodalCreativeAgents #CanvasNativeGeneration #GenerativeModeling #CreativeArtificialIntelligence #MultimodalProductionSystems
GitHub
Hugging Face
The AI community building the future. Hugging Face has 469 repositories available. Follow their code on GitHub.
AI & ML Papers
Photo
🔥 GNM Head: A Generative aNthropometric Model of the human head
📅 Published on Jul 26
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.23687
• PDF: https://arxiv.org/pdf/2607.23687
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#GenerativeModeling #AnthropometricAnalysis #3DHeadModeling #FacialReconstruction #HumanComputerInteraction
💡 The paper introduces a new parametric model called the Generative Anthropometric Model of the human head, or GNM Head. The model is designed to address the limitations of existing publicly available models, which typically only capture the outer geometry of the head and ignore internal structures such as the eyes and mouth. These existing models also often suffer from reduced geometric quality due to low-fidelity input data sets.
The GNM Head model is built on an extensive database of high-resolution 3D scans combined with high-quality anatomy-specific artist-made samples. The model encompasses the head, face, neck, eyeballs, teeth, and tongue, and includes specialized sub-models for the ocular and intra-oral structures.
The paper details the data provenance, model architecture, and performance of the GNM Head model, including its ability to fit target 3D face scans. The results show that the GNM Head model is a significant improvement over existing models, offering a more comprehensive and accurate representation of the human head.
To foster community innovation, the complete GNM framework is made publicly available. The GNM Head model has the potential to serve as a crucial conditioning signal within generative large vision models, allowing for tight spatial control of generated imagery. Overall, the paper contributes a new and improved parametric model of the human head, which can be used in a variety of applications, including computer vision, graphics, and animation.
📅 Published on Jul 26
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.23687
• PDF: https://arxiv.org/pdf/2607.23687
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#GenerativeModeling #AnthropometricAnalysis #3DHeadModeling #FacialReconstruction #HumanComputerInteraction
GitHub
Hugging Face
The AI community building the future. Hugging Face has 469 repositories available. Follow their code on GitHub.
❤2