AI & ML Papers
Photo
🔥 DiffusionBench: On Holistic Evaluation of Diffusion Transformers
📅 Published on Jun 23
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2606.24888
• PDF: https://arxiv.org/pdf/2606.24888
• Project Page: https://end2end-diffusion.github.io/diffusion-bench/
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#DiffusionTransformers #ImageGenerationTasks #TextToImageGeneration #GenerativeModeling #DiffusionBasedArchitectures
💡 The paper introduces a unified framework called NanoGen for training and evaluating diffusion transformers, which are used in image generation tasks. The current evaluation setup for diffusion transformers is limited to class-conditional generation on ImageNet, which may not reflect real progress in generative modeling. The authors argue that text-to-image generation is a more comprehensive task, but it is often skipped due to perceived high costs and inconvenience. However, the authors show that with NanoGen, training and evaluating text-to-image models requires comparable compute to ImageNet.
The NanoGen framework supports various diffusion methods and can be easily configured to train models on both ImageNet and text-to-image tasks. The authors trained 21 latent diffusion models using NanoGen and found that the ranking of methods on ImageNet and text-to-image tasks shows no strong correlation. This suggests that a method that improves performance on ImageNet may not necessarily improve performance on text-to-image generation.
To address this issue, the authors propose a holistic benchmark called DiffusionBench, which summarizes results on both ImageNet and text-to-image tasks. The authors recommend reporting DiffusionBench in place of ImageNet alone, as methods that improve DiffusionBench are more likely to reflect broader progress in generative modeling. The main contribution of the paper is the introduction of NanoGen and DiffusionBench, which provide a more comprehensive evaluation setup for diffusion transformers and can help to advance research in generative modeling.
📅 Published on Jun 23
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2606.24888
• PDF: https://arxiv.org/pdf/2606.24888
• Project Page: https://end2end-diffusion.github.io/diffusion-bench/
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#DiffusionTransformers #ImageGenerationTasks #TextToImageGeneration #GenerativeModeling #DiffusionBasedArchitectures
GitHub
Hugging Face
The AI community building the future. Hugging Face has 469 repositories available. Follow their code on GitHub.
AI & ML Papers
Photo
🔥 Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation
📅 Published on Jul 14
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.13125
• PDF: https://arxiv.org/pdf/2607.13125
• Project Page: https://boogu.org/
🤖 Models citing this paper:
• https://huggingface.co/Boogu/Boogu-Image-0.1-Edit
• https://huggingface.co/Boogu/Boogu-Image-0.1-Base
• https://huggingface.co/Boogu/Boogu-Image-0.1-Turbo
🚀 Spaces citing this paper:
• https://huggingface.co/spaces/RioShiina/ImageGen
• https://huggingface.co/spaces/multimodalart/Boogu-Image
• https://huggingface.co/spaces/Tinchote/ImageGen
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#MultimodalUnderstanding #TextToImageGeneration #OpenSourceAI #MultimodalGeneration #InstructionBasedEditing
💡 The paper introduces Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family that delivers competitive performance in high-quality text-to-image generation, fast inference, instruction-based editing, and bilingual text rendering. The model family consists of Base, Turbo, Edit, and Edit-Turbo variants.
The problem addressed in the paper is that closed-source multimodal systems achieve strong performance through system-level integration rather than a single model, yet their internal practices remain largely undisclosed. The authors demonstrate that targeted improvements in model understanding, data quality, and training pipelines, coupled with genetic inference-time scaling, can substantially enhance generation and editing performance even under highly constrained compute budgets.
The method used to achieve this involves making improvements in model understanding, data quality, and training pipelines. The authors also use genetic inference-time scaling to enhance performance. The model is trained on a dataset of 208.62 million unique images, and the theoretical training cost is approximately 400,000 dollars.
The results show that Boogu-Image-0.1 consistently matches or surpasses other open-source models across standard benchmarks and achieves results approaching leading closed-source systems. The authors share practical discussions and release weights, code, and recipes under Apache 2.0 to advance the open ecosystem for unified multimodal understanding and generation. The code is available on GitHub.
Overall, the paper contributes to the development of open-source multimodal models that can achieve competitive performance with closed-source systems, and provides a valuable resource for the research community.
📅 Published on Jul 14
🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.13125
• PDF: https://arxiv.org/pdf/2607.13125
• Project Page: https://boogu.org/
🤖 Models citing this paper:
• https://huggingface.co/Boogu/Boogu-Image-0.1-Edit
• https://huggingface.co/Boogu/Boogu-Image-0.1-Base
• https://huggingface.co/Boogu/Boogu-Image-0.1-Turbo
🚀 Spaces citing this paper:
• https://huggingface.co/spaces/RioShiina/ImageGen
• https://huggingface.co/spaces/multimodalart/Boogu-Image
• https://huggingface.co/spaces/Tinchote/ImageGen
━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus
#MultimodalUnderstanding #TextToImageGeneration #OpenSourceAI #MultimodalGeneration #InstructionBasedEditing
GitHub
Hugging Face
The AI community building the future. Hugging Face has 469 repositories available. Follow their code on GitHub.