AI & ML Papers
34.1K subscribers
7.38K photos
597 videos
24 files
8.13K links
Advancing research in Machine Learning – practical insights, tools, and techniques for researchers.

Admin: @HusseinSheikho || @Hussein_Sheikho
Download Telegram
AI & ML Papers
Photo
🔥 Representation Distribution Matching for One-Step Visual Generation

💡 The paper introduces Representation Distribution Matching, a method for one-step visual generation that matches feature distributions under pretrained encoders. The goal is to generate high-quality images by comparing the distributions of generated and reference features. The authors identify two key design axes: how the distributions are compared and the representations they are compared in. They conduct controlled studies and find three main results.

First, they show that the Maximum Mean Discrepancy, a classical method that was previously ineffective, becomes a strong and scalable objective when estimated correctly. Second, they find that the batch size of the generated images has a significant impact on performance, with an optimum batch size above 2048, which is much larger than typical batch sizes. Third, they demonstrate that using a single representation can be gamed, resulting in low scores despite visibly fake images, and instead propose using a balanced set of encoders and evaluating with a Sliced-Wasserstein distance over 14 encoders.

The authors combine these findings to develop an improved Representation Distribution Matching method, which they call iRDM. They evaluate iRDM on the ImageNet dataset and achieve state-of-the-art results, with a Sliced-Wasserstein distance of 1.30. Additionally, they use a human-preference proxy, called PickScore, which shows that iRDM is preferred over the previous best one-step generator on 71.2% of matched samples. They also apply the same method to post-train a four-step generator, called FLUX.2, and achieve better results than the original four-step version, with improved performance on GenEval and PickScore, and requiring only 90 GPU-hours. Overall, the paper presents a new method for one-step visual generation that achieves state-of-the-art results and can be used to improve existing generators.


📅 Published on Jul 2

🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.02375
• PDF: https://arxiv.org/pdf/2607.02375
• Project Page: https://alan-lanfeng.github.io/rdm/

🤖 Models citing this paper:
• https://huggingface.co/epfl-vita/flux2-klein-1step-rdm

🚀 Spaces citing this paper:
• https://huggingface.co/spaces/epfl-vita/flux2-klein-1step-demo

━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus

#VisualGeneration #RepresentationLearning #DistributionMatching #ImageSynthesis #DeepLearning
AI & ML Papers
Photo
🔥 Scaling Properties of Text Conditioning in Visual Generation

💡 This paper studies the scaling properties of text conditioning in visual generation, which has rarely been measured due to the difficulty of scaling diffusion loss with the number of tokens in natural language prompts. The authors surprisingly find that the converged diffusion loss scales with the amount of structured language in the prompt. To quantify structured language, they adapt two complementary measures: a white-box like likelihood metric and a black-box attribute metric.

The authors use these metrics to analyze the relationship between the converged diffusion loss and the amount of structured language in the prompt. They find that the converged diffusion loss decreases approximately linearly with the white-box metric and follows a power law with the black-box metric across controlled training runs.

Guided by these scaling properties, the authors improve the diffusability of visual generation models by constructing structured prompts with semantic and geometric annotations derived from images. They also improve promptability by training a prompter through supervised fine-tuning, cold-start, and verifier-gated on-policy distillation.

The resulting system outperforms all evaluated open-weight models on nearly every compositional, reasoning, and world knowledge benchmark, while matching or surpassing the strongest closed-weight models on most evaluations. The paper's contributions include providing a better understanding of the scaling properties of text conditioning in visual generation and developing methods to improve the diffusability and promptability of visual generation models.


📅 Published on Jul 31

🔗 Links:
• GitHub: https://github.com/huggingface
• arXiv: https://arxiv.org/abs/2607.29679
• PDF: https://arxiv.org/pdf/2607.29679
• Project Page: https://heheyas.github.io/context-scaling/

━━━━━━━━━━━━━━━━━━━━━━━━
📢 By: https://xn--r1a.website/PaperNexus

#VisualGeneration #TextConditioning #DiffusionLoss #NaturalLanguageProcessing #StructuredLanguageMetrics