AI & ML Papers
33.4K subscribers
7.17K photos
556 videos
24 files
7.87K links
Advancing research in Machine Learning – practical insights, tools, and techniques for researchers.

Admin: @HusseinSheikho || @Hussein_Sheikho
Download Telegram
Generating an Image From 1,000 Words: Enhancing Text-to-Image With Structured Captions

📝 Summary:
This paper introduces FIBO, a text-to-image model trained on long structured captions to enhance prompt alignment and controllability. It proposes DimFusion for efficient processing and the TaBR evaluation protocol, achieving state-of-the-art results.

🔹 Publication Date: Published on Nov 10

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2511.06876
• PDF: https://arxiv.org/pdf/2511.06876

🔹 Models citing this paper:
https://huggingface.co/briaai/FIBO

Spaces citing this paper:
https://huggingface.co/spaces/galdavidi/FIBO-Mashup
https://huggingface.co/spaces/briaai/FIBO
https://huggingface.co/spaces/briaai/Fibo-local

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#TextToImage #GenerativeAI #DiffusionModels #AI #MachineLearning
Benchmarking Diversity in Image Generation via Attribute-Conditional Human Evaluation

📝 Summary:
This paper introduces a framework to robustly evaluate diversity in text-to-image models. It uses a novel human evaluation template, curated prompts with variation factors, and systematic analysis of image embeddings to rank models and identify diversity weaknesses.

🔹 Publication Date: Published on Nov 13

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2511.10547
• PDF: https://arxiv.org/pdf/2511.10547

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#ImageGeneration #TextToImage #AIDiversity #Benchmarking #HumanEvaluation
Mixture of States: Routing Token-Level Dynamics for Multimodal Generation

📝 Summary:
MoS is a novel multimodal diffusion model that uses a learnable token-wise router for flexible state-based modality interactions. This achieves state-of-the-art text-to-image generation and editing with minimal parameters and computational overhead.

🔹 Publication Date: Published on Nov 15

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2511.12207
• PDF: https://arxiv.org/pdf/2511.12207

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#GenerativeAI #MultimodalAI #DiffusionModels #TextToImage #DeepLearning
UltraFlux: Data-Model Co-Design for High-quality Native 4K Text-to-Image Generation across Diverse Aspect Ratios

📝 Summary:
UltraFlux overcomes diffusion transformer failures at 4K resolution and diverse aspect ratios through data-model co-design. It uses enhanced positional encoding, VAE improvements, gradient rebalancing, and aesthetic curriculum learning to achieve superior 4K text-to-image generation, outperformin...

🔹 Publication Date: Published on Nov 22

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2511.18050
• PDF: https://arxiv.org/pdf/2511.18050
• Project Page: https://github.com/W2GenAI-Lab/UltraFlux
• Github: https://github.com/W2GenAI-Lab/UltraFlux

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#TextToImage #GenerativeAI #4KGeneration #DiffusionModels #AIResearch
Decoupled DMD: CFG Augmentation as the Spear, Distribution Matching as the Shield

📝 Summary:
This study challenges the understanding of Distribution Matching Distillation DMD for text-to-image generation. It reveals that CFG Augmentation is the primary driver of few-step distillation, while distribution matching acts as a regularizer. This new insight enables improved distillation method...

🔹 Publication Date: Published on Nov 27

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2511.22677
• PDF: https://arxiv.org/pdf/2511.22677
• Project Page: https://tongyi-mai.github.io/Z-Image-blog/
• Github: https://github.com/Tongyi-MAI/Z-Image/tree/main

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#TextToImage #GenerativeAI #DiffusionModels #ModelDistillation #AIResearch
Where Culture Fades: Revealing the Cultural Gap in Text-to-Image Generation

📝 Summary:
Multilingual text-to-image models often generate culturally neutral images. This paper identifies specific neurons for cultural information and proposes two strategies: inference-time activation and layer-targeted enhancement. These methods improve cultural consistency while preserving image qual...

🔹 Publication Date: Published on Nov 21

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2511.17282
• PDF: https://arxiv.org/pdf/2511.17282

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#TextToImage #CulturalAI #ResponsibleAI #DeepLearning #AIResearch
Rethinking Prompt Design for Inference-time Scaling in Text-to-Visual Generation

📝 Summary:
PRIS adaptively revises prompts during text-to-visual generation inference to enhance user intent alignment. It reviews visual failures and redesigns prompts using fine-grained feedback, proving that jointly scaling prompts and visuals improves accuracy and quality.

🔹 Publication Date: Published on Dec 3

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2512.03534
• PDF: https://arxiv.org/pdf/2512.03534
• Project Page: https://subin-kim-cv.github.io/PRIS

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#PromptEngineering #TextToImage #GenerativeAI #DeepLearning #AIResearch
DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation

📝 Summary:
DraCo is a novel text-to-image generation method that uses interleaved reasoning with both textual and visual content. It generates low-resolution drafts, verifies semantic alignment, and refines images to address coarse textual planning and rare attribute generation. DraCo significantly outperfo...

🔹 Publication Date: Published on Dec 4

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2512.05112
• PDF: https://arxiv.org/pdf/2512.05112
• Github: https://github.com/CaraJ7/DraCo

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#TextToImage #GenerativeAI #DeepLearning #ComputerVision #AI
Aligned but Stereotypical? The Hidden Influence of System Prompts on Social Bias in LVLM-Based Text-to-Image Models

📝 Summary:
LVLM-based text-to-image models exhibit greater social bias than non-LVLM models, with system prompts identified as the key driver. The paper introduces FairPro, a training-free meta-prompting framework that significantly reduces demographic bias while maintaining text-image alignment.

🔹 Publication Date: Published on Dec 4

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2512.04981
• PDF: https://arxiv.org/pdf/2512.04981
• Github: https://github.com/nahyeonkaty/fairpro

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#AIBias #TextToImage #LVLMs #PromptEngineering #AIFairness
RealGen: Photorealistic Text-to-Image Generation via Detector-Guided Rewards

📝 Summary:
RealGen is a photorealistic text-to-image framework addressing AI artifacts in current models. It uses an LLM for prompt optimization and a diffusion model, enhanced by a Detector Reward mechanism that quantifies artifacts and assesses realism. RealGen significantly outperforms other models, achi...

🔹 Publication Date: Published on Nov 29

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2512.00473
• PDF: https://arxiv.org/pdf/2512.00473
• Project Page: https://yejy53.github.io/RealGen/
• Github: https://yejy53.github.io/RealGen/

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#TextToImage #GenerativeAI #DiffusionModels #AIResearch #ComputerVision
SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder

📝 Summary:
SVG-T2I enables high-quality text-to-image synthesis directly in the Visual Foundation Model feature domain. This scaled framework achieves competitive performance without a variational autoencoder, validating VFM representations for generative tasks.

🔹 Publication Date: Published on Dec 12

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2512.11749
• PDF: https://arxiv.org/pdf/2512.11749
• Github: https://github.com/KlingTeam/SVG-T2I

🔹 Models citing this paper:
https://huggingface.co/KlingTeam/SVG-T2I

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#TextToImage #DiffusionModels #GenerativeAI #VisualFoundationModels #DeepLearning
Directional Textual Inversion for Personalized Text-to-Image Generation

📝 Summary:
Directional Textual Inversion DTI enhances text-to-image personalization by fixing learned token magnitudes and optimizing only their direction. This prevents norm inflation issues of standard Textual Inversion, improving prompt conditioning and enabling smooth interpolation. DTI offers better te...

🔹 Publication Date: Published on Dec 15

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2512.13672
• PDF: https://arxiv.org/pdf/2512.13672
• Project Page: https://kunheek.github.io/dti
• Github: https://github.com/kunheek/dti

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#TextualInversion #TextToImage #GenerativeAI #DeepLearning #AI
Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing

📝 Summary:
This paper proposes a framework using a semantic-pixel reconstruction objective to adapt encoder features for generation. It creates a compact, semantically rich latent space, leading to state-of-the-art image reconstruction and improved text-to-image generation and editing.

🔹 Publication Date: Published on Dec 19

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2512.17909
• PDF: https://arxiv.org/pdf/2512.17909
• Project Page: https://jshilong.github.io/PS-VAE-PAGE/
• Github: https://jshilong.github.io/PS-VAE-PAGE/

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#TextToImage #ImageGeneration #DeepLearning #ComputerVision #AIResearch
1
MineTheGap: Automatic Mining of Biases in Text-to-Image Models

📝 Summary:
MineTheGap automatically finds prompts that cause Text-to-Image models to generate biased outputs. It uses a genetic algorithm and a novel bias score to identify and rank biases, aiming to reduce redundancy and improve output diversity.

🔹 Publication Date: Published on Dec 15

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2512.13427
• PDF: https://arxiv.org/pdf/2512.13427

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#AIbias #TextToImage #GenerativeAI #ResponsibleAI #MachineLearning
Everything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image Models

📝 Summary:
Text-to-image models struggle with complex spatial reasoning due to sparse prompts. This paper introduces SpatialGenEval, a new benchmark with dense prompts, showing models struggle with higher-order spatial tasks. A new dataset, SpatialT2I, helps fine-tune models for significant performance gain...

🔹 Publication Date: Published on Jan 28

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2601.20354
• PDF: https://arxiv.org/pdf/2601.20354
• Github: https://github.com/AMAP-ML/SpatialGenEval

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#TextToImage #SpatialReasoning #GenerativeAI #ComputerVision #AIResearch
DenseGRPO: From Sparse to Dense Reward for Flow Matching Model Alignment

📝 Summary:
DenseGRPO addresses sparse rewards in flow matching models by providing dense, step-wise rewards for intermediate denoising steps. It uses these rewards to adaptively calibrate exploration, improving alignment with human preferences in text-to-image generation.

🔹 Publication Date: Published on Jan 28

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2601.20218
• PDF: https://arxiv.org/pdf/2601.20218

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#AI #MachineLearning #ReinforcementLearning #TextToImage #GenerativeAI
Enhancing Spatial Understanding in Image Generation via Reward Modeling

📝 Summary:
Text-to-image models struggle with complex spatial relationships. This paper introduces SpatialScore, a reward model trained on 80k preference pairs, to evaluate and improve spatial accuracy. It significantly enhances spatial understanding in image generation via reinforcement learning.

🔹 Publication Date: Published on Feb 27

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2602.24233
• PDF: https://arxiv.org/pdf/2602.24233
• Project Page: https://dagroup-pku.github.io/SpatialT2I/
• Github: https://github.com/DAGroup-PKU/SpatialT2I

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#ImageGeneration #TextToImage #SpatialAI #RewardModeling #DeepLearning
Conditioned Activation Transport for T2I Safety Steering

📝 Summary:
Current T2I models generate unsafe content, and linear steering degrades image quality. This paper proposes Conditioned Activation Transport CAT, which uses geometric conditioning and nonlinear transport maps to activate only in unsafe regions. CAT significantly reduces unsafe content generation ...

🔹 Publication Date: Published on Mar 3

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2603.03163
• PDF: https://arxiv.org/pdf/2603.03163
• Github: https://github.com/NASK-AISafety/conditional-activation-transport

Datasets citing this paper:
https://huggingface.co/datasets/NASK-PIB/SafeSteerDataset

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#AISafety #TextToImage #GenerativeAI #DeepLearning #AIethics
CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation

📝 Summary:
CoCo is a code-driven framework for text-to-image generation, using executable code for precise spatial layout and structured image creation. It significantly outperforms natural language CoT methods, enabling more controllable and accurate image synthesis.

🔹 Publication Date: Published on Mar 9

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2603.08652
• PDF: https://arxiv.org/pdf/2603.08652

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#TextToImage #GenerativeAI #AIResearch #CodeDrivenAI #ComputerVision
1
UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models

📝 Summary:
UDM-GRPO integrates Uniform Discrete Diffusion Models with reinforcement learning, solving training instability issues. It optimizes using final samples as actions and reconstructed trajectories. This achieves state-of-the-art performance in text-to-image generation and OCR tasks.

🔹 Publication Date: Published on Apr 20

🔹 Paper Links:
• arXiv Page: https://arxiv.org/abs/2604.18518
• PDF: https://arxiv.org/pdf/2604.18518
• Project Page: https://yovecent.github.io/UDM-GRPO.github.io/
• Github: https://github.com/Yovecent/UDM-GRPO

🔹 Models citing this paper:
https://huggingface.co/Yovecents/URSA-1.7B-IBQ512-UDMGRPO-GenEval
https://huggingface.co/Yovecents/URSA-1.7B-IBQ512-UDMGRPO-PickScore

==================================

For more data science resources:
https://xn--r1a.website/DataScienceT

#DiffusionModels #ReinforcementLearning #GenerativeAI #TextToImage #DeepLearning
1