Réunion
Modèles génératifs : diffusion et flow matching
Axes scientifiques :
- Théorie et méthodes
Organisateurs :
- - Mathurin Massias (LIP - Lyon)
- - Segolene Martin (LIP - Lyon)
Nous vous rappelons que, afin de garantir l'accès de tous les inscrits aux salles de réunion, l'inscription aux réunions est gratuite mais obligatoire.
La réunion sera également accessible en distanciel mais l'inscription est obligatoire
Inscriptions
97 personnes membres du GdR IASIS, et 189 personnes non membres du GdR, sont inscrits à cette réunion.
Capacité de la salle : 150 personnes. Nombre d'inscrits en présentiel : 100 ; Nombre d'inscrits en distanciel : 186
50 Places restantes
Inscriptions closes pour cette journée
Annonce
Les inscriptions réalisées en juillet n’ont pas été prises en compte; il est nécessaire de se réinscrire.
Les demandes de prise en charge de mission par le GdR sont à formuler avant le 25 septembre.
Les modèles génératifs ont connu de récentes avancées spectaculaires, au point que leurs dernières versions sont désormais capables de produire des images et du texte synthétiques presque indiscernables de contenus réels. Parmi les approches ayant contribué à ces progrès, les approches de diffusion et de flow matching occupent une place centrale.
Cette réunion vise à faire le point sur les derniers développements dans le domaine des modèles génératifs.
L’appel à contributions est ouvert aux travaux portant sur des développements théoriques, algorithmiques, ou des applications. Une liste non exhaustive de thèmes inclut :
- modèles « one-step » (consistency, distillation, flow maps)
- alignement et guidance
- mémorisation et généralisation
- diffusion et flow matching pour les données discrètes et le texte
- liens avec le transport optimal
- évaluation des modèles génératifs
Orateur.ice.s invité.e.s
- Vicky Kalogeiton (LIX, Ecole Polytechnique) — Scale is Religion?
- Anna Korba (CREST, ENSAE) — A Unifying View of Variational Generative Wasserstein Flows
- Umut Simsekli (SIERRA, INRIA) — Algorithm- and Data-Dependent Generalization Bounds for Diffusion Models
L’appel à contribution (posters et/ou présentations de 20 min + 5 min de questions) est désormais clos. All the talks will be in English.
Organisateur.ice.s
- Ségolène Martin (OCKHAM, INRIA)
- Mathurin Massias (OCKHAM, INRIA)
La journée bénéficie du soutien de l’Institut Rhônalpin des Systèmes Complexes (IXXI)
Planning
9 h – 9 h 25 Participants welcome + Introduction
9 h 25 – 9 h 50 Short talk 1: Nicolas Dufour
9 h 50 – 10 h 25 Long Talk 1: Rémi Emonet
10 h 25 – 10 h 40 Coffee break
10 h 40 – 11 h 25 LT 2: Anna Korba
11 h 25 – 11 h 50 ST2: Marie Scheid
11 h 50 – 13 h 15 Lunch break on you own
13 h 30 – 15 h 15 Poster session
15 h 15 – 16 h 05 Long Talk 3: Umut Simsekli
16 h 05 – 16 h 15 Coffee break
16 h 15 – 16 h 40 ST3: Raphaël Urfin
16 h 40 – 17 h 05 ST4: Jerôme Garnier Brun
17 h 05 – 17 h 30 ST5: Mariia Vladimirova
Abstracts Long Talks:
Anna Korba (CREST, ENSAE): A Unifying View of Variational Generative Wasserstein Flows
Many modern generative models can be viewed as minimizing divergences between probability distributions, yet they rely on different algorithmic and geometric principles. Wasserstein gradient flows provide a continuous-time formulation for optimizing over distributions, and can be approximated through their implicit discretization via the Jordan-Kinderlehrer-Otto (JKO) scheme. In this work, we present a unified theoretical framework for generative modeling based on Wasserstein gradient flows, which we refer to as Generative Wasserstein Flows (GWF). We show that a broad class of existing methods can be derived as instances of parametric JKO schemes for f-divergence objectives, and we establish equivalences between several recently proposed algorithms. We extend this framework beyond f-divergence to Integral Probability Metrics and squared Maximum Mean Discrepancy, deriving new JKO-based generative algorithms, and clarifying their connections with GANs. We study empirically the impact of the JKO regularization for a wide set of objectives. Finally, we analyze parametric Wasserstein flows, where the dynamics are restricted to distributions induced by parametrized maps. This is a joint work with Paul Caucheteux (CREST, ENSAE, IP Paris) and Clément Bonet (CMAP, Polytechnique, IP Paris).
Vicky Kalogeiton (LIX, École Polytechnique): Scale is Religion? / Cancelled
Intelligent robots do not just respond to commands; they imagine what you meant, what you wanted, what you believed. And they do this while learning from very little, and running on a chip in your living room. In this talk, I will present recent advances in generative modeling that aim to equip embodied agents with efficient models that can run faster, with fewer data, and more efficient models and imagine possible futures under uncertainty.
Umut Simsekli (Inria, Ecole Normale Supérieure): Algorithm- and Data-Dependent Generalization Bounds for Diffusion Models
Score-based generative models (SGMs) have emerged as one of the most popular classes of generative models. A substantial body of work now exists on the analysis of SGMs, focusing either on discretization aspects or on their statistical performance. In the latter case, bounds have been derived, under various metrics, between the true data distribution and the distribution induced by the SGM, often demonstrating polynomial convergence rates with respect to the number of training samples. However, these approaches adopt a largely approximation theory viewpoint, which tends to be overly pessimistic and relatively coarse. In particular, they fail to fully explain the empirical success of SGMs or capture the role of the optimization algorithm used in practice to train the score network. To support this observation, we first present simple experiments illustrating the concrete impact of optimization hyperparameters on the generalization ability of the generated distribution. Then, this talk aims to bridge this theoretical gap by providing the first algorithmic- and data-dependent generalization analysis for SGMs. In particular, we establish bounds that explicitly account for the optimization dynamics of the learning algorithm, offering new insights into the generalization behavior of SGMs. Our theoretical findings are supported by empirical results on several datasets.
Rémi Emonet (LabHC): Mitigating Memorization in Closed-Form Flow Matching with Bootstrap Aggregating
Diffusion and flow matching are a class of generative models that generate new samples by solving ordinary or stochastic differential equations with a learned score/velocity field. Interestingly, the optimal velocity field admits a closed-form formula that can be computed for finite datasets. Generating samples following the optimal velocity field can only reproduce samples from the training set, i.e., memorize the training set. Neural networks, trained to match the velocity field, introduce inductive bias that can partially mitigate the issue. However, these models can still exhibit memorization, unlike the benign overfitting observed in discriminative tasks. In this work, we propose a bootstrap aggregating (bagging) method for flow matching to reduce memorization. By conceptually averaging over resampled training subsets, our approach effectively reduces memorization. We derive a closed-form bagging formulation compatible with exact flow matching, enabling efficient implementation. Experiments confirm reduced memorization and better generalization without architectural changes or auxiliary objectives.
Abstracts Short Talks:
Nicolas Dufour (Kyutai): The FID Lottery: Quantifying Hidden Randomness in Generative-Model Evaluation
he Frechet Inception Distance (FID) is the de facto arbiter of image generation, yet most papers report just a single number from a single trained model using a single sampling seed. How reproducible is that number if we retrain the model, or merely resample from it? In this paper, we treat FID as a random variable on a two-axis panel of training and generation seeds, and measure its variance directly on several hundred SiT networks trained on class-conditional ImageNet 256×256. We report surprising findings: (a) Retraining the model using the same recipe with a different seed moves FID 3.2x more (in Inception feature space) than redrawing samples from a fixed network. (b) That gap is driven by three factors: random initialisation, data ordering, and the per-step Gaussian noise of the flow-matching loss. (c) Increasing compute or model size barely tightens the spread, holding the FID coefficient of variation (CoV) inside a 1-2% band. (d) Per-cell classifier-free-guidance tuning halves the spread but reshuffles which seeds work best, and a lucky training seed reaches the same FID with up to 2x less compute than an unlucky one. Based on these findings, we recommend a new FID evaluation protocol: evaluate under per-cell optimal guidance, treat any FID gap below the empirically measured ~1.3% CoV as inconclusive, and report an error bar over several training seeds rather than a single FID number.
Marie Scheid (CMAP, Ecole polytechnique): Twister Schrödinger Bridge Matching
Over the past few years, diffusion-based Schrödinger bridge models have been proposed to approximate optimal transport dynamics between two prescribed boundary distributions, with successful applications to generative modeling. More precisely, these methods aim to estimate a path measure whose initial and terminal marginals match the two boundary distributions, while minimizing the Kullback-Leibler divergence with respect to a reference Markov process. In this work, we consider the generalized Schrödinger bridge problem, in which the reference process is a twisted Brownian motion, that is, a Feynman-Kac transform of a Brownian motion induced by a time-dependent differentiable potential. Building on the Iterative Markovian Fitting (IMF) paradigm, and in particular on its special case Diffusion Schrödinger Bridge Matching (DSBM), which corresponds to the zero potential case, we introduce Twisted Schrödinger Bridge Matching (TSBM), a diffusion-based method designed to handle both continuous- and discrete-time potentials. Unlike previous approaches, TSBM provides a rigorous extension of the IMF scheme to the generalized Schrödinger bridge problem. This derivation leads to a new bridge-matching loss that depends explicitly on the gradient of the potential and recovers the DSBM objective when the potential vanishes, yielding improved performance. We further introduce trajectory-based variance-reduction techniques that substantially stabilize optimization and may be useful beyond the present setting. Finally, we empirically demonstrate the benefits of TSBM for trajectory inference across increasingly high-dimensional settings, including crowd navigation and single-cell data.
Raphaël Urfin (Ecole Normale Supérieure): Double Descent and Malign Overfitting in Diffusion Models
Conventional wisdom in deep learning holds that overparameterization—having more parameters $p$ than training samples $n$—is benign: larger models generalize better and, even without regularization, interpolating models generalize well, the test error following a double-descent curve. One might expect the same benign overfitting for diffusion models, whose training reduces to regression, i.e. to minimizing a quadratic denoising score-matching loss. Yet the opposite is observed: overfitting here is catastrophic, driving the model into a memorization regime. We resolve this paradox by combining experiments on U-Nets trained on CelebA with a random-features model for which we derive closed-form learning curves. First, we show that at fixed number $m$ of noise realizations per training sample the interpolation peak does occur, at $p\sim nm$ rather than at $p\sim n$ as in standard regression. Since diffusion models are trained with $m\gg1$ in practice, the peak is pushed to very large model sizes, and what is observed at common sizes as the rising branch of a U-shaped curve is in fact the approach to it. Second, beyond the peak overfitting is \emph{malign}: the implicit regularization of training is fully at work, but it drives the model toward the empirical score, which memorizes the training set, rather than toward the true score. A bias–variance decomposition pinpoints the mechanism: past the peak the variance of the score estimator decays, as in regression, and its bias grows, and both saturate at a large value. Nevertheless, overparameterization remains beneficial when paired with regularization: both analytically and numerically, optimally regularized large models—via a ridge penalty or early stopping, respectively—outperform unregularized models of any size.
Jerôme Garnier Brun (Università Bocconi): Biased Generalization in Generative Diffusion
In generative diffusion, generalization is usually assessed through held-out performance, and training is stopped at the minimum of the test loss. We show that this criterion misses a phase of biased generalization, in which the test loss keeps decreasing while the trained model increasingly favors samples anomalously close to its training data. Training the same network on two disjoint datasets and comparing their outputs provides a simple quantitative measure of this bias, which we observe on real images well before any sign of overfitting. In a controlled hierarchical data model, where the exact denoiser is available, we trace its onset to the sequential nature of feature learning: coarse structure is learned early and independently of the samples, while finer features are resolved later in a way that depends on individual training points. Generalization and memorization thus behave as orthogonal rather than opposite axes, and early stopping may be an insufficient safeguard in privacy-critical applications.
Mariia Vladimirova (Criteo AI Lab): Position: Fairness Failure in Generative Models is an Evaluation Problem
Despite groundbreaking advancements in generative models during the last decade, concerns about their lack of fairness, reinforcing societal inequalities and harming marginalized groups, remain under-addressed and difficult to act upon. This position paper argues that fairness failures in generative models, albeit driven by multiple factors, are ultimately stemming from an evaluation problem: fairness findings are rarely comparable across papers or actionable for deployment decisions. This paper diagnoses recurring empirical and conceptual failure modes in current practice and motivates a shift from ad-hoc bias checks to standardized, generative-specific evaluation. We propose Fairness Cards as a minimal reporting artifact that makes evaluation choices explicit (prompt families, counterfactual protocols, metrics, and refusal handling) enabling reproducibility, comparability, and accountability. We conclude with additional recommendations towards a paradigm shift in evaluation standards.
