Section 2: The general principle of generative models
Pick a parametric family, define a divergence, and minimise it using only samples.
I am working through Mathematical Foundations of Generative AI, Prof. Prathosh AP’s public lectures. The playlist is the spine. These notes are my deep dive on each section: a visual when the picture is the point, a detour when a prerequisite is doing real work, and the formula in my own words.
Previously: Section 1: The dataset and the problem.
Section 2: The general principle of generative models
Section 1 left me with a vague goal: estimate and learn to sample from it. This is the recipe that turns that goal into three moves. Every later model in the lectures is this recipe with different choices plugged in.
The statement I am carrying:
(i) Assume a parametric family on , denoted , represented using deep neural networks. (ii) Define and estimate a divergence between and . (iii) Solve an optimisation problem over the parameters of to minimise that divergence.
The example that makes it concrete: draw , pass it through a network , and call the distribution of by the name . Then solve
After that, sampling and computing gives samples from , which is as close to as this family could get.
Four prerequisites sit under that formula: the standard Gaussian , what happens when a random variable is pushed through a function, what a divergence is, and the notation. Each detour is only as long as the step that needs it.
Intuition
I cannot see . I build a machine with knobs, , that turns plain random noise into data-shaped outputs. I keep turning the knobs until the machine’s output distribution is as close as I can get to the real one. The closeness is a divergence. The turning is optimisation.
The load test, continued
The traffic generator from Section 1 takes random seeds and outputs synthetic requests. It has config knobs: endpoint weights, payload-size parameters, and so on.
| Load-testing analogy | Notation |
|---|---|
| Random seed fed into the generator | |
| The generator code | , a neural network |
| Its config knobs | , the network weights |
| The distribution of traffic it produces | |
| A score of how different synthetic traffic is from real logs | the divergence |
| Tuning knobs until that score is near zero |
The hard step is the score. I cannot see the true user behaviour. I only have the log. That is the difficulty the rest of the lectures keep returning to.
Visual
One dimension is enough to see the recipe. The noise is . The network is the simplest one I can write, , so . The green curve is the target. The blue curve is what this generator can produce.
Three things I want in front of me while I move the sliders:
- On the single bump, and drive the divergence to 0. The target is a bell centred at 2 with spread 0.6, and is exactly that bell. The machine reproduces .
- On the two bumps, the divergence never reaches 0, however I set . The closest blue curve is a wide bell laid over both humps. A linear can only shift and stretch one bell. It cannot make two. That is why the principle asks for a deep network: has to be flexible enough that can get near .
- The chart can draw the green curve, and it can compute the divergence, because I built the target in. In the real problem I do not get that picture. I only get samples. That is the first of the four open questions below.
Detours
The standard Gaussian . is the bell centred at with spread . Standard means and . In dimensions, is independent standard bells, one per coordinate. is the identity covariance: each coordinate has variance 1, and no two coordinates are correlated.
It is the input noise because every library can sample it, and it has no structure of its own. Any structure in the output has to come from . So is a random seed: easy to draw, carrying nothing about images.
Pushing a random variable through a function. If is random and , then is random too, with a new distribution determined by .
The simplest case: if , then . Multiplying by stretches the spread. Adding slides the centre. That is what the sliders above are doing.
A deep, nonlinear can bend the output into more than one hump, or into a thin set such as the distribution of faces. The output distribution is different from the distribution of , and it depends on . I name that output distribution . For a deep network I usually cannot write a formula for it. I only know how to sample it: draw , run .
Divergence, which is weaker than a distance. scores how different two distributions are. The two properties I need:
It need not be symmetric. may differ from . That is why the notation uses rather than a comma. The KL divergence in the chart is one example. The next note builds the family they all belong to: Section 3a: f-divergences.
Those two properties are what make “minimise the divergence” a sensible goal. The smallest possible value is 0, and it is reached only when the model matches the data.
. is the smallest value of . is the input that achieves it. For , the min is 1 and the argmin is 3. I want the knob settings, not the score, so the principle uses argmin.
The formula
Read aloud: is the setting of the network weights that makes the divergence between the true data distribution and the distribution of the network’s outputs as small as possible.
| Symbol | What it is | Type / shape | Role |
|---|---|---|---|
| input noise | vector in , usually | the random seed | |
| standard Gaussian | distribution on | easy-to-sample source of randomness | |
| generator | function , a neural net | reshapes noise into data | |
| generator parameters | a vector of weights | the knobs | |
| model distribution | distribution on , implicit | what the generator produces | |
| true distribution | distribution on , unknown | the target | |
| divergence | two distributions to a scalar | the “how far apart” score | |
| optimal parameters | same shape as | the trained generator |
Once training is done, sampling is: draw a fresh , compute . That is a sample from . I have learned to sample from without writing it down.
A worked example small enough to do by hand
Take and , so .
- Draw three noise values: .
- Outputs: , , .
- By the pushforward above, all outputs come from . Centre 3, spread 2, which matches these samples.
For the divergence, a two-outcome world is enough to compute by hand. Let be a fair coin, , and use KL, . KL is formally the next part of the lectures. I only need it here to see numbers move.
- If the model says : .
- If the model says : .
- If the model says : both terms are , so the divergence is exactly 0.
As the model gets closer to the truth, the score falls toward 0, and it hits 0 only at a perfect match. Optimisation is moving in the direction that lowers this number.
Four questions this recipe does not answer
The rest of the lectures is the attempt to answer these.
| Question | Why it is hard | Where I expect the answer |
|---|---|---|
| How do I compute the divergence without formulas for or ? | I only have samples of both | Write the divergence as expectations, then estimate them with sample averages |
| Which divergence? | Different choices behave differently. The wide bell over two bumps is one failure mode | f-divergences first, Wasserstein later |
| How do I choose , and so ? | Flexible enough, and still trainable | GANs, VAEs, diffusion, transformers |
| How do I solve the optimisation? | Millions of parameters, often a min-max game | Gradient descent, then its adversarial variants |
Where the later models sit
| Model | Choice of | Divergence | Optimisation |
|---|---|---|---|
| GAN | , implicit | Jensen–Shannon, an f-divergence | min-max with a discriminator |
| WGAN | , implicit | Wasserstein | min-max with a 1-Lipschitz critic |
| VAE | latent model | KL, through the ELBO | maximise the ELBO |
| DDPM | reverse chain of Gaussians | KL, through the ELBO | regression on the added noise |
| Autoregressive / transformer | KL, which is maximum likelihood | cross-entropy |
One link I want early: minimising is the same as maximum likelihood, because . That is why VAEs, diffusion models, and language models can all train on log-likelihood losses and still be this same recipe.
Questions I want to be able to answer:
- Why can no setting of match the two-bump target, and what has to change about ?
- After training, how do I produce 10 new samples, in one line?
- If , what follows? What if the divergence is small but not zero?
Next: Section 3a: f-divergences.