Skip to content
DineshKumar Sarangapani
All writing

Section 2: The general principle of generative models

Pick a parametric family, define a divergence, and minimise it using only samples.

I am working through Mathematical Foundations of Generative AI, Prof. Prathosh AP’s public lectures. The playlist is the spine. These notes are my deep dive on each section: a visual when the picture is the point, a detour when a prerequisite is doing real work, and the formula in my own words.

Previously: Section 1: The dataset and the problem.

Section 2: The general principle of generative models

Section 1 left me with a vague goal: estimate PxP_x and learn to sample from it. This is the recipe that turns that goal into three moves. Every later model in the lectures is this recipe with different choices plugged in.

The statement I am carrying:

(i) Assume a parametric family on PxP_x, denoted PθP_\theta, represented using deep neural networks. (ii) Define and estimate a divergence between PxP_x and PθP_\theta. (iii) Solve an optimisation problem over the parameters of PθP_\theta to minimise that divergence.

The example that makes it concrete: draw z∼N(0,I)z \sim \mathcal{N}(0, I), pass it through a network gθ:Z→Xg_\theta : \mathcal{Z} \to \mathcal{X}, and call the distribution of x=gθ(z)x = g_\theta(z) by the name PθP_\theta. Then solve

θ∗=arg⁡min⁡θ  D(Px ∥ Pθ)\theta^* = \arg\min_\theta \; \mathcal{D}(P_x \,\|\, P_\theta)

After that, sampling z∼N(0,I)z \sim \mathcal{N}(0, I) and computing gθ∗(z)g_{\theta^*}(z) gives samples from Pθ∗P_{\theta^*}, which is as close to PxP_x as this family could get.

Four prerequisites sit under that formula: the standard Gaussian N(0,I)\mathcal{N}(0, I), what happens when a random variable is pushed through a function, what a divergence is, and the arg⁡min⁡\arg\min notation. Each detour is only as long as the step that needs it.

Intuition

I cannot see PxP_x. I build a machine with knobs, θ\theta, that turns plain random noise into data-shaped outputs. I keep turning the knobs until the machine’s output distribution is as close as I can get to the real one. The closeness is a divergence. The turning is optimisation.

The load test, continued

The traffic generator from Section 1 takes random seeds and outputs synthetic requests. It has config knobs: endpoint weights, payload-size parameters, and so on.

Load-testing analogyNotation
Random seed fed into the generatorz∼N(0,I)z \sim \mathcal{N}(0, I)
The generator codegθg_\theta, a neural network
Its config knobsθ\theta, the network weights
The distribution of traffic it producesPθP_\theta
A score of how different synthetic traffic is from real logsthe divergence D(Px ∥ Pθ)\mathcal{D}(P_x \,\|\, P_\theta)
Tuning knobs until that score is near zeroθ∗=arg⁡min⁡θD\theta^* = \arg\min_\theta \mathcal{D}

The hard step is the score. I cannot see the true user behaviour. I only have the log. That is the difficulty the rest of the lectures keep returning to.

Visual

One dimension is enough to see the recipe. The noise is z∼N(0,1)z \sim \mathcal{N}(0, 1). The network is the simplest one I can write, gθ(z)=az+bg_\theta(z) = a z + b, so θ=(a,b)\theta = (a, b). The green curve is the target. The blue curve is what this generator can produce.

Three things I want in front of me while I move the sliders:

  1. On the single bump, a=0.6a = 0.6 and b=2b = 2 drive the divergence to 0. The target is a bell centred at 2 with spread 0.6, and x=0.6z+2x = 0.6z + 2 is exactly that bell. The machine reproduces PxP_x.
  2. On the two bumps, the divergence never reaches 0, however I set (a,b)(a, b). The closest blue curve is a wide bell laid over both humps. A linear gθg_\theta can only shift and stretch one bell. It cannot make two. That is why the principle asks for a deep network: gθg_\theta has to be flexible enough that PθP_\theta can get near PxP_x.
  3. The chart can draw the green curve, and it can compute the divergence, because I built the target in. In the real problem I do not get that picture. I only get samples. That is the first of the four open questions below.

Detours

The standard Gaussian N(0,I)\mathcal{N}(0, I). N(μ,σ2)\mathcal{N}(\mu, \sigma^2) is the bell centred at μ\mu with spread σ\sigma. Standard means μ=0\mu = 0 and σ=1\sigma = 1. In kk dimensions, N(0,I)\mathcal{N}(0, I) is kk independent standard bells, one per coordinate. II is the k×kk \times k identity covariance: each coordinate has variance 1, and no two coordinates are correlated.

It is the input noise because every library can sample it, and it has no structure of its own. Any structure in the output has to come from gθg_\theta. So z∼N(0,I)z \sim \mathcal{N}(0, I) is a random seed: easy to draw, carrying nothing about images.

Pushing a random variable through a function. If zz is random and x=g(z)x = g(z), then xx is random too, with a new distribution determined by gg.

The simplest case: if z∼N(0,1)z \sim \mathcal{N}(0, 1), then x=az+b∼N(b,a2)x = az + b \sim \mathcal{N}(b, a^2). Multiplying by aa stretches the spread. Adding bb slides the centre. That is what the sliders above are doing.

A deep, nonlinear gg can bend the output into more than one hump, or into a thin set such as the distribution of faces. The output distribution is different from the distribution of zz, and it depends on gθg_\theta. I name that output distribution PθP_\theta. For a deep network I usually cannot write a formula for it. I only know how to sample it: draw zz, run gθg_\theta.

Divergence, which is weaker than a distance. D(P ∥ Q)\mathcal{D}(P \,\|\, Q) scores how different two distributions are. The two properties I need:

D(Px ∥ Pθ)≥0,D(Px ∥ Pθ)=0  ⟺  Px=Pθ\mathcal{D}(P_x \,\|\, P_\theta) \ge 0, \qquad \mathcal{D}(P_x \,\|\, P_\theta) = 0 \iff P_x = P_\theta

It need not be symmetric. D(P ∥ Q)\mathcal{D}(P \,\|\, Q) may differ from D(Q ∥ P)\mathcal{D}(Q \,\|\, P). That is why the notation uses ∥\| rather than a comma. The KL divergence in the chart is one example. The next note builds the family they all belong to: Section 3a: f-divergences.

Those two properties are what make “minimise the divergence” a sensible goal. The smallest possible value is 0, and it is reached only when the model matches the data.

arg⁡min⁡\arg\min. min⁡θf(θ)\min_\theta f(\theta) is the smallest value of ff. arg⁡min⁡θf(θ)\arg\min_\theta f(\theta) is the input θ\theta that achieves it. For f(θ)=(θ−3)2+1f(\theta) = (\theta - 3)^2 + 1, the min is 1 and the argmin is 3. I want the knob settings, not the score, so the principle uses argmin.

The formula

θ∗=arg⁡min⁡θ  D(Px ∥ Pθ),Pθ:=distribution of gθ(z),  z∼N(0,I)\theta^* = \arg\min_\theta \; \mathcal{D}\big(P_x \,\|\, P_\theta\big), \qquad P_\theta := \text{distribution of } g_\theta(z),\; z \sim \mathcal{N}(0, I)

Read aloud: θ∗\theta^* is the setting of the network weights that makes the divergence between the true data distribution and the distribution of the network’s outputs as small as possible.

SymbolWhat it isType / shapeRole
zzinput noisevector in Rk\mathbb{R}^k, usually k≪dk \ll dthe random seed
N(0,I)\mathcal{N}(0, I)standard Gaussiandistribution on Rk\mathbb{R}^keasy-to-sample source of randomness
gθg_\thetageneratorfunction Rk→Rd\mathbb{R}^k \to \mathbb{R}^d, a neural netreshapes noise into data
θ\thetagenerator parametersa vector of weightsthe knobs
PθP_\thetamodel distributiondistribution on Rd\mathbb{R}^d, implicitwhat the generator produces
PxP_xtrue distributiondistribution on Rd\mathbb{R}^d, unknownthe target
D(⋅ ∥ ⋅)\mathcal{D}(\cdot \,\|\, \cdot)divergencetwo distributions to a scalar ≥0\ge 0the “how far apart” score
θ∗\theta^*optimal parameterssame shape as θ\thetathe trained generator

Once training is done, sampling is: draw a fresh z∼N(0,I)z \sim \mathcal{N}(0, I), compute gθ∗(z)g_{\theta^*}(z). That is a sample from Pθ∗≈PxP_{\theta^*} \approx P_x. I have learned to sample from PxP_x without writing it down.

A worked example small enough to do by hand

Take k=d=1k = d = 1 and gθ(z)=2z+3g_\theta(z) = 2z + 3, so θ=(a,b)=(2,3)\theta = (a, b) = (2, 3).

  • Draw three noise values: z=−1,0,1z = -1, 0, 1.
  • Outputs: gθ(−1)=1g_\theta(-1) = 1, gθ(0)=3g_\theta(0) = 3, gθ(1)=5g_\theta(1) = 5.
  • By the pushforward above, all outputs come from Pθ=N(3,22)P_\theta = \mathcal{N}(3, 2^2). Centre 3, spread 2, which matches these samples.

For the divergence, a two-outcome world is enough to compute by hand. Let PxP_x be a fair coin, (0.5,0.5)(0.5, 0.5), and use KL, ∑iPx(i)ln⁡Px(i)Pθ(i)\sum_i P_x(i) \ln \frac{P_x(i)}{P_\theta(i)}. KL is formally the next part of the lectures. I only need it here to see numbers move.

  • If the model says Pθ=(0.9,0.1)P_\theta = (0.9, 0.1): 0.5ln⁡0.50.9+0.5ln⁡0.50.1=0.5(−0.588)+0.5(1.609)≈0.5110.5 \ln\frac{0.5}{0.9} + 0.5 \ln\frac{0.5}{0.1} = 0.5(-0.588) + 0.5(1.609) \approx 0.511.
  • If the model says Pθ=(0.6,0.4)P_\theta = (0.6, 0.4): 0.5ln⁡0.50.6+0.5ln⁡0.50.4=0.5(−0.182)+0.5(0.223)≈0.0200.5 \ln\frac{0.5}{0.6} + 0.5 \ln\frac{0.5}{0.4} = 0.5(-0.182) + 0.5(0.223) \approx 0.020.
  • If the model says Pθ=(0.5,0.5)P_\theta = (0.5, 0.5): both terms are 0.5ln⁡1=00.5 \ln 1 = 0, so the divergence is exactly 0.

As the model gets closer to the truth, the score falls toward 0, and it hits 0 only at a perfect match. Optimisation is moving θ\theta in the direction that lowers this number.

Four questions this recipe does not answer

The rest of the lectures is the attempt to answer these.

QuestionWhy it is hardWhere I expect the answer
How do I compute the divergence without formulas for PxP_x or PθP_\theta?I only have samples of bothWrite the divergence as expectations, then estimate them with sample averages
Which divergence?Different choices behave differently. The wide bell over two bumps is one failure modef-divergences first, Wasserstein later
How do I choose gθg_\theta, and so PθP_\theta?Flexible enough, and still trainableGANs, VAEs, diffusion, transformers
How do I solve the optimisation?Millions of parameters, often a min-max gameGradient descent, then its adversarial variants

Where the later models sit

ModelChoice of PθP_\thetaDivergenceOptimisation
GANgθ(z)g_\theta(z), implicitJensen–Shannon, an f-divergencemin-max with a discriminator
WGANgθ(z)g_\theta(z), implicitWassersteinmin-max with a 1-Lipschitz critic
VAElatent model ∫Pθ(x∣z) p(z) dz\int P_\theta(x \mid z)\, p(z)\, dzKL, through the ELBOmaximise the ELBO
DDPMreverse chain of GaussiansKL, through the ELBOregression on the added noise
Autoregressive / transformer∏tPθ(xt∣x<t)\prod_t P_\theta(x_t \mid x_{<t})KL, which is maximum likelihoodcross-entropy

One link I want early: minimising KL(Px ∥ Pθ)\mathrm{KL}(P_x \,\|\, P_\theta) is the same as maximum likelihood, because arg⁡min⁡KL(Px ∥ Pθ)=arg⁡max⁡E[log⁡Pθ(x)]\arg\min \mathrm{KL}(P_x \,\|\, P_\theta) = \arg\max \mathbb{E}[\log P_\theta(x)]. That is why VAEs, diffusion models, and language models can all train on log-likelihood losses and still be this same recipe.

Questions I want to be able to answer:

  1. Why can no setting of (a,b)(a, b) match the two-bump target, and what has to change about gθg_\theta?
  2. After training, how do I produce 10 new samples, in one line?
  3. If D(Px ∥ Pθ)=0\mathcal{D}(P_x \,\|\, P_\theta) = 0, what follows? What if the divergence is small but not zero?

Next: Section 3a: f-divergences.