Skip to content
DineshKumar Sarangapani
All writing

Section 1: The dataset and the problem

The hidden distribution, the dataset I actually get, and why generative modeling starts there.

I am working through Mathematical Foundations of Generative AI, Prof. Prathosh AP’s public lectures. The playlist is the spine. These notes are my deep dive on each section: a visual when the picture is the point, a detour when a prerequisite is doing real work, and the formula in my own words.

Section 1: The dataset and the problem

Everything later in the playlist, GANs, VAEs, diffusion, transformers, state-space models, and preference tuning, starts from the same setup. I want this part solid before I touch a model.

The setup I am carrying:

Data D={x1,x2,…,xn}∼iidPxD = \{x_1, x_2, \dots, x_n\} \overset{\text{iid}}{\sim} P_x (unknown), with xi∈Rdx_i \in \mathbb{R}^d, where dd is the dimensionality of the data. Example: xix_i is an image, 400×400×3400 \times 400 \times 3, so d=480,000d = 480{,}000. Goal: estimate PxP_x and learn to sample from it.

Intuition

Somewhere there is a hidden recipe, a probability distribution, that produced every cat photo, every English sentence, every speech clip. I never get the recipe. I get a pile of examples it produced. Generative modeling is using that pile to reconstruct the recipe well enough to cook up new examples from the same kitchen.

An analogy I can check

Think of a production web service whose request process I cannot inspect. All I have is an access log of 1,000 requests. I want a load-testing tool that produces synthetic traffic indistinguishable from real users: same mix of endpoints, same payload sizes, same timing.

Load-testing analogyNotation
The real, unknown user behaviourPxP_x, the true data distribution
One logged requestxix_i, one data point
The fields in a log linethe dd coordinates of xi∈Rdx_i \in \mathbb{R}^d
The access logdataset DD
The traffic generatorthe generative model PθP_\theta
Synthetic requests it emitssamples from PθP_\theta

Two consequences. I do not need to write the user-behaviour formula down. I need a generator whose output is statistically indistinguishable. That is why a GAN never writes PxP_x down at all. And replaying the log is not the goal. I want new requests, not copies.

Visual

A hidden 2-D distribution plays the role of PxP_x. At first I only see the dots, the dataset. Reveal PxP_x to see what is normally invisible, and drag nn to see how more data reveals the shape.

Each “Draw a new dataset” gives a different DD and the same green PxP_x underneath. The dataset is random. The distribution is the fixed thing I am after. With n=5n = 5 the shape is a guess. With n=1000n = 1000 the two clumps are obvious.

Detours I needed before the formula

Random variable, and random vector. A random variable is a quantity whose value is decided by chance, like a die roll X∈{1,…,6}X \in \{1,\dots,6\}. A random vector is the same idea with several numbers at once, such as height and weight of a randomly chosen person, X∈R2X \in \mathbb{R}^2. Capital XX means the random thing in general. Lowercase xix_i means one specific value that actually came out.

Distribution and density. PxP_x tells me how likely each possible value is. For a die it is a table: each face has probability 1/6. For continuous values, such as pixel intensities, it is a density px(x)p_x(x), a heat map over space. High density means values land there often. In the chart above, the green shading is that density.

Two facts I will need constantly:

px(x)≥0∫Rdpx(x) dx=1p_x(x) \ge 0 \qquad \int_{\mathbb{R}^d} p_x(x)\, dx = 1

The second says all probability mass integrates to 1. I will use PxP_x loosely for both the distribution and its density. That is standard, and I will be explicit when the difference matters.

iid, independent and identically distributed.

  • Identically distributed: every xix_i comes from the same PxP_x. Every log line came from the same user population, not some from production and some from a test environment.
  • Independent: knowing x3x_3 tells me nothing extra about x7x_7. The first roll of a die does not influence the second.

In notation, xi⊥xjx_i \perp x_j and xi∼Pxx_i \sim P_x.

The formula

D={x1,x2,…,xn}∼iidPx,xi∈RdD = \{x_1, x_2, \dots, x_n\} \overset{\text{iid}}{\sim} P_x, \qquad x_i \in \mathbb{R}^d

Read aloud: the dataset DD is a collection of nn points. Each point is an independent draw from the same unknown distribution PxP_x. Each point is a vector of dd real numbers.

SymbolWhat it isType / shapeRole
DDthe dataseta set of nn vectorsthe only thing I actually observe
nnnumber of samplesintegermore data, a better picture of PxP_x
xix_ione data pointvector in Rd\mathbb{R}^dan image, a sentence embedding, a signal
dddimensionalityintegerhow many numbers describe one point
∼\sim”is drawn from”relationlinks data to its source
iidindependent and identically distributedassumptionmakes learning from samples legitimate
PxP_xtrue data distributiondistribution on Rd\mathbb{R}^dthe hidden target, unknown

Why “unknown” is the whole problem. If I knew PxP_x, I would sample from it and stop. Every model in the playlist is a different trick for working with PxP_x through samples alone.

Why iid matters. It licenses replacing an expectation, an average over PxP_x, with an average over the dataset. That is the law of large numbers:

1n∑i=1nh(xi)  → n→∞   Ex∼Px[h(x)]\frac{1}{n}\sum_{i=1}^n h(x_i) \;\xrightarrow{\,n\to\infty\,}\; \mathbb{E}_{x\sim P_x}[h(x)]

That line is the engine behind every training loss later. The chart shows a small version: the sample mean settles as nn grows.

Why dd is so large. A 400×400 RGB image is one point in R480,000\mathbb{R}^{480{,}000}, one coordinate per pixel per colour channel. Real images occupy a thin region of that space. Random pixel vectors look like television static. This is the manifold hypothesis, and it comes back when naive GAN training falls over.

Two jobs hiding in one goal

Given DD, estimate PxP_x and learn to sample from it. Those are different jobs, and different models emphasise different ones.

JobMeaningWho emphasises it
Density estimationanswer “how likely is this xx?“autoregressive models, VAEs through the ELBO
Samplingproduce a new x∼Pxx \sim P_xGANs, which never write a density down, and diffusion

A GAN can sample well and still be unable to tell me p(x)p(x) for a given image. I want to remember that split.

A worked example small enough to do by hand

Let d=2d = 2 and n=3n = 3. Each point is hours slept and coffees drunk for a random person:

D={x1,x2,x3}={(7,1), (6,2), (8,0)}D = \{x_1, x_2, x_3\} = \{(7, 1),\ (6, 2),\ (8, 0)\}

Each xi∈R2x_i \in \mathbb{R}^2, so d=2d = 2. Here x2=(6,2)x_2 = (6, 2) means 6 hours of sleep and 2 coffees. iid means three different people, surveyed separately, from the same population.

I want E[coffees]\mathbb{E}[\text{coffees}] under the unknown PxP_x. I cannot integrate against a distribution I do not have. iid lets me estimate it from the sample: 13(1+2+0)=1\tfrac{1}{3}(1 + 2 + 0) = 1.

The generative goal is a machine that outputs new pairs such as (7.5,0.5)(7.5, 0.5): plausible, not copied from DD, and following the same pattern. More sleep tends to go with less coffee. That dependency between coordinates is what PxP_x captures, and a per-column average does not.

The same picture with d=480,000d = 480{,}000 and nn in the millions is image generation.

Where the rest of the playlist sits, as I understand it now

StretchModelHow it attacks “learn PxP_x from DD“
EarlyGAN / divergence minimisationpush z∼N(0,I)z \sim \mathcal{N}(0, I) through gθg_\theta and shrink an f-divergence to PxP_x
NextVAElatent-variable Pθ(x)=∫Pθ(x,z) dzP_\theta(x) = \int P_\theta(x, z)\,dz, maximise the ELBO
ThenWGANthe same game, with the Wasserstein distance
ThenDDPMadd noise to x0x_0, learn to reverse it
ThenAutoregressive / transformerfactorise Pθ(x)=∏tPθ(xt∣x<t)P_\theta(x) = \prod_t P_\theta(x_t \mid x_{<t})
ThenSSM / Mambaanother sequence model
LateRLHF / DPOmove the learned distribution toward preferences

The next section turns the goal into a recipe: pick a parametric family PθP_\theta, define a divergence, and minimise it using only samples.

Questions I want to be able to answer before I move on:

  1. If a dataset mixes 5,000 phone photos and 5,000 medical scans, which part of iid breaks, and why does that hurt?
  2. What is dd for a 28×28 grayscale digit? For a 64×64 RGB image?
  3. Once I have DD, do I know PxP_x exactly? What changes as n→∞n \to \infty?

Next: Section 2: The general principle of generative models.